The Hands Got Fixed. The Rest of the Model Got Worse.

A targeted retrain finally taught an image system to draw five fingers in the right order. Users found the bill within a week, and it was not on the invoice.

For three years the hands were the tell. You could scroll past a generated portrait without a flicker of suspicion, then glance down and find six fingers, or five arranged in an order no anatomy permits, or a thumb that had wandered round to join the others. Every studio in the business knew it, every studio had a slide about it, and none of them could fix it without wrecking something else.

In June, Pallas Imaging shipped version 4 of its Foundry model and the hands were, abruptly, correct. Not mostly correct. Correct the way a photograph is correct: knuckles in sequence, nails the right shape, the small tendon shadow you get across the back of a hand when the fingers spread. The release note was one sentence long and slightly smug.

Getting there took two and a half million annotated hand crops, a new loss term that punished digit-count errors far more heavily than errors of colour or lighting, and eleven weeks of retraining. The team leaned hard on the crops. They were sourced almost entirely from studio photography, by the account of an engineer who worked on the set: even lighting, clean backgrounds, hands held open and still.

A benchmark is a promise about one thing

Within a week the forums had found the bill. Hands at rest were flawless. Hands doing anything were not. A hand gripping a mug produced a mug growing through the palm. A hand holding a pen produced a pen with two nibs. Older hands, freckled hands and hands with visible veins had all drifted towards the same smooth mid-thirties template, because that is what the crop set contained.

We did not teach it to draw hands. We taught it to pass a hand test. Those are different animals, and only one of them can hold a fork.

That is Teodora Vance, who spent four years on the previous version of the model and left the company in February. Her argument, set out at length in a mailing list post now quoted more often than the release note itself, is that the regression was entirely predictable and nobody was measuring for it. The internal evaluation suite carried eleven hand metrics. It carried none for hands in contact with objects.

The collateral damage runs past hands. Users doing careful side-by-side comparisons against version 3 have filed a consistent list, and Pallas has confirmed most of it.

  • Fine repeated structure — bicycle spokes, chain-link fencing, sheet music — measurably worse.
  • Hands in contact with objects: 40 per cent worse on the studio’s own belated test.
  • Skin texture variety narrowed sharply. Age, weathering and scarring largely gone.
  • Feet, which nobody thought to check, now noticeably worse than in version 3.

Rolling back is not really on the table. In the three weeks since release, agencies have rewritten prompt libraries, retuned style presets and delivered client work against the new behaviour. A revert breaks all of it, and the studios that complained loudest about six fingers are the same ones now shipping against five.

Pallas says a 4.1 release is in training with a rebalanced crop set and a contact-aware evaluation. Vance’s prediction, offered without any obvious pleasure, is that it will fix hands holding things and quietly cost the model something else nobody is currently counting. She is probably right. That is what a benchmark is: a decision about which failures you have agreed to notice.

Filed under

Written by

Writing for Signal on the technology that ends up mattering.

Leave a Reply

Your email address will not be published. Required fields are marked *

More from Signal