Let's Start With the Rejection Letter
Four fictional groups of candidates. Every one with a 90% chance of succeeding in the job. Forty rounds of hiring, thirty runs per condition. And by the end, the model had decided this group belonged in one sort of role and that group belonged somewhere else.
Princeton and the University of Chicago ran GPT, Claude and Gemini through 40 rounds of simulated hiring with four fictional candidate groups, each with an identical 90% success probability. The models still sorted them unequally, inventing stereotypes from their own decision history rather than inheriting them from training data. A couple of early wins with one group, a pattern read into noise, then the pattern reinforcing itself round after round. The paper's line is that large language models "can also invent novel biases that influence human and agent behavior". The researchers tried every intervention a sensible buyer would ask for. Reason step by step. Turn up the randomness. Cut how much of its own history the model can see. But nothing moved. The only thing that worked was instructing the model to optimise explicitly for diversity, and that pushed allocations towards random even where groups genuinely were not equal. Which puts a hole in the reassurance most of us have been accepting. Bias audits, fairness benchmarks, the compliance pack that comes with the platform: all of it was designed to catch bias that arrived WITH the data. The bias in this study was generated on the premises, mid-process, by a system doing what it was built to do. The researchers are careful about scope, and so should we. The tests used synthetic scenarios and invented labels, not live enterprise systems. This is no proof that any particular product is discriminating today. It is narrower and…