Merlin published a blinded, citation by citation evaluation of seven OpenAI, Anthropic and Google models available in Alchemy, each writing cited reports from a public opioid litigation document collection (Mallinckrodt, McKesson, Insys, Valeant). An eight pass automated grader checked 777 citations and 454 cited sentences, audited its own verdicts with two graders from other providers, and found no fabricated facts; six of seven models were statistically tied on accuracy, and the most expensive model cost over nine times the cheapest without ranking higher. The article also names GPT-6 Astra and Claude Fable 5.1 as the newest models added to Alchemy, without giving an addition date.
Buyer relevance. Buyers get a disclosed method for checking citation grounding and fabrication across the models they can pick inside Alchemy, which supports choosing a cheaper model for citation heavy reports. The test is vendor run and model names are masked, so buyers should ask for the answer key and the attached full report.