Drapeau’s AI Deep Dive: The Control Lab Tested for Flavor

Published on: 

Pharma manufacturing spent 90 years learning that producing correct outputs isn't the same as catching your own errors. AI has only learned the first.

In June 1937, a salesman for the S. E. Massengill Company reported demand in the southern states for sulfanilamide in liquid form. The company's chief chemist found that it dissolved in diethylene glycol. The control laboratory tested the resulting preparation for flavor, appearance, and fragrance, and found it satisfactory.1

More than 100 people died, many of them children, across 15 states. The solvent was a close chemical relative of antifreeze.

The part that matters is not that the company was careless. The part that matters is that the control lab ran its tests. It ran the tests it had designed. It documented a passing result. The quality process as that company understood quality could not have caught the thing that killed people, because nobody had built it to ask whether the solvent was poison.

The 1938 Federal Food, Drug, and Cosmetic Act followed within a year.2 Proof of safety before marketing. That statute is still the basis of American drug regulation.

I have spent my career inside the machinery that grew out of that disaster and the ones after it: Deviation management, corrective and preventive action (CAPA), change control, audit trails, data integrity. It is unglamorous work, and most people outside the industry assume it is paperwork. It is not paperwork. It is the formal answer to a specific question, refined across 90 years and written largely in response to body counts. The question is this: How does an organization full of competent people acting in good faith discover that it is wrong, and then be compelled to change?

I have been reading the AI safety literature for two years with a growing and specific discomfort.

The field has built extraordinary systems for producing correct outputs and has not built the thing that makes a system correctable, and those are different objects with different architectures.

Correctness is a property of an output. Correctability is a property of an organization. You can have the first without the second for a long time, and the Massengill control lab is what that looks like from the inside on the day before it matters.

What a Correction Loop Actually Contains

CAPA is not one regulation. It is a concept that recurs across every regulated industry that has had to learn this lesson, codified differently in pharmaceutical quality systems, device quality systems, aviation, and nuclear power.3 Strip the vocabulary away and it has four parts. Each one is separately necessary, and the loop is worthless if any one is missing.

Detection. Something must surface the failure, and it has to surface failures the designers did not anticipate, because the anticipated ones are already handled.

Attribution. You have to be able to trace the failure to a cause. Not a correlate, a cause, specific enough to act on.

Mandated change. Somebody has to be obligated to fix it, on a timeline, with the fix documented and approved by someone who did not write it.

Effectiveness check. And then, later, you have to go back and demonstrate that the change actually worked. This is the step everyone wants to skip and the step regulators fixate on, because a corrective action nobody verified is a corrective action nobody made.

Now run current AI practice against those four.

Detection exists in weak form. Evaluations, benchmarks, red teaming (structured adversarial testing in which a team tries to make a model fail in ways its designers did not anticipate). These methods catch the anticipated failure modes, which is real value, and I do not want to dismiss it. They catch almost nothing that nobody thought to look for, and they are run by the same organization that built the system.

Attribution is the limiting step. Developers can often identify a failure mode. What they generally cannot do is trace a particular frontier-model output back to the training examples, parameter updates, or design decisions that produced it, with enough confidence to ground a corrective action. Partial methods exist in constrained settings, and data attribution is an active research field.4,5 Nothing in it yet resembles the traceability a serious quality system would require.

Mandated change is largely absent for general-purpose frontier models. The organization fixes problems when it decides to fix them, on its own schedule, by its own criteria. Sectoral rules are beginning to move: The European Union's AI Act imposes postmarket monitoring and serious-incident reporting duties on certain high-risk systems.6 None of it yet amounts to a public correction loop for model behavior.

Effectiveness checks are voluntary and private. Some providers do rerun evaluations, monitor production incidents, and build tests from failures they discovered. What does not exist is a standard, auditable, externally enforceable protocol requiring a provider to demonstrate that a defined corrective action eliminated the original failure rather than relocating it.

The loop is open at three points out of four.

You Cannot Correct What You Cannot Trace

This is the part I know best and the part I think the AI field has most underestimated.

Data integrity requirements in pharmaceutical manufacturing run on a principle usually abbreviated ALCOA: Records must be attributable, legible, contemporaneous, original, and accurate. The extended version adds complete, consistent, enduring, and available. FDA and the United Kingdom's Medicines and Healthcare products Regulatory Agency both published guidance formalizing these principles within the past decade,7,8 and the framework drives an enormous amount of how regulated manufacturing actually operates.

Attributable comes first for a reason. It is not a bureaucratic nicety. If a batch fails and you cannot determine which operator, which equipment, which lot of raw material, and which procedural step produced the failure, then you cannot fix it. You can only throw the batch away and hope. Everything downstream in the correction loop depends on attribution, which is why the industry spends staggering amounts of money on audit trails that most days nobody reads.

Frontier models are not currently traceable at the granularity a corrective system would require. Gradient descent over trillions of tokens distributes the influence of training data across billions of parameters in a way current methods cannot reliably reconstruct for a particular output.5,9

This means the AI field is operating with the requirement my industry treats as foundational not merely unmet but, at frontier scale, not yet technically satisfiable.

I want to be careful here, because this is the point at which a generalist starts telling engineers how to do their jobs, and I am not going to do that. I do not know how to make a large model attributable. It may not be possible in the current paradigm. What I can say with confidence, from 90 years of accumulated regulatory experience, is what happens to an industry that cannot attribute its failures. It cannot close the correction loop at the cause. It adds output filters, retrains, or substitutes one mechanism for another, and it cannot demonstrate that it fixed the thing that produced the failure. It builds elaborate inspection at the output stage and calls that quality, which is precisely what the Massengill control lab was doing.

Nothing Gets Recalled

Science has a retraction mechanism. It is slow, it is weak, it is embarrassing to everyone involved, and it exists.

Pharmaceutical manufacturing has one too, and it is considerably better. A Class I recall requires notification, retrieval, reconciliation of what was distributed against what came back, and a public record.10 FDA publishes enforcement reports weekly.11 If a batch is bad, there is a defined process for reaching the product already in the world.

A provider can withdraw a model, disable a feature, publish a correction, or notify customers who use its application programming interface, and providers do all of these. None of that is a recall.

If a model asserts something false to 10 million people, there is no way to identify who received it. There is no way to retrieve the documents, decisions, and code it entered. There is no reconciliation of what went out against what came back. There is no durable public record of the event. The next version simply does not make that error, or makes it less often, and nobody is told about the previous one.

Advertisement

Isaac Newton's alchemy at least had the decency to sit in a drawer.

The Part Where My Own Industry Fails

I would not trust this essay if it did not include this section, and you should not either.

Quality systems fail. They fail in a specific and recognizable way, and I have watched it happen: They become theater. Documentation that demonstrates compliance rather than producing quality. Deviations closed by writing a more persuasive narrative rather than finding a cause. Root cause analyses that arrive at human error, which is not a root cause, but the place investigations go when they are being terminated early. Effectiveness checks performed as a formality against criteria loose enough to pass.

The failure mode is that the correction loop keeps running as ritual after it has stopped detecting anything. The forms get filled out. The signatures get collected. Nothing gets found.

Red teaming and system cards (published documents in which a developer reports a model's capabilities, limitations, and evaluation results; also called model cards) are on that trajectory right now, and I say that as someone who wants them to work. The moment an evaluation becomes the thing you must pass rather than the thing that tells you what is broken, it stops being detection and becomes documentation. My industry took roughly 40 years to learn that lesson and still relearns it during every inspection cycle. AI is almost seven years into the same arc and moving faster.

There is a second failure worth naming. Correction machinery has a conservative bias. It makes change expensive, and expensive change suppresses good ideas along with bad ones. Pharmaceutical manufacturing is genuinely slower and more risk-averse than it needs to be in places, and some of that cost lands on patients who wait. Anyone who tells you correction architecture is free is selling something.

What I Think Should Exist

Both of my previous essays on this subject—"Newton's Other Million Words" and "Calculus Ran on Gossip"—ended by refusing to prescribe. That refusal was honest then. It would be evasion here, so here is the claim.

The correct unit of oversight for AI systems is not the model. It is the quality system that produces and maintains the model.

This principle is the single most useful thing my industry learned, and it took the deaths to learn it. You cannot inspect quality into a product at the end. Testing every tablet is impossible and would not work anyway, because the test only asks the questions you thought to ask. Pharmaceutical oversight therefore cannot rest on testing individual tablets at the end. It makes the manufacturing process, and the evidence that the process remains in a state of control, a central object of regulation in its own right.3,12

Evaluations are release testing. They are the control lab checking flavor, appearance, and fragrance. Useful, necessary, and categorically incapable of catching the thing nobody specified.

Three things follow, and I will state them as concretely as I can.

A model provider should maintain a formal deviation record. When a systematic failure is identified, whether internally or externally, it gets an entry: what failed, what was investigated, what change was made, and what evidence shows the change worked. Not a blog post. A record, retained, with a defined retention period, available to an auditor.

There should be an external body with the authority to compel investigation. Not to approve models, which I think would fail for reasons the field's critics have argued well,13 but to require that a credible report of systematic failure be investigated and answered. Self-reported safety work is the pre-1938 arrangement, and we know how that arrangement performs.

And attribution should be treated as a research priority with the weight the industry currently gives capability. Not because interpretability is intellectually interesting, though it is, but because every other part of a correction loop is load-bearing on it. A field that cannot attribute its failures cannot compound its fixes. It can only accumulate, which brings me back to where the previous essay ended.

I hold these positions with different confidence. The first I am confident about, because it is mostly a matter of discipline and costs little. The second I am less sure of, because regulatory capture is real and a badly designed body is worse than none. The third is a research bet I am not qualified to evaluate on technical grounds, only on the grounds that everything else depends on it.

What I Am Not Claiming

I am not claiming pharmaceutical quality systems should be ported to AI. The domains differ in ways that matter. Drug manufacturing produces physical units in discrete batches with defined specifications; a language model produces an unbounded space of outputs with no specification at all. Much of what makes GMP work depends on that discreteness and does not survive the translation.

I am not claiming my industry has this solved. We do not. We have a system that is better than what preceded it, purchased at a price paid in advance by people who died, and it still fails regularly enough to keep me employed.

And I am not claiming correction machinery produces truth. It does not. It produces the capacity to stop being wrong in a particular way twice, which is a smaller thing and the only thing on offer.

Harold Watkins, the chemist who formulated the elixir, took his own life before the trial. Samuel Massengill said publicly that there was no error in the manufacture of the product and that he felt no responsibility on the part of the company. Under the law as it stood in 1937, he was largely correct. No statute required safety testing of a new drug, and the only charge available was misbranding, because the word elixir implied an alcohol-based solution.1

That is what it looks like to build a system that produces outputs without building the system that catches them. Everyone follows the procedure. The procedure passes. The procedure was never asking the right question, and nobody finds out until the thing has already left the building.


References

1. Ballentine C. Taste of raspberries, taste of death: the 1937 Elixir Sulfanilamide incident. FDA Consumer. June 1981. Accessed October 1, 2026. View source

2. Wax PM. Elixirs, diluents, and the passage of the 1938 Federal Food, Drug and Cosmetic Act. Ann Intern Med. 1995;122(6):456-461. doi:10.7326/0003-4819-122-6-199503150-00009. View source

3. International Council for Harmonisation. ICH Q10: pharmaceutical quality system. ICH.org. Published June 2008. Accessed October 1, 2026. View source

4. Koh PW, Liang P. Understanding black-box predictions via influence functions. In: Proceedings of the 34th International Conference on Machine Learning. Vol 70. PMLR; 2017:1885-1894. Accessed October 1, 2026. View source

5. Grosse R, Bae J, Anil C, et al. Studying large language model generalization with influence functions. arXiv. Preprint posted online August 7, 2023. Accessed October 1, 2026. doi:10.48550/arXiv.2308.03296. View source

6. European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), articles 72-73. Official Journal of the European Union. Published July 12, 2024. Accessed October 1, 2026. View source

7. US Food and Drug Administration. Data integrity and compliance with drug CGMP: questions and answers; guidance for industry. FDA.gov. Published December 2018. Accessed October 1, 2026. View source

8. Medicines and Healthcare products Regulatory Agency. Guidance on GxP data integrity. GOV.UK. Published March 9, 2018. Updated September 27, 2021. Accessed October 1, 2026. View source

9. Basu S, Pope P, Feizi S. Influence functions in deep learning are fragile. In: International Conference on Learning Representations. ICLR; 2021. Accessed October 1, 2026. View source

10. Product recalls, including removals and corrections, 21 CFR Part 7. eCFR. Accessed October 1, 2026. View source

11. US Food and Drug Administration. Enforcement reports. FDA.gov. Accessed October 1, 2026. View source

12. Current good manufacturing practice for finished pharmaceuticals, 21 CFR §211. eCFR. Accessed October 1, 2026. View source

13. Kapoor S, Narayanan A. Licensing is neither feasible nor effective for addressing AI risks. AI Snake Oil. June 10, 2023. Accessed October 1, 2026. View source