Study Design: The Answer Is Chosen Before the Question Is Asked
Stage × AI, issue 1 of ~38. One stage per issue: what the stage really does, where AI helps, where it must not, and one buildable pattern.
Confidence: hypothesis. The failure modes here are drawn from ICH E9(R1)'s own rationale for the estimand framework and widely reported industry patterns, not firsthand deployment. The buildable pattern at the end is not: it is an unbuilt design. Argue with it.

What this stage really does
Before Issue 2 writes a protocol, someone decides what trial to run: the endpoint and how it is measured, the comparator, the population, the duration, the randomization scheme Issue 9 will execute, the estimand: the precise question, down to what happens when patients take rescue medication, discontinue treatment, or experience other intercurrent events. ICH E9(R1) gave that last decision a name and a framework precisely because, in the addendum's own words, “avoiding or over-simplifying the process of discussing and constructing an estimand risks misalignment between trial objectives, trial design, data collection and method of analysis,” and because “it is necessary to address intercurrent events when describing the clinical question of interest.” That's exactly the readout-stage debate this essay opens with.
Two properties define the stage. First, design fixes the answer-space. Every downstream stage in this series, twenty-three essays of execution discipline, can only deliver the answer the design made available. A trial with the wrong comparator, a too-short duration, or an endpoint the disease does not move is not rescued by flawless conduct; it is a perfectly manufactured answer to a question nobody needed. Execution errors are recoverable in principle. Design errors are the trial.
Second, design decisions are near-irreversible bets made at maximum ignorance. The moment of greatest freedom, before first patient in, is the moment of least data. Changing an endpoint or population mid-study costs amendments, credibility, and sometimes the trial; so the stage's real discipline is forcing the expensive arguments forward, into the room where they are still cheap.
Consider a readout meeting, years after design, where the primary result hinged on how to treat patients who had taken rescue medication. Two camps formed: treat the intercurrent event one way, the effect stands; the other way, it shrinks toward noise. Both sides argued from the protocol; the protocol was silent, because the question had never been asked when asking was free. The “analysis debate” was archaeology: excavating a design decision that was never made, under the worst possible conditions, after everyone could compute which answer their preferred method produced. Issue 28's decide-before-seeing discipline exists exactly because of rooms like that one; the estimand framework is that discipline pushed all the way to the start.
The failure currency of the stage: trials that answer the wrong question precisely, endpoints chosen by precedent rather than by the disease, designs whose complexity fails at real sites (Issue 16's burden is designed here), and the end-of-study estimand debate, always the most expensive place to hold it.
Where AI helps
The pattern holds: AI helps where the work is precedent mining, simulation, and stress-testing wearing a scientific-judgment costume: the judgment stays human; the evidence base under it gets much bigger.
Design precedent assembly. What endpoints, durations, comparators, and populations did the last hundred trials in this indication use, and what happened: regulatory reception, enrollment reality, effect sizes seen. Cited, so the design team argues from the field's actual record instead of its remembered highlights.
Trial simulation. The design run as a digital twin before it runs as a trial: virtual populations against candidate designs, power under realistic dropout (Issue 3's calibrated priors, not hopeful ones), sensitivity to enrollment mix, operating characteristics of adaptive rules. The Issue 10 wind-tunnel move, aimed at the design itself.
Estimand stress-testing. For each candidate estimand: enumerate the intercurrent events this population actually generates (from precedent and real-world data), and check the design answers the question under each. The rescue-medication debate, held at design time, with a model generating the awkward cases nobody volunteers.
Complexity and burden scoring. Visits, procedures, decision points per patient and per site, scored against what comparable designs achieved operationally: the design reviewed as something sites and patients must live, before Issue 16 pays for it.
Regulatory precedent retrieval. What agencies said about similar designs (advice letters, precedent approvals, published disputes), assembled with citations for the team deciding how much novelty to spend where.

None of this designs the trial. All of it makes the expensive arguments happen while they are still cheap.
Where it must not go
Choosing the question. What is worth studying, in whom, against what, these are scientific and ethical commitments owned by the people accountable for them. A model can enumerate options and their precedent; the moment it ranks “best design” it is substituting an objective function nobody examined for a judgment everybody must defend.
Optimizing for approval odds alone. Precedent mining has a dark gradient: designs converge on whatever passed before, safest endpoint, narrowest population, likeliest yes. Followed blindly, it is a machine for scientific timidity, optimizing the industry toward questions already answered.
Simulating with flattering priors. A trial simulation is only as honest as its dropout, enrollment, and effect assumptions, and every one of those has an advocacy-friendly setting. Simulations carry their priors on their face, sourced and Issue 3-calibrated, or they are rhetoric with confidence intervals.
Deferring the estimand. Any tooling that lets the team ship a design with “analysis details to follow” has automated the war story. The intercurrent-event questions are answered at design time, in the design record, by name. Silence is not a decision. It is a debate scheduled for the worst possible moment.

One buildable pattern: the design pre-mortem
The stage's cheapest hour is before commitment. The pattern spends it deliberately.

- Simulate against honest priors. Every candidate design runs in the wind tunnel with Issue 3-calibrated dropout and enrollment priors, sources attached. A design that only works under flattering assumptions is rejected by its own simulation record.
- Enumerate the intercurrent events. From precedent and real-world data: what this population does: rescue, discontinuation, death, crossover. Each one gets an estimand answer, at design time, logged. The rescue-medication debate, held for pennies.
- Argue against the precedent record. The assembled field record is the sparring partner: where this design deviates from precedent, the deviation is deliberate and documented; where it follows precedent, someone has checked the precedent was right.
- Log every contested choice. Endpoint, comparator, population, duration: the options considered, the reasoning, the dissent, by name: the Issue 2 decision-log flywheel starting one stage earlier, where the decisions are largest.
- Rehearse the readout. Write both topline summaries, the win and the miss, before first patient in, against the simulated data. If the team cannot agree what the miss would mean, the design is not done; that disagreement is the war story, caught early.
What to measure first: estimand questions answered at design versus surfaced at analysis (the war story's metric), simulation prior accuracy scored against the trial's actuals (the Issue 3 loop closing on design), amendment count traceable to design gaps, design-record completeness when Issue 33's questions arrive, and the honesty metric: readout debates that were really unmade design decisions, counted and named as such. Until those numbers exist somewhere public, this stays labeled a hypothesis.
Three Starting Points
A word on what this section is: drawn from public vendor materials and industry reporting on where governed-data and multi-agent tooling are heading, not from any employer's actual or planned build. Treat it as a map of a category, not a build order or an endorsed architecture.
The pattern above treats precedent mining, simulation, and estimand stress-testing as three separate passes run by one generalist tool. That was reasonable when each pass meant a different bespoke script. The infrastructure to run them as one coordinated system now ships as a platform feature, not a research prototype.
Greenfield. Multi-agent delegation, a lead agent decomposing a scope and routing sub-tasks to specialist workers running in parallel, is now a mainstream, shipped feature: Anthropic, OpenAI, AWS, Microsoft, and Google have each shipped or expanded a version of it in the last year, with IBM's still in private preview at last public update; OpenAI is far enough along that it's already retiring its first-generation agent-building tools in favor of a second, and AWS has closed its first-generation offering to new customers and moved it to maintenance mode (references below). Applied to the design pre-mortem, that turns three sequential passes into three concurrent ones: one worker mining precedent for a candidate endpoint, one running the simulation sweep across honest dropout and enrollment priors, one enumerating intercurrent events from precedent and real-world data, reconciled by a lead agent that flags where the three disagree, not one that averages them into false consensus. The precedent corpus itself belongs in a governed catalog with lineage back to the trial it came from, the same shift already underway for the TMF (Issue 33) and the protocol (Issue 2).
Reinventing. The assumption worth dropping industry-wide: that a trial simulation is a one-time gate run once before the protocol is finalized, then archived. Organizations experimenting with continuous simulation re-run the design's operating characteristics every time a precedent search or an intercurrent-event enumeration turns up something new, closer to how a test suite re-runs on every commit than how a power calculation traditionally gets signed off once and filed.
Traditional: don't know where to start. Don't start by wiring up multi-agent orchestration. Start by indexing the field's actual record, endpoints, comparators, populations, outcomes from the last hundred trials in the indication, into one queryable, sourced corpus. Precedent mining, simulation calibration, and estimand stress-testing are all retrieval on top of that index; none of them has anything real to check without it.

This is an architectural proposal, not a regulator-mandated stack. One industry-watcher's hypothesis, not a description of anyone's actual build.
References
Frontier labs and platforms (2025–2026, same bench as Issue 33):
- Anthropic, “How we built our multi-agent research system” (June 2025)
- OpenAI, “New tools for building agents” (March 2025). The Agents SDK replacing Swarm; “Introducing AgentKit” (October 2025). Agent Builder wind-down announced June 3, 2026, unavailable from November 30, 2026
- AWS, “Amazon Bedrock now supports multi-agent collaboration” (March 2025); “Amazon Bedrock AgentCore is now generally available” (October 2025); “Amazon Bedrock Agents Classic maintenance mode” . Classic Bedrock Agents closed to new customers July 30, 2026
- Microsoft, “New and improved multi-agent orchestration, connected experiences, and faster prompt iteration” . Copilot Studio multi-agent orchestration and A2A, general availability April 2026
- Google Cloud, “Build and manage multi-system agents with Vertex AI” (April 2025). The Agent Development Kit and Agent Engine, Google's lead/worker orchestration product; the same post introduces the Agent2Agent (A2A) protocol for cross-agent interoperability, a related but separate capability
- IBM Newsroom, “Think 2026: IBM Delivers the Blueprint for the AI Operating Model” (May 2026). watsonx Orchestrate's agentic control plane, in private preview at last public update
Regulatory sources
- ICH, “Addendum on Estimands and Sensitivity Analysis in Clinical Trials to the Guideline on Statistical Principles for Clinical Trials E9(R1)” (2019). The estimand framework and intercurrent-event enumeration this essay's failure modes and buildable pattern are drawn from.
- ICH, “E9: Statistical Principles for Clinical Trials” (1998). The parent guideline the R1 addendum extends; randomization, blinding, and the broader statistical-design framework this essay's decisions sit inside.
- U.S. Food and Drug Administration, “Assessing the Credibility of Computational Modeling and Simulation in Medical Device Submissions — Guidance for Industry and FDA Staff” (final, issued Nov 17, 2023). A device-CM&S credibility-assessment framework cited here as an adjacent public benchmark for the simulation-honesty discipline behind this essay's “Trial simulation” claim, not as a drug-trial design mandate.
- U.S. Food and Drug Administration, “Adaptive Designs for Clinical Trials of Drugs and Biologics — Guidance for Industry” (final, Nov 2019). The regulatory framework behind this essay's mention of “operating characteristics of adaptive rules.”
Last tended 2026-07-26 · Corrections welcome — evidence in, evidence out.
Under the Hood · 5
5 capability-angle deep-dives generalizing this essay's pattern for AI and software engineers.
A decision made before the outcome is visible can't be reverse-engineered from it.
Architecture Decision Records for choices whose reasoning would otherwise survive only in memory.
Retrieval that hunts for disagreement, not confirmation.
A simulation with a flattering input is a forecast that was told the answer.
A premortem is honest doubt with social permission attached.