Research archive · Case 7 in full
Three days of flawless AI execution,
and the moment a human called stop.
A collaborative study on the lab's own unpublished data: three days, five changes of direction, 24 candidates all dead, and a final human veto. We turned the whole record into a post-mortem — how research actually works in the AI era, what the human's role is, and what the human needs to master. The case summary is here.
The short answer
The real division of labour in AI-era research: execution can be delegated to AI; direction judgment cannot. In this study the AI handled environments, code, statistics, figures and self-checks to publication standard — but the judgment of whether the study was worth doing, or worth continuing, was never once raised by the AI. Of the five direction changes, the first four were the AI changing posture inside the wall; only the fifth was the human pointing at the wall itself.
human + clinician AI + bioinformatician AI
each fully recorded
direction-level AI initiations: zero
criteria locked in advance
The opening: a textbook start
Unlike the six earlier cases, where the human said a single sentence at the start, this study was collaborative from the first minute: the human set the rules and held the verdict, a clinician-role AI (the reviewer) owned design and acceptance, and a bioinformatician-role AI (the executor) ran the analyses. Before work began, the human fixed three principles: the study must be novel; the reviewer has the power to rule a direction unviable and demand a new question; no fabrication — report honestly.
The founding quality was impeccable too: the first design produced through the bio-design process had pilot signals (computed by the executor itself, not quoted), competitor papers checked one by one online, translational relevance — and a red line written in advance: if the core signal turned out to be driven mainly by changes in tissue cell composition, the hypothesis would be honestly retired, "no adjusting criteria to save the hypothesis".
One detail only gained its full weight later: the human had stated in the setup that this dataset had already been explored manually once, and no good direction had been found. At kick-off that sentence was treated as background. It was actually a base rate: the prior that one more round of exploration succeeds on a dataset already carefully explored without yielding a direction is not high. The outcome three days later agreed with that prior.
Five changes of direction: the full timeline
Every jump shared one structure: not "a new question was chosen" but "the previous one died, and its cause of death was promoted into the new topic". The first two were legitimate research moves; the third was the turning point; the fourth formalised the drift; the fifth was ended by the human.
| Node | What happened | Triggered by | Verdict |
|---|---|---|---|
| Founding Day 1 | A metabolic hypothesis with concordant changes across all three omics layers: pilot signals + competitor checks + a pre-written red line | AI proposed human approved | ✅ A high-quality opening |
| Change ① Day 1→2 | Red line triggered: the core signal proved to be driven mainly by cell-composition change. Hypothesis honestly retired per the pre-locked criterion; pivot to an independent physiological signal in the data | AI, per criteria | ✅ Legitimate — refuted against locked criteria, then redirected; normal exploration |
| Change ② Day 2 | The new signal failed to replicate in an external cohort against criteria locked in advance (winner's curse), and the evidence pointed back to the same confounder. Downgraded to a "descriptive atlas" | AI, per criteria | ⚠️ Criteria correctly enforced — but nobody asked why the second direction died of the same cause |
| Change ③ Day 2→3 | Review found the last remaining positive candidate rested on pseudo-replication; it was voluntarily retracted — 24 candidates now at zero. "Why we cannot find anything" was promoted into the topic: positioning switched from discovery to methodological cautionary tale | AI proposed nobody intercepted | ❌ The turning point — from here, every negative became supporting evidence |
| Change ④ Day 3 | The reviewer role hand-wrote a publication design; bio-design formalised it — the two versions' core claims identical word for word. Along the way, a novelty check of first-rate search quality reached the opposite of the right conclusion | AI | ❌ The drift formalised. bio-design excels at arguing for a given direction, not at generating one |
| Change ⑤ End of day 3 | After three consecutive challenges the human vetoed the entire direction. Challenged, both AIs re-assessed independently and agreed. The study returned to topic selection | Human | 🔴 The only direction-level veto in the whole run |
The human's verdict, verbatim: "So the positioning is 'here are the pitfalls of this kind of design'? That can get into an SCI journal? Do not make things up — assess objectively whether this can actually be published" and "this direction has no publication value, no novelty, and solves no clinical problem".
Why neither AI caught it: three mechanisms
Mechanism 1: an unfalsifiable container
The pivotal move inside change ③ was one sentence of re-encoding in the design document: "this is not 'nothing found' — the load-bearing results are positive and recomputable". It translated "we found nothing" into "we positively proved that nothing can be found" — after which 24 dead candidates were no longer an alarm but a component of the argument. A study that cannot be killed by its own data will not be stopped by its own executors. The same mechanism appeared a second time in the novelty check: "scooped" was renamed twice into "independent replication, substantial reinforcement".
Mechanism 2: acceptance criteria with no fourth tier
The design's acceptance criteria were the standard three tiers (🟢 supported / 🟡 partial / 🔴 unsupported) — all three assume the study continues: 🔴 meant downgrade, never terminate. The reviewer's charter explicitly granted "the power to rule the direction unviable", but none of the three tiers mapped onto "change the question": the power existed with no moment that would ever trigger it. Likewise "report honestly, publish honestly" had been welded into one phrase — reporting honestly is an integrity question whose answer is always yes; publishing is a value question needing independent assessment. Welded together, the second question passed on the legitimacy of the first.
Mechanism 3: the stricter the execution, the better it hides a direction problem
Every execution flaw the reviewer caught over three days was real: arguing "the same" from p=0.899 is a desk-reject error; the headline figure collapsed twenty-fold when one extreme sample was removed; a batch of "evidence" turned out to be pseudo-replication with effective degrees of freedom ≈ 1 (the classic treatment is Hurlbert 1984); five checks were "green theatre" that could never turn red. All of it correct — and all of it polishing a study that should not have been done. Rigorous verification consumed nearly all the attention and kept emitting a strong signal of "we are doing serious science", which lowered rather than raised the odds of questioning the direction. The executor verified whether the numbers were right; the reviewer verified whether reports matched artifacts; nobody verified whether the thing was worth doing — no amount of internal consistency adds up to external value.
Methodological references: pseudo-replication — Hurlbert 1984, Pseudoreplication and the Design of Ecological Field Experiments; winner's curse and effect inflation — Ioannidis 2008, Why Most Discovered True Associations Are Inflated; the statistics of compositional data — Aitchison 1986, The Statistical Analysis of Compositional Data.
The human's 17 interventions: a ledger
After the collaboration ended, every human intervention was counted and classified — each one marks a position where nobody inside the system was responsible.
| Type | Count | Representative quote (verbatim) | Could the AI have caught it? |
|---|---|---|---|
| Progress confirmations | ~6 | "Keep watching" / "Execution done, awaiting your reply" | — |
| Platform defects found | 3 | "Why is it running in a directory unrelated to the project?" | No |
| Pace and dispatch corrections | 2 | "Is this too fast? Shouldn't it finish one task, get checked, and only then move to the next?" | No |
| Output rules established | 3 | "Figures must support the study and be publishable, and must be assessed from an SCI reviewer's standpoint — remember this rule!" | No |
| Direction challenges and veto | 3 | "This direction has no publication value, no novelty, and solves no clinical problem" | No — and all at the very end |
Three readings worth returning to: ① every direction-level intervention was initiated by the human — zero by the AIs; ② the vetoes clustered at the very end — for three days no internal mechanism ever asked "should this study continue"; ③ the only human sentence the AI wrote into its long-term rules was "do not give up before it is settled, and do not start finalising" — direction discipline constrained only "don't stop", with no symmetric "stop when stopping is right". That rule was correct at the time (it prevented premature wrap-up), but it also institutionally released both AIs from the duty of proposing a stop.
In hindsight three legitimate stopping points were crossed: the first collision with the confounding wall ("this wall is fatal to this class of question — should we still ask this class of question?"), the second direction dying of the same cause ("this is a problem with how we ask, not with the candidate"), and the zero-positive moment ("re-select the topic now, instead of writing the zero itself into the thesis"). What the three moments share: the same cause of death. The second time you hit the same wall, change the type of question — do not keep swapping candidates inside the wall.
Doing research with AI: what the human needs to master
A checklist distilled from this failure — each row maps to a real failure point in the record, and each is written as a mechanism that gets triggered, not a principle. This study's setup actually contained all the right principles; all three anticipated the failure that followed, and none of them fired — because principles were never wired to a moment that would check them.
| Capability | Concrete practice | The failure it maps to |
|---|---|---|
| Ask about value, not only correctness | Periodically ask one question unrelated to quality: "If every figure were flawless, would this paper be accepted?" — and ask it before the figures are finished; sunk cost pollutes the answer afterwards | The question was first asked after five figures were done; the answer was no |
| Install a trigger for stopping | Add a fourth tier to acceptance criteria: "the study does not stand". Three accumulated refutations, a zero-positive state, or any repositioning forces an independent direction review. "Don't give up early" needs its symmetric twin, "don't turn late" | 🔴 only ever downgraded; the veto power had no trigger |
| Recognise re-encoding | Force a review whenever you hear "this is not nothing-found", "the negative is itself the finding", "scooped is really independent replication", "our increment is being more rigorous". The test: if no result can stop the study, it is no longer a scientific study | Failure was renamed into success twice |
| Check novelty as net increment | Never ask "has anyone done exactly this" — "no single paper scoops the whole combination" is almost always true. Demand a net-increment table: for each component, the closest prior work and whether the increment is first report or "a more rigorous replication". If every row is the latter, novelty fails | A novelty check with first-rate searching and the opposite conclusion |
| Same cause of death = change the question | When candidates keep dying of one cause, the problem is the class of question, not the candidates. Switch question types (e.g. from mechanistic to predictive) instead of swapping candidates inside the wall | Two directions died at one wall; nobody changed the question |
| Control the pace | Unattended defaults should be slower, not faster: dispatch → await the report → verify independently → only then dispatch the next. One task at a time, no batching. With no human to catch you, the right default is more caution | The AI read "autonomous" as "maximise throughput" and had to be corrected |
| Ask whether numbers are stable, not just right | For every headline number, beyond "is it computed correctly", ask "is it stable": remove one extreme sample and see how much it moves. Anything shifting by an order of magnitude cannot headline a main figure | The headline figure collapsed twenty-fold on one removed sample — both prior checks had only asked "is it right" |
| Treat interventions as measurements | When a collaboration ends, count every human intervention and ask of each: "why did this have to come from a human?" Each one is a capability gap — turn it into a mechanism for the next round | Section 04 of this page is that very exercise |
One more rule, contributed by the executor AI itself, deserves its own line: "a result that is too clean is a sign the check never ran — when everything is green, first test it against one fact you know should be red; if it does not turn red, the check did not execute."
Taking this collaboration home: six moves
- Before work starts, fix three things in writing: acceptance criteria (with a fourth tier, "the study does not stand"), trigger conditions for direction review (three accumulated refutations / a zero-positive state / any repositioning — any one fires it), and the honest-reporting red line.
- Turn direction exploration into documents with bio-design: only once the direction is settled does it produce the design.md proposal and the machine-readable TOPIC.yml — criteria are locked at this step, in advance.
- Hand execution to bio-analyze, one task at a time: dispatch → await its report → verify independently → accept or return → only then dispatch the next. Unattended defaults should be slower, not faster.
- When a trigger fires, stop and ask three questions: did the dead candidates share one cause of death? What is the first question a reviewer who knows nothing of the process would ask? What does this study solve right now, in one sentence that does not contain "we proved it cannot be done"? — never answered by the role currently executing.
- Ask about value before the figures are finished: "if every figure were flawless, would this paper be accepted?" Asked afterwards, sunk cost pollutes the answer.
- At the end, count every human intervention: ask of each "why did this have to come from a human", and turn every gap into a trigger for the next round.
A veto is not a waste
The direction was vetoed; the process assets are real: the habit of locking criteria in advance and writing every verdict to disk; a full audit of statistical units across the project (which is what caught the pseudo-replication and retired the last positive candidate); the executor's honest record of at least 8 written self-corrections; a positive-control culture of "no information ≠ refuted"; and a reproducibility census that voluntarily disclosed, on a main figure, that a sizeable share of evidence files had no generating script.
Two outputs are even more tangible: the platform and skill defects this collaboration exposed were fixed one by one and regression-tested; and the "what the human needs to master" checklist above, which has been encoded into triggered collaboration rules and installed into the workflow of subsequent studies. The failed direction is sunk; every mechanism the failure exposed has become machinery.
Frequently asked questions
Why publish a failed study?
Because it shows the real division of labour in AI-era research better than any success story. On the execution side the AI reached a very high standard — 24 candidates tested against pre-locked criteria, a positive candidate voluntarily retracted after a pseudo-replication check, 8 written self-corrections. But over three days every direction-level challenge came from the human. This failure marks precisely where humans are irreplaceable today, which is exactly what anyone doing research with AI needs to know.
Why are the disease and the data details not disclosed?
The data are the lab's own unpublished sequencing data, and the study will continue after re-orientation. Until publication, the disease, the sampling design, sample counts, genes and all analysis results stay private. This page discloses only the collaboration process and methodology — none of that contains data, so publishing it does not compromise the future paper.
Can AI learn to call a stop by itself?
Partly — but through mechanisms, not expectations. In Case 6 the AI did hold the stop-loss line, because that project had installed an up-front adjudication of whether continuing was worthwhile. This project had no such trigger: the rules said "do not wrap up prematurely" but had no symmetric "stop when stopping is right", so 24 dead candidates never tripped an alarm. The lesson is now encoded as triggered rules: three accumulated refutations, a zero-positive state, or any change of positioning forces an independent direction review.
Does this contradict Case 6, where the AI decided when to stop?
No — together they complete the conclusion: whether AI holds direction discipline depends on whether a triggerable mechanism was installed. Case 6 had an up-front adjudication and the AI honestly stopped; this study had principles without triggers (the setup said the study must be novel and granted veto power over unviable directions, but nothing ever forced those checks to run), so every principle failed. Write triggers, not principles — the hardest engineering conclusion the two cases give when read side by side.
Last updated: · All process facts come from the project's process archive (decision logs, interaction records, direction-drift analysis, collaboration rules); the unpublished data themselves are outside the scope of this page.