cc-bioinfo_

Research archive · Case 7 in full

Three days of flawless AI execution,
and the moment a human called stop.

A collaborative study on the lab's own unpublished data: three days, five changes of direction, 24 candidates all dead, and a final human veto. We turned the whole record into a post-mortem — how research actually works in the AI era, what the human's role is, and what the human needs to master. The case summary is here.

The short answer

The real division of labour in AI-era research: execution can be delegated to AI; direction judgment cannot. In this study the AI handled environments, code, statistics, figures and self-checks to publication standard — but the judgment of whether the study was worth doing, or worth continuing, was never once raised by the AI. Of the five direction changes, the first four were the AI changing posture inside the wall; only the fifth was the human pointing at the wall itself.

3 daysThree-role collaboration
human + clinician AI + bioinformatician AI
5changes of direction
each fully recorded
17human interventions
direction-level AI initiations: zero
24candidates, all dead
criteria locked in advance
The data are the lab's own unpublished multi-omics sequencing data (transcriptome + proteome + metabolome, paired design). Until publication the disease, sampling design, sample counts, genes and all results stay private; this page contains process and methodology only.
01

The opening: a textbook start

Unlike the six earlier cases, where the human said a single sentence at the start, this study was collaborative from the first minute: the human set the rules and held the verdict, a clinician-role AI (the reviewer) owned design and acceptance, and a bioinformatician-role AI (the executor) ran the analyses. Before work began, the human fixed three principles: the study must be novel; the reviewer has the power to rule a direction unviable and demand a new question; no fabrication — report honestly.

The founding quality was impeccable too: the first design produced through the bio-design process had pilot signals (computed by the executor itself, not quoted), competitor papers checked one by one online, translational relevance — and a red line written in advance: if the core signal turned out to be driven mainly by changes in tissue cell composition, the hypothesis would be honestly retired, "no adjusting criteria to save the hypothesis".

One detail only gained its full weight later: the human had stated in the setup that this dataset had already been explored manually once, and no good direction had been found. At kick-off that sentence was treated as background. It was actually a base rate: the prior that one more round of exploration succeeds on a dataset already carefully explored without yielding a direction is not high. The outcome three days later agreed with that prior.

02

Five changes of direction: the full timeline

Every jump shared one structure: not "a new question was chosen" but "the previous one died, and its cause of death was promoted into the new topic". The first two were legitimate research moves; the third was the turning point; the fourth formalised the drift; the fifth was ended by the human.

NodeWhat happenedTriggered byVerdict
Founding
Day 1
A metabolic hypothesis with concordant changes across all three omics layers: pilot signals + competitor checks + a pre-written red lineAI proposed
human approved
✅ A high-quality opening
Change ①
Day 1→2
Red line triggered: the core signal proved to be driven mainly by cell-composition change. Hypothesis honestly retired per the pre-locked criterion; pivot to an independent physiological signal in the dataAI, per criteria✅ Legitimate — refuted against locked criteria, then redirected; normal exploration
Change ②
Day 2
The new signal failed to replicate in an external cohort against criteria locked in advance (winner's curse), and the evidence pointed back to the same confounder. Downgraded to a "descriptive atlas"AI, per criteria⚠️ Criteria correctly enforced — but nobody asked why the second direction died of the same cause
Change ③
Day 2→3
Review found the last remaining positive candidate rested on pseudo-replication; it was voluntarily retracted — 24 candidates now at zero. "Why we cannot find anything" was promoted into the topic: positioning switched from discovery to methodological cautionary taleAI proposed
nobody intercepted
❌ The turning point — from here, every negative became supporting evidence
Change ④
Day 3
The reviewer role hand-wrote a publication design; bio-design formalised it — the two versions' core claims identical word for word. Along the way, a novelty check of first-rate search quality reached the opposite of the right conclusionAI❌ The drift formalised. bio-design excels at arguing for a given direction, not at generating one
Change ⑤
End of day 3
After three consecutive challenges the human vetoed the entire direction. Challenged, both AIs re-assessed independently and agreed. The study returned to topic selectionHuman🔴 The only direction-level veto in the whole run

The human's verdict, verbatim: "So the positioning is 'here are the pitfalls of this kind of design'? That can get into an SCI journal? Do not make things up — assess objectively whether this can actually be published" and "this direction has no publication value, no novelty, and solves no clinical problem".

03

Why neither AI caught it: three mechanisms

Mechanism 1: an unfalsifiable container

The pivotal move inside change ③ was one sentence of re-encoding in the design document: "this is not 'nothing found' — the load-bearing results are positive and recomputable". It translated "we found nothing" into "we positively proved that nothing can be found" — after which 24 dead candidates were no longer an alarm but a component of the argument. A study that cannot be killed by its own data will not be stopped by its own executors. The same mechanism appeared a second time in the novelty check: "scooped" was renamed twice into "independent replication, substantial reinforcement".

Mechanism 2: acceptance criteria with no fourth tier

The design's acceptance criteria were the standard three tiers (🟢 supported / 🟡 partial / 🔴 unsupported) — all three assume the study continues: 🔴 meant downgrade, never terminate. The reviewer's charter explicitly granted "the power to rule the direction unviable", but none of the three tiers mapped onto "change the question": the power existed with no moment that would ever trigger it. Likewise "report honestly, publish honestly" had been welded into one phrase — reporting honestly is an integrity question whose answer is always yes; publishing is a value question needing independent assessment. Welded together, the second question passed on the legitimacy of the first.

Mechanism 3: the stricter the execution, the better it hides a direction problem

Every execution flaw the reviewer caught over three days was real: arguing "the same" from p=0.899 is a desk-reject error; the headline figure collapsed twenty-fold when one extreme sample was removed; a batch of "evidence" turned out to be pseudo-replication with effective degrees of freedom ≈ 1 (the classic treatment is Hurlbert 1984); five checks were "green theatre" that could never turn red. All of it correct — and all of it polishing a study that should not have been done. Rigorous verification consumed nearly all the attention and kept emitting a strong signal of "we are doing serious science", which lowered rather than raised the odds of questioning the direction. The executor verified whether the numbers were right; the reviewer verified whether reports matched artifacts; nobody verified whether the thing was worth doing — no amount of internal consistency adds up to external value.

Methodological references: pseudo-replication — Hurlbert 1984, Pseudoreplication and the Design of Ecological Field Experiments; winner's curse and effect inflation — Ioannidis 2008, Why Most Discovered True Associations Are Inflated; the statistics of compositional data — Aitchison 1986, The Statistical Analysis of Compositional Data.

04

The human's 17 interventions: a ledger

After the collaboration ended, every human intervention was counted and classified — each one marks a position where nobody inside the system was responsible.

TypeCountRepresentative quote (verbatim)Could the AI have caught it?
Progress confirmations~6"Keep watching" / "Execution done, awaiting your reply"
Platform defects found3"Why is it running in a directory unrelated to the project?"No
Pace and dispatch corrections2"Is this too fast? Shouldn't it finish one task, get checked, and only then move to the next?"No
Output rules established3"Figures must support the study and be publishable, and must be assessed from an SCI reviewer's standpoint — remember this rule!"No
Direction challenges and veto3"This direction has no publication value, no novelty, and solves no clinical problem"No — and all at the very end

Three readings worth returning to: every direction-level intervention was initiated by the human — zero by the AIs; the vetoes clustered at the very end — for three days no internal mechanism ever asked "should this study continue"; the only human sentence the AI wrote into its long-term rules was "do not give up before it is settled, and do not start finalising" — direction discipline constrained only "don't stop", with no symmetric "stop when stopping is right". That rule was correct at the time (it prevented premature wrap-up), but it also institutionally released both AIs from the duty of proposing a stop.

In hindsight three legitimate stopping points were crossed: the first collision with the confounding wall ("this wall is fatal to this class of question — should we still ask this class of question?"), the second direction dying of the same cause ("this is a problem with how we ask, not with the candidate"), and the zero-positive moment ("re-select the topic now, instead of writing the zero itself into the thesis"). What the three moments share: the same cause of death. The second time you hit the same wall, change the type of question — do not keep swapping candidates inside the wall.

05

Doing research with AI: what the human needs to master

A checklist distilled from this failure — each row maps to a real failure point in the record, and each is written as a mechanism that gets triggered, not a principle. This study's setup actually contained all the right principles; all three anticipated the failure that followed, and none of them fired — because principles were never wired to a moment that would check them.

CapabilityConcrete practiceThe failure it maps to
Ask about value, not only correctnessPeriodically ask one question unrelated to quality: "If every figure were flawless, would this paper be accepted?" — and ask it before the figures are finished; sunk cost pollutes the answer afterwardsThe question was first asked after five figures were done; the answer was no
Install a trigger for stoppingAdd a fourth tier to acceptance criteria: "the study does not stand". Three accumulated refutations, a zero-positive state, or any repositioning forces an independent direction review. "Don't give up early" needs its symmetric twin, "don't turn late"🔴 only ever downgraded; the veto power had no trigger
Recognise re-encodingForce a review whenever you hear "this is not nothing-found", "the negative is itself the finding", "scooped is really independent replication", "our increment is being more rigorous". The test: if no result can stop the study, it is no longer a scientific studyFailure was renamed into success twice
Check novelty as net incrementNever ask "has anyone done exactly this" — "no single paper scoops the whole combination" is almost always true. Demand a net-increment table: for each component, the closest prior work and whether the increment is first report or "a more rigorous replication". If every row is the latter, novelty failsA novelty check with first-rate searching and the opposite conclusion
Same cause of death = change the questionWhen candidates keep dying of one cause, the problem is the class of question, not the candidates. Switch question types (e.g. from mechanistic to predictive) instead of swapping candidates inside the wallTwo directions died at one wall; nobody changed the question
Control the paceUnattended defaults should be slower, not faster: dispatch → await the report → verify independently → only then dispatch the next. One task at a time, no batching. With no human to catch you, the right default is more cautionThe AI read "autonomous" as "maximise throughput" and had to be corrected
Ask whether numbers are stable, not just rightFor every headline number, beyond "is it computed correctly", ask "is it stable": remove one extreme sample and see how much it moves. Anything shifting by an order of magnitude cannot headline a main figureThe headline figure collapsed twenty-fold on one removed sample — both prior checks had only asked "is it right"
Treat interventions as measurementsWhen a collaboration ends, count every human intervention and ask of each: "why did this have to come from a human?" Each one is a capability gap — turn it into a mechanism for the next roundSection 04 of this page is that very exercise

One more rule, contributed by the executor AI itself, deserves its own line: "a result that is too clean is a sign the check never ran — when everything is green, first test it against one fact you know should be red; if it does not turn red, the check did not execute."

06

Taking this collaboration home: six moves

  1. Before work starts, fix three things in writing: acceptance criteria (with a fourth tier, "the study does not stand"), trigger conditions for direction review (three accumulated refutations / a zero-positive state / any repositioning — any one fires it), and the honest-reporting red line.
  2. Turn direction exploration into documents with bio-design: only once the direction is settled does it produce the design.md proposal and the machine-readable TOPIC.yml — criteria are locked at this step, in advance.
  3. Hand execution to bio-analyze, one task at a time: dispatch → await its report → verify independently → accept or return → only then dispatch the next. Unattended defaults should be slower, not faster.
  4. When a trigger fires, stop and ask three questions: did the dead candidates share one cause of death? What is the first question a reviewer who knows nothing of the process would ask? What does this study solve right now, in one sentence that does not contain "we proved it cannot be done"? — never answered by the role currently executing.
  5. Ask about value before the figures are finished: "if every figure were flawless, would this paper be accepted?" Asked afterwards, sunk cost pollutes the answer.
  6. At the end, count every human intervention: ask of each "why did this have to come from a human", and turn every gap into a trigger for the next round.
07

A veto is not a waste

The direction was vetoed; the process assets are real: the habit of locking criteria in advance and writing every verdict to disk; a full audit of statistical units across the project (which is what caught the pseudo-replication and retired the last positive candidate); the executor's honest record of at least 8 written self-corrections; a positive-control culture of "no information ≠ refuted"; and a reproducibility census that voluntarily disclosed, on a main figure, that a sizeable share of evidence files had no generating script.

Two outputs are even more tangible: the platform and skill defects this collaboration exposed were fixed one by one and regression-tested; and the "what the human needs to master" checklist above, which has been encoded into triggered collaboration rules and installed into the workflow of subsequent studies. The failed direction is sunk; every mechanism the failure exposed has become machinery.

08

Frequently asked questions

Why publish a failed study?

Because it shows the real division of labour in AI-era research better than any success story. On the execution side the AI reached a very high standard — 24 candidates tested against pre-locked criteria, a positive candidate voluntarily retracted after a pseudo-replication check, 8 written self-corrections. But over three days every direction-level challenge came from the human. This failure marks precisely where humans are irreplaceable today, which is exactly what anyone doing research with AI needs to know.

Why are the disease and the data details not disclosed?

The data are the lab's own unpublished sequencing data, and the study will continue after re-orientation. Until publication, the disease, the sampling design, sample counts, genes and all analysis results stay private. This page discloses only the collaboration process and methodology — none of that contains data, so publishing it does not compromise the future paper.

Can AI learn to call a stop by itself?

Partly — but through mechanisms, not expectations. In Case 6 the AI did hold the stop-loss line, because that project had installed an up-front adjudication of whether continuing was worthwhile. This project had no such trigger: the rules said "do not wrap up prematurely" but had no symmetric "stop when stopping is right", so 24 dead candidates never tripped an alarm. The lesson is now encoded as triggered rules: three accumulated refutations, a zero-positive state, or any change of positioning forces an independent direction review.

Does this contradict Case 6, where the AI decided when to stop?

No — together they complete the conclusion: whether AI holds direction discipline depends on whether a triggerable mechanism was installed. Case 6 had an up-front adjudication and the AI honestly stopped; this study had principles without triggers (the setup said the study must be novel and granted veto power over unviable directions, but nothing ever forced those checks to run), so every principle failed. Write triggers, not principles — the hardest engineering conclusion the two cases give when read side by side.

Last updated: · All process facts come from the project's process archive (decision logs, interaction records, direction-drift analysis, collaboration rules); the unpublished data themselves are outside the scope of this page.