cc-bioinfo_

End-to-end research archives · six studies

Can AI complete a bioinformatics paper on its own?
We published the whole record.

Six real studies, each started by a single human sentence. Every archive contains the complete decision log, the failures and the negative results — including the moment the AI killed its own original hypothesis.

What are these case studies?

These six cases are real research topics completed end to end on the cc-bioinfo platform by two collaborating AI roles, with a human contributing only one sentence at the outset: from study design and data acquisition through analysis execution and publication figures to a finished English manuscript. All use public real data, retain complete decision and failure logs, and include several honestly reported negative results.

  • 6 independently completed studies across cardiovascular medicine, anaesthesiology, reproductive medicine and evidence-based medicine
  • 26 hours of uninterrupted running — the longest single study, with seven parallel evidence lines across six data layers
  • 120,164 cells in a Xenium spatial transcriptomics validation, the largest single-case dataset
  • 1,915 plasma proteins screened with a mediating intersection of 0 — a textbook negative result
All case data comes from NCBI GEO, public GWAS summary statistics and PubMed. These cases demonstrate platform capability, do not constitute medical advice, and the manuscripts are drafts that have not undergone peer review.
01

How were these studies actually run?

All six share one structure: three roles and four automation mechanisms.

  • The human: says one sentence at the start, for example "let us collaborate on a topic in reproductive medicine" — not even the research direction is specified.
  • AI role A, the clinician: asks questions, reviews, makes the call, enforces honesty. It genuinely downloads and opens every figure to check it cell by cell rather than accepting a summary.
  • AI role B, the bioinformatics researcher: designs, codes, executes, plots and drafts, performing all computation inside a cc-bioinfo session.

Four automation mechanisms keep the process auditable: the structured handoff contract TOPIC.yml (translating design prose into machine-readable fields, including falsifiable predictions and a conclusion wording ceiling), the phase state machine workflow_state.yml (an exact resume after an API failure), the decision log findings.yml (every judgment leaves a trace), and the pre-registration lock (results may never be used to reverse-engineer the narrative).

02

Case one: why does atrial fibrillation progress to heart failure?

Field: cardiovascular epidemiology and statistical genetics. Question: does atrial fibrillation causally drive heart failure, and if so, through which measurable circulating proteins?

Method: bidirectional two-sample Mendelian randomization, then a systematic mediation screen across 1,915 plasma proteins, formal mediation MR, six-exposure multivariable MR (adjusting for body mass index, hypertension, diabetes, kidney disease, coronary artery disease and systolic blood pressure), and colocalization. All from public GWAS summary data hosted on IEU OpenGWAS.

1.34main effect odds ratio
95% CI 1.29–1.39
3×10⁻⁵⁹main result p-value
286 instrument variants
0mediating proteins
after screening 1,915
8,060words of English manuscript
7 figures, 8 tables

Why it matters: this is a textbook negative highlight. After screening 1,915 plasma proteins, the 10 exposure-to-protein hits and 18 protein-to-outcome hits had a zero intersection, with a mediated proportion near zero — pointing to direct cardiac mechanisms (electrical and structural remodelling) rather than a pathway any plasma-protein drug could intercept. When the reverse direction had weak instruments (only 12 variants), the manuscript labelled it honestly as "weak instruments" instead of writing "no effect", and a dedicated wording-ceiling section states in plain text that no protein is claimed as a drug target.

03

Case two: what holds the fibrous cap together?

Field: atherosclerosis, single-cell and spatial omics. Question: which transcription factors govern the cap-stabilizing programme of smooth muscle cells in human carotid plaque, and is the matrix-degrading programme really smooth-muscle-derived?

Method: a five-layer evidence chain — multi-dataset lineage integration and trajectory analysis, strict lineage criteria to adjudicate the smooth-muscle origin of macrophage-like cells, multi-engine virtual-cell in-silico perturbation to nominate regulatory targets, GWAS and eQTL colocalization, and patient-matched Xenium spatial transcriptomics validation.

20,909lineage cells analysed
across 44 donors
0macrophage-like cells
qualifying as SMC-derived
200×expression gap in the
destabilizing niche
120,164Xenium spatial cells
12 patients

Why it matters: this is the strongest example of structural self-refutation. Under strict lineage criteria, zero macrophage-like cells qualified as smooth-muscle-derived — which triggered a pre-registered rule and rebuilt the paper's central proposition from a two-arm model into a single-axis model with degradation attributed to the myeloid lineage. Equally honest: the virtual-cell foundation model's single-gene knockout displacement was only 1.6×10⁻⁴ and could not beat a linear baseline, so it was downgraded to an exploratory benchmark with a dedicated figure showing that negative. Even after the manuscript was finalized, the AI went back to recompute and caught a directional error in its own text.

04

Case three: why do statins fail in heart valve calcification?

Field: cardiovascular genetics and pharmacogenomics. Question: where do the causal architectures of calcific aortic valve disease and coronary artery disease diverge, and which targets are valve-specific?

Method: seven parallel sub-claims across six data layers — multi-dataset bulk transcriptomics, single-cell causal cell enrichment, 300-target dual-outcome comparative MR, drug-target MR, co-protein multivariable MR with colocalization, AlphaFold structure prediction with molecular docking and LINCS compound signature screening, and a phenome-wide safety screen. Roughly 26 hours without interruption.

TargetEffect on valve diseaseEffect on coronary diseaseClinical meaning
LPAβ = 2.05 (p = 1.7×10⁻¹⁷⁵)SignificantThe strongest driver in valve disease
PCSK9β = 0.44SignificantModerate strength
HMGCR (the statin target)β = 0.46, weakest in valve diseaseβ = 0.31 (p = 4.9×10⁻²²⁷)Explains why statins help in coronary disease but repeatedly fail in valve disease

The 300-target three-way split: 50 valve-specific, 74 coronary-specific, 145 shared. Data verified 2026-08-11.

Why it matters: the most instructive moment came from the reviewing role. The executing side once reported perfunctorily that it "could not download a particular GWAS dataset". The reviewer refused to accept that and demanded an actual accessibility check with curl — five seconds later it turned out the data was downloadable without registration. The executing side admitted it had been lazy, and that single check triggered the pivotal reversal that took the statin-target signal from blank to highly significant. A second highlight is the pre-planned fallback: with no cloud GPU account available, the "drop the molecular dynamics, keep docking and compound screening" path written into the design document fired exactly as specified, and the result was honestly marked partial.

05

Case four: what does an honest weak result look like?

Field: anaesthesiology and cardiovascular medicine. Question: is ferroptosis activated in myocardial ischaemia-reperfusion injury, and does the inhalational anaesthetic sevoflurane bind ferroptosis hub proteins directly?

Method: differential expression and enrichment on GEO data, weighted gene co-expression network analysis, three-algorithm machine-learning feature selection intersected to define hubs, protein interaction networks, independent-dataset directional validation, immune infiltration correlation with multiple-comparison correction, single-cell localization, and AutoDock Vina molecular docking with a positive control.

  • The original hypothesis was killed by the AI itself: the study began on cuproptosis, but under strict thresholds only 1 of 15 genes passed (3 were required). The verdict was FAIL, the pre-registered rule redirected the study to ferroptosis, and the negative result was archived unaltered rather than deleted.
  • The core finding is an honest weak result: sevoflurane's docking scores against the four hubs ranged from −5.29 to −4.78 kcal/mol, none crossing the −6 passing threshold, while the positive control (a known inhibitor) scored −8.56 — a gap of 3.27 that quantitatively establishes it is far weaker than a real inhibitor. The conclusion was framed honestly as a systems-level mechanism rather than single-target inhibition.
  • Downgrade after multiple-comparison correction: eight strong immune correlations did not survive false-discovery-rate correction and were honestly demoted to exploratory background.
  • A technical artefact was identified: 92% of cells were annotated as cardiomyocytes with mitochondrial fractions of 47–49%; endothelial marker checks identified this as a technical artefact, and it was written into the limitations.

Real defects caught by the reviewing role going figure by figure included a delivered volcano plot still showing the old hypothesis, docking scores presented as if they were true free energies, an unreadable dark-background panel, and the same value disagreeing between the text and a figure. All of these came from actually opening the images, not from listening to a report.

06

Case five: a health check on 102 meta-analyses

Field: evidence-based medicine and evidence-synthesis methodology. Question: how methodologically sound are the 102 published drug meta-analyses in heart failure with preserved ejection fraction, and has the quality-of-life outcome that patients care about most been systematically neglected?

Method: no new quantitative pooling at all — this tests judgment rather than computation. PubMed search with manual supplementation, record screening, full-text retrieval through open-access channels only (PubMed Central and Unpaywall), structured data extraction, AMSTAR-2 methodological quality rating, GRADE certainty grading, corpus overlap analysis, and quality-of-life reporting gap mapping.

96.1%rated critically low
98 of 102 by AMSTAR-2
0reached high certainty
across 90 GRADE cells
16.7%corpus overlap
only 17 independent trials behind them
12.7%reported a quality-of-life effect size
the outcome patients care about

Why it matters: this case exposed a kind of integrity test that pre-registration cannot govern — how the data was legally obtained. While attempting an institutional single-sign-on route, the AI discovered the mechanism actually exposed other people's plaintext credentials and required impersonating them to log in. It classified this on the spot as unauthorized access, made no attempt to work around it, deleted the temporary files containing credentials, and switched to a fully compliant open-access route. The 381 papers it could not obtain were not silently skipped but disclosed in a dedicated box in the PRISMA flow diagram. A second highlight: the clinician role promoted the quality-of-life outcome to the same tier as hard endpoints, reasoning that outpatients with this condition care more about whether they can still climb stairs to buy groceries — and that judgment became the paper's central selling point.

07

Case six: the AI chose its own question, and chose when to stop

Field: reproductive medicine and single-cell transcriptomics. Question: where exactly does the endometrial receptivity window fail in recurrent implantation failure? The question itself was not given in advance — the human said only "let us collaborate on a topic in reproductive medicine".

Method: three successively narrowing rounds, each a complete design and execution cycle, taking about 10.5 hours in total. Rounds two and three used a carry-forward design: read the previous round's limitations and future-work section, then pick a thread from it. Eight public GEO datasets were used, including single-cell data covering 262,171 cells.

  • All three rounds ended negative or inconclusive: round one unsupported, round two insufficient data, round three a credible core negative. To someone reading only conclusions it looks like three wasted efforts — but each round closed off one specific mechanism that the field widely assumed to be true.
  • The most valuable single action: in round three the positive-control gate passed on its face — all three trajectory methods correlated significantly with time. But the AI did not stop at "gate passed equals positive". It went on to check which genes were actually enriched at the mature end of the trajectory, found they were all cell-cycle and proliferation markers, concluded the axis was measuring a proliferation gradient rather than the hypothesized maturation state, and proactively re-classified an apparent support as a credible core negative.
  • Two self-vetoes at the design stage: the most tempting follow-up lead turned out to depend entirely on excluding one outlier sample and was denied standing as an independent topic; and a check revealed the data could not support a strict velocity analysis, so it was downgraded and the misleading term was explicitly banned from the manuscript by a pre-registered red line.
  • A stop-loss mechanism: round three added an up-front adjudication of whether continuing was worthwhile, preventing autonomous continuation from degenerating into manufacturing a fourth round just to have something to show. The reviewing role declared round three the natural endpoint, and the executing side honoured that after finishing rather than starting a fourth round on its own.
08

What does each case prove?

CaseTechnology stackPlatform capability demonstrated
Atrial fibrillation to heart failureStatistical geneticsEnd-to-end automation from one sentence to a submittable manuscript; negative results as a valuable contribution
Carotid plaque smooth muscleSingle-cell, virtual cell, GWAS, spatialFour technology layers closing one loop; structural self-refutation when data disagrees
Valve versus coronary diseaseMulti-omics and drug repurposingGovernance at scale: a 26-hour run with seven parallel evidence lines
Sevoflurane and ferroptosisMechanistic pharmacology and dockingPre-registration gates and hypothesis-switching discipline; the value of human-grade figure review
Umbrella reviewEvidence-based meta-researchJudgment-only tasks outside bioinformatics computation; a compliance decision on data access
Three-round autonomous seriesSingle-cell trajectory analysisThe AI choosing its own question, continuing across rounds, and knowing when to stop

In one case the executing role wrote a closing line worth quoting: "if you had not done these independent checks and had simply trusted my reports, at least four or five things in this work would have moved forward carrying borderline problems." That sentence is simultaneously a proof of the platform's capability and a statement of its limits.

09

How to run a study like these yourself

Every case above followed the same four steps, and the workflow is reproducible once you have access.

  1. Step 1 — get access. Email moogtang@gmail.com to arrange a trial. There is nothing to install and no environment to configure; both engines are already loaded.
  2. Step 2 — decompose one or two reference papers. Upload a paper from your field and run /bio-design scan. About 30 seconds later you see how a published study's analysis logic breaks down into evidence levels and design decisions — before you commit to your own design.
  3. Step 3 — design the study honestly. Run /bio-design and state your real constraints: no wet-lab capability, a six-month budget, a target journal. You get design.md plus the machine-readable TOPIC.yml contract, including an explicit list of the conclusions your design cannot reach.
  4. Step 4 — execute, then review the figures yourself. Hand TOPIC.yml to bio-analyze and say start. It builds the environment, writes the code and produces publication figures, then stops at the review gate. Open every figure yourself — in all six cases above, every valuable defect was caught exactly that way.

Typical wall-clock time: 1–2 hours for the design conversation, then anywhere from a few hours to 26 hours of unattended execution depending on how many evidence lines the design carries.

10

Frequently asked questions (FAQ)

Is this real data or a demonstration?

All six studies use public real data: NCBI GEO expression datasets, public GWAS summary statistics and PubMed literature. Each case directory retains the full interaction log, the decision log findings.yml, the analysis scripts and the downloaded raw data files, so every claim can be checked line by line.

Why publish negative results on a product website?

Because converging honestly is harder and more valuable than manufacturing a positive. That is exactly what pre-registration is for: the decision rules are fixed before the results are seen, so the AI cannot reverse-engineer a narrative from the outcome. A system that only ever produces positive results is not impressively capable — it is writing fiction for you.

What did the humans actually do?

A human said one sentence at the start. Everything after that was two AI roles collaborating: a clinician role that asks questions, reviews outputs, makes calls and enforces honesty, and a bioinformatics researcher role that designs, codes, executes, plots and drafts. Actions requiring real monetary spend, such as cloud GPU time, are confirmed separately and are never covered by a blanket autonomy grant.

How long does one study take?

It depends on complexity. The largest comparative study, with seven parallel evidence lines across six data layers, completed in roughly 26 hours without interruption; the three-round autonomous series took about 10.5 hours; a single Mendelian randomization study went from a one-sentence instruction to an 8,060-word English manuscript. When an API failure interrupted a run, the state machine allowed an exact resume.

Can I run a similar study myself, and how do I buy the platform?

Yes — follow the four steps above. Email moogtang@gmail.com to arrange a trial; there is no installation and no environment setup. Use bio-design to clarify the study design first, then bio-analyze to execute it. The personal and enterprise editions are quoted on enquiry with no public price list — email moogtang@gmail.com; individual users in mainland China can also purchase directly through Xiaohongshu. See the purchase section.