Give CellForge a single-cell dataset and a research question. It reads the literature, argues with itself until a group of expert agents converge on an architecture, writes the training code, and runs it.
Quickstart Β· What it designed Β· Benchmarks Β· Architecture Β· Docs Β· Paper
Building a virtual cell model is a months-long loop: read the perturbation-modeling literature, pick an architecture, adapt it to your assay, write the training code, debug the CUDA errors, tune, evaluate, repeat.
CellForge collapses that loop into a single command.
cellforge --dataset-path data/datasets/adamson.h5ad \
--task "Predict single-cell gene expression after CRISPRi knockdown in K562."What comes back is not a chat transcript. It is a research plan with citations, a runnable result.py, and a trained model β and on the six perturbation benchmarks in the paper, the models it designed are competitive with or better than hand-built state of the art.
| PCC 0.9883 | Adamson CRISPRi β best of all methods compared |
| MSE_DE 0.1736 | Norman combinatorial CRISPRa, on differentially expressed genes |
| 14 / 15 | taskβcriterion combinations where blinded LLM judges ranked CellForge first |
| 6 datasets, 3 modalities | scRNA-seq, scATAC-seq, CITE-seq β genetic, chemical, and cytokine perturbations |
| 4β8 GPU-hours | end to end, question β trained model, on a single GPU |
Note
CellForge does not ship pretrained weights. It ships the process that produces them. Every architecture below was specified by the agents, not selected from a menu.
1. Install
git clone https://github.com/gersteinlab/CellForge.git
cd CellForge
conda create -n cellforge python=3.11 -y && conda activate cellforge
pip install -e .2. Add one API key
cp .env.example .env
# Open .env and set at least one of:
# OPENAI_API_KEY / ANTHROPIC_API_KEY / DEEPSEEK_API_KEY / QWEN_API_KEY / LLAMA_API_KEY3. Check your workspace
cellforge --init # create data/{datasets,analyses,plans,codes,literature}
cellforge --doctor # verify config, dataset dir, literature dir, LLM key, Python version4. Get a dataset
python scripts/download_datasets.py adamson --out data/datasets/5. Run it
cellforge \
--dataset-path data/datasets/adamson.h5ad \
--task "Predict post-perturbation gene expression in K562 cells after CRISPRi knockdown. \
Report MSE, PCC, R2 overall and restricted to differentially expressed genes."You will find:
data/
βββ analyses/<dataset>/ # structured task analysis + retrieved literature with provenance
βββ plans/<dataset>/ # the research plan the expert agents converged on
βββ codes/<dataset>/ # result.py β verified, runnable training code
Want to also train the model it wrote? Add the opt-in execution stage:
cellforge --phase autorun --dataset-path data/datasets/adamson.h5ad --executor local --workers 2π Full walkthrough: docs/QUICKSTART.md Β· Running on a cluster: docs/QUICKSTART.md#running-on-slurm
CellForge is three stages plus an opt-in fourth. Each maps to a directory in the package, and each writes a human-readable artifact you can inspect, edit, or approve before the next stage runs.
ββββββββββββββββββββββββββββββββββββββββββββββββ
dataset (.h5ad) β β
+ β β TASK ANALYSIS cellforge/Task_Analysis/
task description β βββββββββββββ β
ββββββββββββΆβ Dataset Analyst Β· Problem Investigator β
β Baseline Assessor Β· Refinement Agent β
β + literature retrieval (local Β· PubMed Β· β
β Crossref Β· Semantic Scholar Β· Qdrant) β
β β β
β analysis report β
β β β
β β‘ METHOD DESIGN cellforge/Method_Design/
β βββββββββββββ β
β Data-modeling expert β
β Single-cell biology expert β³ debate β
β Deep-learning expert β³ review β
β Training expert β³ revise β
β β central CRITIC (area chair) β
β β β
β research plan βββ π€ checkpoint β
β β β
β β’ CODE GENERATION cellforge/Code_Generation/
β βββββββββββββ β
β coding agent edits result.py in place β
β β deterministic verifier (file / syntax / β
β CLI / contract) β bounded repair Γ5 β
β β β
β result.py βββ π€ checkpoint β
β β β
β β£ AUTORUN (opt-in) cellforge/autorun/ β
β task-wise split β local or Slurm sbatch β
ββββββββββββββββββββββββββββββββββββββββββββββββ
β
trained model + metrics
The part that makes it work: agents that disagree.
Method Design is not a prompt chain. Four domain experts each draft a proposal, then peer-review every other proposal, while a central Critic plays conference area chair. Each agent carries a coordination score updated every round:
Debate runs until every agent clears
Two human checkpoints, by design. Nothing touches a GPU unattended: you approve the research plan after Method Design, and the training script before submission.
π Deep dive: docs/ARCHITECTURE.md
A complex run costs roughly 80k prompt / 400k completion tokens, plus 4β8 GPU-hours if you train what it writes. You should be able to look at the deliverables first.
examples/outputs/ is a complete worked bundle for the Adamson CRISPRi task β the task analysis, the research plan, and the training script, laid out exactly as a run leaves them:
| Stage | Artifact | |
|---|---|---|
| Task Analysis | task_analysis_report.md |
dataset, problem, baselines, agent refinement round |
| Method Design | research_plan.md |
architecture + protocol β π€ your approval gate |
| Code Generation | result.py |
runnable training script β π€ approval before any GPU job |
| Verification | verification.json |
the deterministic, non-LLM acceptance check |
cd examples/outputs/adamson_crispri/workspace
python result.py --help # no third-party dependencies needed
python result.py --selftestImportant
The two documents are reference artifacts, not transcripts of a live run. They are written to match the schemas the code serialises, section for section. result.py is real, working, verified code, and its metrics.json is the genuine output of really running it β on synthetic data, so those numbers are a smoke test and nothing more.
Every file's provenance is stated individually in PROVENANCE.md. Real benchmark numbers are in docs/RESULTS.md.
Have you run the pipeline for real? scripts/export_example_run.py packages and scrubs a run into a bundle β a real one should replace this, and we would take that PR gladly.
Six datasets in, six architectures out. None were templates β each was specified by the agents, then implemented and trained end to end. Full model cards: docs/MODELS.md.
| Model | Designed for | The idea the agents landed on |
|---|---|---|
| CPA-X | Adamson CRISPRi (scRNA-seq) | Compositional perturbation autoencoder with a disentangled perturbation embedding and adversarial covariate removal |
| scGen-X | Norman combinatorial CRISPRa | Latent-space arithmetic extended to pairs of simultaneously activated genes, with an explicit interaction term for non-additive effects |
| ChemCellFlow | Srivatsan drug response | Sinkhorn conditional optimal transport coupled to a 6-layer normalizing flow β dose-aware and chemically conditioned |
| CPA-Traj | Schiebinger cytokine time course | Trajectory-aware VAE that conditions on time as a continuous covariate rather than a class label |
| totalGAT | Papalexi CITE-seq | Graph attention over the gene network, cross-attention between RNA and protein, separate decoder heads per modality |
| ChromDDPM | Liscovitch-Brauer scATAC-seq | Denoising diffusion over the 200k-peak accessibility profile, conditioned on the perturbation |
The trajectory-aware encoder in CPA-Traj and the diffusion denoiser in ChromDDPM have no counterpart in the seed literature corpus. They came out of the debate.
Six datasets, seven readouts, five-fold cross-validation, three independent runs, perturbation-centric averaging. Baselines: CPA, scGen, CondOT, Biolord, scGPT, GEARS, STATE, ChemCPA, CellFlow, random forest, linear regression, and the unperturbed control.
| Benchmark | CellForge model | Headline result |
|---|---|---|
| Adamson CRISPRi | CPA-X | PCC 0.9883 β best of all methods compared |
| Norman combinatorial CRISPRa | scGen-X | MSE_DE 0.1736 |
| Schiebinger cytokine time course | CPA-Traj | DEG recall 0.535 |
| Papalexi CITE-seq (protein) | totalGAT | protein recall 0.420 |
Against other autonomous systems. Blinded LLM judges (Claude 3.7, DeepSeek-R1, OpenAI o1, Qwen-plus, Llama 3.1) scored CellForge against OpenAI Deep Research, Perplexity Deep Research, Gemini Deep Research, Biomni, and a single-LLM baseline across 15 taskβcriterion combinations. CellForge ranked first in 14 of 15. Inter-judge agreement was Pearson 0.88β0.93; human expert ratings tracked the judge panel at r = 0.87 β versus r = 0.53 for the system's own internal confidence, which is precisely why the Critic is external.
| Dataset | Perturbation | Modality | Cells / features | Accession |
|---|---|---|---|---|
| Adamson 2016 | CRISPRi | scRNA-seq | 111k / 33k genes | GSE90546 |
| Norman 2019 | combinatorial CRISPRa | scRNA-seq | 84k / 17k genes | GSE133344 |
| Srivatsan 2020 | small molecules | scRNA-seq | 81k / 18k genes | GSE139944 |
| Schiebinger 2019 | cytokine time course | scRNA-seq | 65k / 17k genes | scPerturb |
| Papalexi 2021 | CRISPR | CITE-seq | 171k / 18k genes + 200 proteins | scPerturb |
| Liscovitch-Brauer 2021 | CRISPR | scATAC-seq | 58k / 200k peaks | scPerturb |
Preprocessed .h5ad files for all six are mirrored by scPerturb at DOI 10.5281/zenodo.13350497.
python scripts/download_datasets.py --list # show everything available
python scripts/download_datasets.py norman papalexi # fetch specific onesπ Preprocessing, splits, and DEG definitions: docs/DATASETS.md
| You want to swap⦠| How |
|---|---|
| LLM provider | Set OPENAI_API_KEY, ANTHROPIC_API_KEY, DEEPSEEK_API_KEY, QWEN_API_KEY, or LLAMA_API_KEY β or point CUSTOM_API_URL at any OpenAI-compatible endpoint (vLLM, Ollama, TGI, OpenRouter) |
| Coding agent | --codegen-backend codex; the backend registry in cellforge/Code_Generation/registry.py takes a new backend in about 40 lines |
| Literature corpus | Drop PDFs in CELLFORGE_LITERATURE_DIR; set QDRANT_ENABLED=true for vector search over them |
| Retrieval providers | PubMed, Crossref, and Semantic Scholar are pluggable in cellforge/retrieval/providers.py |
| Compute | --executor local, or --executor slurm with --partition / --gres / --mem / --slurm-time |
| Cost ceiling | METHOD_DESIGN_MAX_ROUNDS, METHOD_DESIGN_MAX_EXPERTS, METHOD_DESIGN_MAX_TOKENS_PER_CALL |
Per end-to-end run, from the paper:
| Simple task | Complex task | |
|---|---|---|
| Prompt tokens | ~40k | ~80k |
| Completion tokens | ~200k | ~400k |
| Generated model size | 10β30M parameters | 10β30M parameters |
| Wall clock (1 GPU) | ~4h | ~8h |
For a cheap smoke test: MODEL_NAME=gpt-4o-mini with METHOD_DESIGN_MAX_ROUNDS=2 and METHOD_DESIGN_MAX_EXPERTS=2.
When it fails, this is how β measured across runs: computation execution error 41% Β· invalid type or unsupported operation 23% Β· error-recovery failure 16% Β· model misconfiguration 6% Β· data access 5% Β· other system-level 5% Β· hallucinated structures 4%. Outright invented architecture is the rarest failure; the verifier and repair loop catch most of it. See docs/FAQ.md.
- Nothing runs on a GPU without you. The research plan and the training script are both human-approval checkpoints.
- Generated code is verified before it is published. Files, Python syntax, CLI surface, and the acceptance contract are checked deterministically; failures loop back to the same agent workspace for up to five bounded repair attempts.
- The coding agent runs in an OS sandbox (
workspace-write), with isolated credentials:CODEX_AUTH_MODE=localstrips provider API keys from the subprocess. Task Analysis and Method Design credentials are never silently reused for code generation. - Agent traces may contain model-generated commands and output. They are written with owner-only permissions under
<output_dir>/.cellforge_workspaces/. Review a workspace before publishing it.
| Quickstart | Install, configure, first run, Slurm |
| Example outputs | What the pipeline actually hands you, stage by stage |
| Architecture | The stages, agent roster, coordination score, ablations |
| Results | Full benchmark tables, all metrics, all baselines, judge study |
| Model cards | The six generated architectures in detail |
| Datasets | Sources, preprocessing, splits, DEG definitions |
| FAQ | Cost, failure modes, limitations, troubleshooting |
| Roadmap | Where this is going, and what to help with |
| Contributing | Dev setup, tests, PR conventions |
Good first issues, in rough order of how much they help:
- Add a dataset. A new perturbation dataset with a loader and a split definition is the single most valuable contribution β it directly widens the benchmark.
- Add a baseline. More comparators make the evaluation harder to argue with.
- Add a code-generation backend. The registry is small on purpose.
- Add a retrieval provider. bioRxiv, OpenAlex, and Europe PMC are all unclaimed.
- Report a failure. A run that produced a bad plan, with the plan attached, is genuinely useful data.
pip install -e ".[dev]"
python -m pytest tests -vRead CONTRIBUTING.md first. Be decent: CODE_OF_CONDUCT.md.
@article{tang2025cellforge,
title = {CellForge: Agentic Design of Virtual Cell Models},
author = {Tang, Xiangru and Yu, Zhuoyun and Chen, Jiapeng and Cui, Yan and
Shao, Daniel and Wang, Weixu and Wu, Fang and Zhuang, Yuchen and
Shi, Wenqi and Huang, Zhi and Cohan, Arman and Lin, Xihong and
Theis, Fabian and Krishnaswamy, Smita and Gerstein, Mark},
journal = {arXiv preprint arXiv:2508.02276},
year = {2025},
url = {https://arxiv.org/abs/2508.02276}
}A CITATION.cff is included, so GitHub's Cite this repository button works too.
MIT β see LICENSE.
Built at the Gerstein Lab, Yale University.
If CellForge saved you a month of architecture search, a β is a nice way to say so.