engineer2: ScientistTwo-style goal-driven research
Input: a scientific problem G (a SOTA paper and its code). Output: improved paper P+ and reproducible codebase C+.
Keep state across modules: h_best, E_best (results), C_best (code), E_base/C_base (baseline), the trace log R of every idea tried, and E_abl.
Use subagents for the roles named below where useful.
Pipeline
- Setup
- Versioning (applies throughout)
- Limitations
- Seed Ideas
- Idea Evaluation
- Idea Refinement
- Ablation
- Drafting & Review-Rebuttal
- Meta-Review
Module: Setup
Before starting, ask the user these questions in a single message. Show the suggested value with each one, and use it if the user accepts:
- Goal: what are you trying to achieve with this work (e.g. target metric/threshold, deployment constraint, speed/memory budget, intended use)? (no suggestion; the user must supply this). Save the answer verbatim to
goal.mdat project root, and re-read it whenever a Goal rating is to be made. Maintain this file in concise instruction bullets whenever user provides more details of the goal. - Input G: which SOTA paper file and baseline repo path? (no suggestion; the user must supply these)
- N_seed (seed pool size)? Suggest 8.
- N0 (seed ideas evaluated in round 0)? Suggest 2.
- Subset: which part of the benchmark? Suggest the 1-2 smallest datasets on the primary metric, using under ~10% of full-run compute.
- Good/Bad rule: Good = beats
E_baseon most subset datasets on the primary metric; Bad = worse on most of them; otherwise Engineer. OK? - "Better" for the verification gates: prefer
E_newon the primary metric averaged across datasets, with no dataset regressing beyond noise. OK? - N_p (ablation plans)? Suggest one per component (3-6). N_t (rebuttal tasks)? Suggest ≤3.
- Output location? Suggest
research_out/(R.jsonl, E_*.json, code/, paper/paper.md + paper/figures/; write FAILURE.md if the run terminates).
Finally, store all the above project setup except already in goal.md into project_setup.md at project root and re-read it if you miss the context. Be concise and only bulleted instructions. Update this file when more policies are decided by either user or yourself, like data file paths.
Module: Versioning
Applies throughout. Give each idea an <idea_id-codename>, e.g. 07-kernel-earlystop (a sequential id plus a short slug).
Before starting a new version, copy (snapshot) the current version into <output folder>/<idea_id-codename>/v<n>/. A snapshot holds the idea text, results, code, the whole paper/ directory (including figures/) and the review. Number v<n> per idea, and start a new one for each refinement or paper draft. Engineer reruns stay in the same version. Paper drafts go under h_best's id.
Never overwrite a snapshot. Snapshot rejected versions too, and log their paths in R. If a refinement fails a gate, restore the working tree from the best snapshot.
When conceiving a new idea, scan these folders to avoid repeating yourself.
Module: Limitations
- As the agent Limitation Extractor, list the limitations of the SOTA method on G.
- As the agent Limitation Verifier, judge whether the list is enough to guide novel improvements.
- If it is not enough, extract the missing limitations and verify again. Stop when the verifier confirms all actionable limitations are covered, or after 16 extract→verify cycles (the first one counts).
Module: Seed Ideas
Save untested ideas into seeds.md.
- Write an initial idea h0 that targets the limitations.
- As the agent Promising Checker, score its potential success rate against the goal. Retrieve 2 related papers via web search and compare against them.
- As the agent Idea Generator, add distinct ideas to deal with the limitations until you have N_seed candidates.
- Sort the seed pool H0 by the potential success rate against the goal, highest first.
Module: Idea Evaluation
Treat this as one callable unit, Coder(G, h) -> (h, E, C, decision, feedback), run by another agent.
- Baseline (run once, before Idea Refinement; every
Codercall reuses it): reproduce G's main experiments on a representative benchmark subset to getE_baseandC_base. - Subset run: implement h by modifying
C_base, then run it on the subset. - Subset Critic: compare against
E_baseand decide, applying Setup's Good/Bad rule:Bad: worse than the baseline. Discard the idea.Good: better than the baseline. Approve it for scale-up.Engineer: neither; it needs tuning or code fixes. As the agent Subset Engineer, apply the critic's feedback, rerun, and send it back to the critic.
- Allow at most N_eng=2 engineer+rerun cycles after the first critic call. If the idea is still not
Goodafter that, mark itBadand prune it. - Full-set: for
Goodideas, adapt the code to the full benchmark (all datasets and metrics). RunC_baseon the full benchmark once (or use G's reported numbers) as the reference. A Full-Set Critic decidesGood/Bad/Engineerby the same rule, and the Full-Set Engineer fixes the idea for at most N_eng cycles. Its final decision isGoodorBad.
Module: Idea Refinement
- Round 0: run
Coderon the top-N0 seed ideas and log every trace inR. - Round k≥1: as the agent Idea Evolver, read all earlier traces (Good results and Bad failure logs) and propose N_k evolved ideas.
- For exploration, add the next N_e unevaluated seed ideas from H0, in novelty order. Default: N_k=1 evolved and N_e=1 seed idea per round.
- Run
Coderon each candidate and append the results toR. - Stop once S=4 ideas have a final (full-set) decision of
Good, or after K=4 rounds in total (rounds 0..3). If no idea isGoodby round K, terminate the whole pipeline and report failure. - As the agent Selector, compare full-benchmark metrics and logs across all
Goodideas. Seth_best,E_best,C_bestto the best one.
Module: Ablation
- As the agent Ablation Planner, write N_p executable component-level ablation plans for
h_best. - Run each plan by modifying
C_best. Collect the results asE_abl. - As the agent Ablation Critic, inspect the component breakdown and decide
Good(optimal) orRefine, with a critique. - On
Refine, act as the Full-Set Engineer agent: use the critique (e.g. drop redundant or harmful components) to produceh_new,E_new,C_new. - Verification gate: as the agent Result Comparison role, replace the best state with the new one only if
E_newis strictly better thanE_best. If you replace it, redo the ablation planning and runs on the newh_best, for reporting only; this cannot trigger a further refinement once N_abl is spent. If it fails, restore the best snapshot. - Repeat until the critic says
Good, or for at most N_abl=1 refinement.
Module: Drafting & Review-Rebuttal
- As the agent Initial Drafter, combine
h_best,E_best, andE_ablinto a full manuscript,paper/paper.md, in Markdown. Save figures as PNG files inpaper/figures/and link them with relative paths, e.g.. - As the agent Goal Reviewer, read
goal.mdand the paper. Write strengths, weaknesses, and questions with respect to the goal, and give a Goal rating from 1 to 10: how well the method, results, and paper achieve the user's goal (10 = fully achieved and convincingly demonstrated; 8 = achieved with minor gaps; 5 = partially achieved; 1 = not addressed). Justify the rating against each part of the goal. - If the Goal rating is below 8: as the Rebuttal Planner, turn the review's goal gaps into N_t supplementary experiment tasks. As the agent Rebuttal Coder, run each task on
C_bestto getE_reb. - As the agent Paper Enhancer, fold the review and
E_rebinto the paper by revising claims, tables, and figures. Then review it again. - Repeat until the Goal rating is at least 8, or after at most N_peer=2 Reviewer calls in total (the first one counts).
Module: Meta-Review
- As the agent Meta-Reviewer, read
goal.md, the paper, and the latest review (with its Goal rating). DecideAcceptorRefine, with a meta-critique. - On
Accept: export, i.e. leave the finalpaper/paper.md+paper/figures/andcode/(=C_best) in the output folder. Done. - On
Refine(a critical algorithmic or empirical weakness, or a major part of the goal not met): act as the Full-Set Engineer and use the meta-critique to produceh_new,E_new,C_new. - Verification gate: if
E_newis strictly better thanE_best, update the best state, then rerun Ablation (reporting only, no new refinement) and Drafting & Review-Rebuttal. Otherwise discard the change, restore the best snapshot, and export; the run ends. On any exit other than Accept, export P+ = the latest paper built fromC_best, and C+ =C_best. - Repeat until
Accept, or for at most N_meta=1 refinement. After that refinement, export without running another meta-review.