Developer 3s Key Decision Metrics
Real-world working agents capable of reading diverse files, coordinating terminal tools, and producing verifiable deliverables require tasks grounded in realistic multi-file workspaces. Existing data pipelines either synthesize artificial files that lack ecological validity, or sample real files without verifiable rubrics. Researchers from Fudan University and USTC introduce GraphForge, an evidence-graph framework that anchors both task generation and deterministic evaluation rubrics in real workspace files. Fine-tuning Qwen3.6-27B on only 2,169 GraphForge trajectories catapults GDPVal to 1445.7 (+65.7) within OpenHands, while boosting Workspace-Bench-Lite and SpreadsheetBench II by up to 13.7 points under Claude Code.
Key Takeaways
- ✓Pioneers GraphForge to synthesize real-world workspaces grounded in topological evidence graphs with deterministic rubrics
- ✓Drives OpenHands GDPVal to 1445.7 (+65.7 points) by fine-tuning Qwen3.6-27B on only 2,169 curated trajectories
- ✓Elevates Workspace-Bench-Lite to 63.7 and SpreadsheetBench II to 24.0 under Claude Code runtime environments
Heavy Claude Code use: compare subscription limits and API bills
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
核心背景与行业痛点
Modern working agents—such as Devin, OpenHands, and Claude Code—are tasked with reading heterogeneous files, orchestrating tools, and delivering functional outputs across enterprise workspaces. Training these agents demands tasks situated in rich multi-file workspaces paired with deterministic verification rubrics. Existing data generation recipes face an acute tradeoff: fully synthetic files generated by LLMs lack real-world messiness and diversity, while crawling organic repositories leaves tasks without automated, task-specific ground-truth verifiers, precluding reliable post-training alignment.
架构亮点与底层机制
Researchers from Fudan University and USTC introduce GraphForge (arXiv:2609.38923), grounding workspace synthesis in an evidence graph:
- Occupation-Grounded Workspace Seeding: Assembles workspaces of authentic real files spanning code, structured spreadsheets, and configuration scripts seeded by occupational domain requirements.
- Topological Evidence Graph Construction: Maps dependency references, variable schemas, and cross-file relationships into an explicit semantic evidence graph.
- Dual-Anchored Formulation: Simultaneously generates task statements and deterministic evaluation rubrics directly from the evidence graph, anchoring every evaluation criterion to verifiable workspace ground-truths.
- Initial Rollout & Revision Pipeline: Evaluates task feasibility with an execution probe and employs an automated revision agent to refine tasks and repair rubrics prior to trajectory collection.
权威 Benchmark 与实测跑分对比
Benchmarked across demanding interactive agent environments:
- 65.7-Point Leap on OpenHands GDPVal: Fine-tuning Qwen3.6-27B on only 2,169 GraphForge trajectories pushes GDPVal to 1445.7 (+65.7) within the OpenHands harness.
- Up to 13.7-Point Boost Under Claude Code: Catapults Workspace-Bench-Lite to 63.7 (+7.7) and SpreadsheetBench II to 24.0 (+13.7) in the Claude Code runtime.
- Rejection Fine-Tuning (RFT) Gains: Rejection fine-tuning using evidence-anchored rubrics yields consistent secondary improvements, demonstrating the fidelity of GraphForge's automated reward signal.
开发者实战落地与开箱指南
GraphForge datasets, synthesis engines, and checkpoints are publicly accessible. For AI teams building enterprise automation agents and full-stack software engineers, GraphForge establishes a reproducible framework to synthesize high-fidelity multi-file environments, accelerating agentic capabilities without human verification bottlenecks.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.