AI systems now write research papers autonomously — yet phantom references, method-code misalignment, and unreproducible scores have become endemic failures undermining scientific integrity.
Title: Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence
Google Cloud's Science One Framework treats verifiability as a first-class architectural constraint through the Chain-of-Evidence principle, solving the trustworthiness crisis in autonomous AI research at its root.
🔍 Highlight 1 — Chain-of-Evidence (CoE): two foundational principles
Completeness: every claim carries a recorded evidence chain. Correctness: each chain genuinely supports its claim. These two principles eliminate phantom references entirely — baseline systems showed rates up to 21% — while achieving best-in-class method-code alignment across all evaluated systems. The key difference from prior work: evidence chains are constructed at claim-generation time, not retrofitted as a post-hoc check.
🏗 Highlight 2 — Three-module architecture
Problem Investigator builds citation graphs from up to 100 full-text PDFs via Semantic Scholar API, grounding every reference in retrieved data rather than model memory. Discovery Engine explores parallel solution branches while keeping immutable records of all raw evaluator outputs. Paper Writer and Claim Verifier binds every factual claim to specific workspace artifacts and conservatively reconciles misalignments rather than deleting them — preserving scientific transparency.
🏆 Highlight 3 — MLE-Bench and Parameter-Golf results
Across five Kaggle competitions covering medical imaging, fine-grained recognition, and 3D perception: two Gold Medals and two Silver Medals. Won the 3D Object Detection task where every baseline system failed completely. On Parameter-Golf — a live LLM training competition under strict hardware and file-size constraints — achieved state-of-the-art as of April 27, 2026, while baselines could not produce valid submissions at all.
Rigor and capability don't trade off. Science One outperformed five state-of-the-art systems including AI Scientist v2, AutoResearchClaw, and DeepScientist, setting a new standard for verifiable autonomous research.
#
AIResearch# #
AutonomousScience#