-
Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis
Authors:
Zijiao Chen,
Nicholas Lu,
Xinhui Li,
Jocelyn A. Ricard,
Ce Ju,
Huan H. Wang,
Christian Kindermann,
Jeanette A. Mumford,
Steven Dillmann,
James Kent,
Alejandro de la Vega,
Sanmi Koyejo,
Vince D. Calhoun,
Joshua W. Buckholtz,
Juan Helen Zhou,
Steffen Bollmann,
Russell A. Poldrack
Abstract:
AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria. We present Brain Researcher, an agentic research harness operating in a neuroimag…
▽ More
AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria. We present Brain Researcher, an agentic research harness operating in a neuroimaging researcher's computational environment under rules for admissible analyses, required checks and claim scope. In benchmarks, Brain Researcher increased first-choice tool-selection accuracy across seven models by 70.2 percentage points (23.3% without it versus 93.6% with it) and verifiable grounding from 4.6% to 22.0%. In collaborator-led and self-evolving studies, multiverse analyses exposed analytic-choice sensitivity, and scientific review classified claims as accepted, qualified, revised, blocked, rejected or deferred. By linking decisions to evidence and provenance, Brain Researcher embeds methodological judgment within the workflow, not after it.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Early Exploration of the Scientific Discovery Space for the Habitable Worlds Observatory
Authors:
Courtney D. Dressing,
Danica Adams,
Evelyne Alecian,
Gagandeep Anand,
Giada Arney,
Sarah Gomes Aroucha Barbosa,
Martin Barstow,
Joanna K. Barstow,
Rachael L. Beaton,
Eduardo Bendek,
Svetlana Berdyugina,
Julie Biedermann,
Sarah Blunt,
Sanchayeeta Borthakur,
Kara Brugman,
Joseph N. Burchett,
Eric Burns,
Jenna M. Cann,
Ludmila Carone,
Cody A. Carr,
Richard Cartwright,
Renyue Cen,
Jean-yves Chaufray,
Pin Chen,
Lígia F Coelho
, et al. (302 additional authors not shown)
Abstract:
The Habitable Worlds Observatory (HWO) is a future NASA flagship mission concept identified by the Astro2020 Decadal Survey as the highest priority for large space missions. HWO should conduct "transformative astrophysics" and search for biosignatures in the atmospheres of approximately 25 potentially Earth-like planets. To further the early-stage development of HWO, NASA formed the Science, Techn…
▽ More
The Habitable Worlds Observatory (HWO) is a future NASA flagship mission concept identified by the Astro2020 Decadal Survey as the highest priority for large space missions. HWO should conduct "transformative astrophysics" and search for biosignatures in the atmospheres of approximately 25 potentially Earth-like planets. To further the early-stage development of HWO, NASA formed the Science, Technology, Architecture Review Team (START). In turn, START invited the scientific community to join working groups to explore the potential discovery space. In this paper, we present 70 science cases that resulted from this process. The cases address four scientific pillars: growth of galaxies (15 cases), evolution of the elements (13 cases), solar systems in context (32 cases), and living worlds (10 cases). Combined, they would address 27 of the 30 science questions and discovery areas identified by Astro2020. The 140 observing programs needed for the 70 investigations encompass a rich variety of spectroscopic (for 87% of science cases) and photometric (for 30%) observations extending from the UV to the NIR. Additionally, high-contrast and polarimetric capabilities would be needed for 34% and 27% of science cases, respectively. Access to UV wavelengths is critical: 83% of science cases need data at wavelengths <400 nm, and 26% extend to <100 nm. In the NIR, 26% of science cases need observations at wavelengths >=2000 nm. Pursuing the full portfolio of science would also necessitate precise astrometry for planet mass measurement, rapid response capabilities, a large instantaneous field of regard, non-sidereal tracking, saturation mitigation strategies, and high dynamic range.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
OpenThoughts-Agent: Data Recipes for Agentic Models
Authors:
Negin Raoof,
Richard Zhuang,
Marianna Nezhurina,
Etash Guha,
Atula Tejaswi,
Ryan Marten,
Charlie F. Ruan,
Tyler Griggs,
Alexander Glenn Shaw,
Hritik Bansal,
E. Kelly Buchanan,
Artem Gazizov,
Reinhard Heckel,
Chinmay Hegde,
Sankalp Jajee,
Daanish Khazi,
Emmanouil Koukoumidis,
Xiangyi Li,
Hange Liu,
Shlok Natarajan,
Harsh Raj,
Nicholas Roberts,
Ethan Shen,
Nishad Singhi,
Michael Siu
, et al. (25 additional authors not shown)
Abstract:
Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically target a single benchmark, leaving open the question of how to train models that generalize across diverse agentic tasks. The OpenThoughts-Agent (OT-Agent) project…
▽ More
Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically target a single benchmark, leaving open the question of how to train models that generalize across diverse agentic tasks. The OpenThoughts-Agent (OT-Agent) project addresses this gap with a fully open data curation pipeline for training agentic models. We conduct more than 100 controlled ablation experiments to systematically investigate each stage of the pipeline, yielding insights on the importance of task sources and diversity. We then assemble a training set of 100K examples from our pipeline and fine-tune Qwen3-32B on this dataset, which yields an average accuracy of 44.8% across seven agentic benchmarks and a 3.9 percentage point improvement over the strongest existing open data agentic model (Nemotron-Terminal-32B, 40.9%). Moreover, our training data exhibits strong scaling properties, outperforming alternative open datasets at every training set size in compute-controlled comparisons. We publicly release our training sets, data pipeline, experimental data, and models at openthoughts.ai to support future open research on agentic model training.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Authors:
Jan Batzner,
Sree Harsha Nelaturu,
Damian Stachura,
Anastassia Kornilova,
Jon Crall,
Tommaso Cerruti,
Yanan Long,
Yifan Mai,
Sanchit Ahuja,
Asaf Yehudai,
Marek Šuppa,
John P. Lalor,
Oluwagbemike Olowe,
Jatin Ganhotra,
Brian H. Hu,
Eliya Habba,
Andrew M. Bean,
Chang Liu,
Sander Land,
Steven Dillmann,
Aniketh Garikaparthi,
Elron Bandel,
Saki Imai,
James Edgell,
Wm. Matthew Kennedy
, et al. (23 additional authors not shown)
Abstract:
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First, results are saved in incompatible formats, scattered across leaderboards, papers, blog posts, evaluation harness logs, and custom repositories. Second, results are created by different evaluation frameworks, which prod…
▽ More
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First, results are saved in incompatible formats, scattered across leaderboards, papers, blog posts, evaluation harness logs, and custom repositories. Second, results are created by different evaluation frameworks, which produce divergent scores for nominally identical evaluations and record metadata inconsistently, hindering comparison, cross-community evaluation science, cost reduction, and reuse. We introduce Every Eval Ever, the first shared schema and community-crowdsourced repository for AI evaluation results. The schema standardizes how evaluations are represented in a unified, single JSON document. It is source-agnostic by design, ingesting results from evaluation harnesses and papers alike, and optionally stores per-instance outputs for fine-grained analysis. We contribute: (i) a community-governed metadata schema with a companion instance-level schema, the first standardization effort of its kind; (ii) automatic converters from popular formats, evaluation harnesses, and leaderboards to the unified schema; and (iii) a crowdsourced community database hosted on Hugging Face, currently spanning to date 22,235 models, 2,273 unique benchmarks, and 31 evaluation formats.
△ Less
Submitted 12 June, 2026;
originally announced June 2026.
-
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Authors:
Rishi Desai,
Jesse Hu,
Joan Cabezas,
Neel Harsola,
Pratyush Shukla,
Roey Ben Chaim,
Adnan El Assadi,
Omkaar Mukund Kamath,
Fenil Faldu,
Prannay Hebbar,
Jiankai Sun,
Yiyuan Li,
Pramod Srinivasan,
Ishan Gupta,
Christopher Settles,
Daniel Wang,
Derek Chen,
Pranav Raja,
Albert Liu,
Marek Šuppa,
Nevasini Sasikumar,
Luyang Kong,
Erik Quintanilla,
Xiangyi Li,
Ivan Bercovich
, et al. (1 additional authors not shown)
Abstract:
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent benchmarks largely evaluate short-form tasks, such as single pull requests, small tickets, or 5-10 minute exercises, limiting our ability to measure agents' capabilities in planning, long-context understanding, and memory…
▽ More
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent benchmarks largely evaluate short-form tasks, such as single pull requests, small tickets, or 5-10 minute exercises, limiting our ability to measure agents' capabilities in planning, long-context understanding, and memory use. We introduce SWE-Marathon, a benchmark of 20 long-horizon tasks spanning software engineering and adjacent technical domains. Each task consists of a unique executable environment, a human-written reference solution, and a multi-layer verification suite. Logged agent attempts average 27.2M total tokens, making SWE-Marathon substantially longer-horizon than existing SWE and command-line agent benchmarks. Current frontier coding agents solve fewer than 30% of tasks. Failures often arise from poor self-verification, self-reported infeasibility, and premature termination. We also observe reward-hacking behavior in 13.8% of rollouts, where agents attempt to exploit the environment or verifier to bypass the intended workflow. SWE-Marathon includes adversarial review of test suites and execution environments, as well as multi-layer checks designed to prevent shortcut solutions. We release SWE-Marathon, evaluation code, and agent trajectories at https://swe-marathon.org/.
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
Toward decision-aware AI for LSST-scale time-domain astronomy
Authors:
C. R. Bom,
A. Mahabal,
F. Bianco,
P. Darc,
B. Fraga,
R. Bonito,
S. Chaini,
M. W. Coughlin,
S. Dillmann,
F. Fontinele Nunes,
A. Gomboc,
N. Hernitschek,
X. Li,
F. Z. Majidi,
A. I. Malz,
A. Melandri,
V. Petrecca,
S. Piranomonte,
M. Rabus,
F. Ragosta,
O. Razim,
M. C. Romão,
N. Sarin,
A. Sasli,
V. A. Srećković
, et al. (5 additional authors not shown)
Abstract:
The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will generate approximately (10^7) alerts per night, pushing time-domain astronomy beyond pipelines that treat discovery as a static labeling problem. We argue that LSST is better understood as a partially observed dynamical environment, in which scientific return depends on the quality of follow-up decisions made under uncerta…
▽ More
The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will generate approximately (10^7) alerts per night, pushing time-domain astronomy beyond pipelines that treat discovery as a static labeling problem. We argue that LSST is better understood as a partially observed dynamical environment, in which scientific return depends on the quality of follow-up decisions made under uncertainty and finite observational resources. The central challenge is therefore to maintain evolving, uncertainty-aware representations of astrophysical sources and to select actions that maximize long-term scientific value. We propose that foundation models trained on heterogeneous time-domain data can learn survey-scale representations of source state, while decision-theoretic policies support principled, auditable allocation of follow-up resources. Embedded within human-supervised agentic systems, these components position AI as part of the operational inference loop rather than as a downstream predictive tool. The way such systems represent belief, optimize utility, and expose their reasoning will shape observational efficiency, the distribution of scientific agency, including who participates in discovery and the scientific questions that receive priority.
△ Less
Submitted 3 June, 2026;
originally announced June 2026.
-
AsymmetryZero: A Framework for Operationalizing Human Expert Preferences as Semantic Evals
Authors:
Tadhg Looram,
Lucas Nuzzi,
Kyle Waters,
Steven Dillmann
Abstract:
Much of the focus in RL today is on evaluation design: building meaningful evals that serve simultaneously as benchmarks and as well-defined reward signals for post-training. Yet, many real-world tasks are governed by subjective, procedural, and domain-specific requirements that are difficult to encode as exact-match targets or open-ended preference judgments frequently used in RL pipelines today.…
▽ More
Much of the focus in RL today is on evaluation design: building meaningful evals that serve simultaneously as benchmarks and as well-defined reward signals for post-training. Yet, many real-world tasks are governed by subjective, procedural, and domain-specific requirements that are difficult to encode as exact-match targets or open-ended preference judgments frequently used in RL pipelines today. In this work, we present AsymmetryZero, a framework for operationalizing human expert preferences as semantic evals. AsymmetryZero represents each task as a stable evaluation contract that makes grading criteria explicit: what is being graded, how each criterion is judged, and how criterion-level decisions are aggregated into a task outcome. The same contract can be executed using Inspect for model-only evaluations, as well as the Harbor Framework for agentic evaluations, enabling comparable scores and shared audit artifacts across both settings. We argue that the central challenge in post-training today is the faithful encoding of expert requirements into the evaluation itself. To that end, we present a study using Harbor that holds task contracts fixed and compares a five-model frontier jury against a five-model compact jury across four frontier-class solvers (Claude Opus 4.6, GPT-5.4, Grok-4.20, Gemini-3.1-Pro). We find that criterion-level frontier-vs-compact agreement ranges from $75.9\%$ to $89.6\%$ (strict common-subset agreement: $77.8\%$ to $92.1\%$), while compact juries exhibit substantially higher internal dissent (3--2 split rate $28.7\%$--$32.4\%$) than frontier juries ($6.1\%$--$11.5\%$). Verifier traces further show that compact juries reduce per-criterion judging cost to roughly $4.2\%$--$5.6\%$ of frontier and latency to roughly $21.7\%$--$27.1\%$, even as aggregated task-level outcomes often remain comparatively stable.
△ Less
Submitted 15 April, 2026;
originally announced May 2026.
-
COMPOSITE-Stem
Authors:
Kyle Waters,
Lucas Nuzzi,
Tadhg Looram,
Alessandro Tomasiello,
Ariel Ghislain Kemogne Kamdoum,
Bikun Li,
Damien Sileo,
Egor Kretov,
Francesco Fournier-Facio,
Georgios Soloupis,
Haile Kassahun,
Hew Wolff,
Jiaqi Cai,
Lianghui Li,
Marc Roth,
Mohinder Naiya,
Naixu Guo,
Qicheng Tang,
Richard Wheeler,
Samuele Sala,
Serguei Popov,
Steven Dillmann,
Yuqi Li
Abstract:
AI agents hold growing promise for accelerating scientific discovery; yet, a lack of frontier evaluations hinders adoption into real workflows. Expert-written benchmarks have proven effective at measuring AI reasoning, but most at this stage have become saturated and only measure performance on constrained outputs. To help address this gap, we introduce COMPOSITE-STEM, a benchmark of 70 expert-wri…
▽ More
AI agents hold growing promise for accelerating scientific discovery; yet, a lack of frontier evaluations hinders adoption into real workflows. Expert-written benchmarks have proven effective at measuring AI reasoning, but most at this stage have become saturated and only measure performance on constrained outputs. To help address this gap, we introduce COMPOSITE-STEM, a benchmark of 70 expert-written tasks in physics, biology, chemistry, and mathematics, curated by doctoral-level researchers. Our benchmark combines exact-match grading and criterion-based rubrics with an LLM-as-a-jury grading protocol, allowing more flexible assessment of scientifically meaningful outputs. Using an adapted multimodal Terminus-2 agent harness within the Harbor agentic evaluation framework, we evaluate four frontier models. The top-performing model achieves 21%, demonstrating that COMPOSITE-STEM captures capabilities beyond current agent reach. All tasks are open-sourced with contributor permission to support reproducibility and to promote additional research towards AI's acceleration of scientific progress in these domains.
△ Less
Submitted 16 April, 2026; v1 submitted 10 April, 2026;
originally announced April 2026.
-
SLSim: a strong lensing population simulation package
Authors:
Narayan Khadka,
Simon Birrer,
Henry Best,
Paras Sharma,
Katsuya T. Abe,
Xianzhe Tang,
Carly Mistick,
Felipe Urcelay,
Emrecan M. Sonmez,
Nikki Arendse,
Sydney Erickson,
Jacob O. Hjortlund,
Phil Holloway,
Alan Huang,
Rahul Karthik,
Mia Lamontagne,
Vibhore Negi,
Justin R. Pierel,
Bruno Sanchez,
Aysu Ece Saricaoglu,
Anowar Shajib,
Yixuan Shao,
Padma Venkatraman,
Bryce Wedig,
Aadya Agrawal
, et al. (23 additional authors not shown)
Abstract:
Gravitational lensing offers unique insights into cosmology by bending light around massive objects. Strong gravitational lensing, in particular, produces magnified and often multiple images of distant sources, crucial for precise cosmological measurements and understanding the distribution of dark matter in the universe. Current studies are limited by the number of strong gravitational lenses. Fr…
▽ More
Gravitational lensing offers unique insights into cosmology by bending light around massive objects. Strong gravitational lensing, in particular, produces magnified and often multiple images of distant sources, crucial for precise cosmological measurements and understanding the distribution of dark matter in the universe. Current studies are limited by the number of strong gravitational lenses. From upcoming cosmological surveys, we anticipate observing a several orders of magnitude increase in the number of lenses, for both static and transient phenomena. However, detecting and analyzing these events from vast surveys like Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) presents significant challenges. To prepare for these challenges, we introduce SLSim, a versatile simulation tool tailored for the Vera C. Rubin Observatory. SLSim integrates advanced astrophysical models with computational efficiency to generate synthetic strong lens populations under realistic observational conditions. SLSim simulates static and variable lensing scenarios, essential for cosmological studies, training and testing lens search and data analysis pipelines. This paper details SLSim,'s design and implementation, emphasizing its modularity and capabilities across various astrophysical regimes. Validation against observational data and existing simulations confirms SLSim's accuracy in reproducing observed lensing phenomena. SLSim is publicly available at https://github.com/LSST-strong-lensing/slsim, and we anticipate continued development and expansion of its capabilities. Users are encouraged to check the repository for updates and to contribute to ongoing community efforts in strong lensing simulations.
△ Less
Submitted 2 June, 2026; v1 submitted 17 March, 2026;
originally announced March 2026.
-
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Authors:
Xiangyi Li,
Yimin Liu,
Wenbo Chen,
Bingran You,
Zonglin Di,
Yifeng He,
Shenghan Zheng,
Kyoung Whan Choe,
Jiankai Sun,
Shuyi Wang,
Chujun Tao,
Binxu Li,
Xuandong Zhao,
Hejia Geng,
Xiaojun Wu,
Junwei Zhou,
Xiaokun Chen,
Hanwen Xing,
Yubo Li,
Qunhong Zeng,
Di Wang,
Yuanli Wang,
Roey Ben Chaim,
Penghao Jiang,
Haotian Shen
, et al. (53 additional authors not shown)
Abstract:
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark whose current inventory contains 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Our latest aggregate evaluation ru…
▽ More
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark whose current inventory contains 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Our latest aggregate evaluation runs the 87-task benchmark under matched no-Skills and curated-Skills conditions for 18 model-harness configurations. Curated Skills raise the average pass rate from 33.9% to 50.5% (+16.6 percentage points; 25.5% normalized gain), with configuration-level gains ranging from +4.1 to +25.7 pp. Focused Skills with at most three modules outperform larger or exhaustive bundles, and smaller models with Skills can match larger models without them. SkillsBench establishes paired evaluation as the foundation for rigorous measurement of Skill efficacy on agentic, expertise-heavy work.
△ Less
Submitted 14 June, 2026; v1 submitted 13 February, 2026;
originally announced February 2026.
-
Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration
Authors:
LSST Dark Energy Science Collaboration,
Eric Aubourg,
Camille Avestruz,
Matthew R. Becker,
Biswajit Biswas,
Rahul Biswas,
Boris Bolliet,
Adam S. Bolton,
Clecio R. Bom,
Raphaël Bonnet-Guerrini,
Alexandre Boucaud,
Jean-Eric Campagne,
Chihway Chang,
Aleksandra Ćiprijanović,
Johann Cohen-Tanugi,
Michael W. Coughlin,
John Franklin Crenshaw,
Juan C. Cuevas-Tello,
Juan de Vicente,
Seth W. Digel,
Steven Dillmann,
Mariano Javier de León Dominguez Romero,
Alex Drlica-Wagner,
Sydney Erickson,
Alexander T. Gagliano
, et al. (41 additional authors not shown)
Abstract:
The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful…
▽ More
The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.
△ Less
Submitted 20 January, 2026;
originally announced January 2026.
-
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Authors:
Mike A. Merrill,
Alexander G. Shaw,
Nicholas Carlini,
Boxuan Li,
Harsh Raj,
Ivan Bercovich,
Lin Shi,
Jeong Yeon Shin,
Thomas Walshe,
E. Kelly Buchanan,
Junhong Shen,
Guanghao Ye,
Haowei Lin,
Jason Poulos,
Maoyu Wang,
Marianna Nezhurina,
Jenia Jitsev,
Di Lu,
Orfeas Menis Mastromichalakis,
Zhiwei Xu,
Zizhao Chen,
Yue Liu,
Robert Zhang,
Leon Liangyu Chen,
Anurag Kashyap
, et al. (60 additional authors not shown)
Abstract:
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems f…
▽ More
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\% on the benchmark and conduct an error analysis to identify areas for model and agent improvement. We publish the dataset and evaluation harness to assist developers and researchers in future work at https://www.tbench.ai/ .
△ Less
Submitted 16 January, 2026;
originally announced January 2026.
-
Investigating the Dark Energy Constraint from Strongly Lensed AGN at LSST-Scale
Authors:
Sydney Erickson,
Martin Millon,
Padmavathi Venkatraman,
Tian Li,
Philip Holloway,
Phil Marshall,
Anowar Shajib,
Simon Birrer,
Xiang-Yu Huang,
Timo Anguita,
Steven Dillmann,
Narayan Khadka,
Kate Napier,
Aaron Roodman,
The LSST Dark Energy Science Collaboration
Abstract:
Strongly lensed Active Galactic Nuclei (AGN) with an observable time delay can be used to constrain the expansion history of the Universe through time-delay cosmography (TDC). As the sample of time-delay lenses grows to statistical size, with $\mathcal{O}$(1000) lensed AGN forecast to be observed by the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), there is an emerging opportun…
▽ More
Strongly lensed Active Galactic Nuclei (AGN) with an observable time delay can be used to constrain the expansion history of the Universe through time-delay cosmography (TDC). As the sample of time-delay lenses grows to statistical size, with $\mathcal{O}$(1000) lensed AGN forecast to be observed by the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), there is an emerging opportunity to use TDC as an independent probe of dark energy. To take advantage of this statistical sample, we implement a scalable hierarchical inference tool which computes the cosmological likelihood for hundreds of strong lenses simultaneously. With this new technique, we investigate the cosmological constraining power from a simulation of the full LSST sample. We start from individual lenses, and emulate the full joint hierarchical TDC analysis, including image-based modeling, time-delay measurement, velocity dispersion measurement, and external convergence prediction. We fully account for the mass-sheet and mass-anisotropy degeneracies. We assume a sample of 800 lenses, with varying levels of follow-up fidelity based on existing campaigns. With our baseline assumptions, within a flexible $w_0w_a$CDM cosmology, we simultaneously forecast a $\sim$2.5% constraint on H0 and a dark energy figure of merit (DE FOM) of 6.7. We show that by expanding the sample from 50 lenses with IFU kinematics to include 750 lenses with plausible LSST time-delay measurements, we improve the forecasted DE FOM by nearly a factor of 3, demonstrating the value of incorporating this portion of the sample. We also investigate different follow-up campaign strategies, and find significant improvements in the DE FOM with additional stellar kinematics measurements and higher-precision time-delay measurements. We also demonstrate how the redshift configuration of time-delay lenses impacts constraining power in $w_0w_a$CDM.
△ Less
Submitted 14 July, 2026; v1 submitted 17 November, 2025;
originally announced November 2025.
-
The Advanced X-ray Imaging Satellite (AXIS) Community Science Book
Authors:
Michael Koss,
Nafisa Aftab,
Steven W. Allen,
Roberta Amato,
Hongjun An,
Igor Andreoni,
Timo Anguita,
Riccardo Arcodia,
Thomas Ayres,
Matteo Bachetti,
Maria Cristina Baglio,
Arash Bahramian,
Marco Balboni,
Ranieri D. Baldi,
Solen Balman,
Aya Bamba,
Eduardo Banados,
Tong Bao,
Iacopo Bartalucci,
Antara Basu-Zych,
Rebeca Batalha,
Lorenzo Battistini,
Franz Erik Bauer,
Andy Beardmore,
Werner Becker
, et al. (373 additional authors not shown)
Abstract:
The AXIS Community Science Book represents the collective effort of 592 scientists worldwide to define the transformative science enabled by the Advanced X-ray Imaging Satellite (AXIS), a next-generation X-ray mission selected by NASA's Astrophysics Probe Program for Phase A study. AXIS will advance the legacy of high-angular-resolution X-ray astronomy with ~1.5'' imaging over a wide 24' field of…
▽ More
The AXIS Community Science Book represents the collective effort of 592 scientists worldwide to define the transformative science enabled by the Advanced X-ray Imaging Satellite (AXIS), a next-generation X-ray mission selected by NASA's Astrophysics Probe Program for Phase A study. AXIS will advance the legacy of high-angular-resolution X-ray astronomy with ~1.5'' imaging over a wide 24' field of view and an order of magnitude greater collecting area than Chandra in the 0.3-12 keV band. Combining sharp imaging, high throughput, and rapid response capabilities, AXIS will open new windows on virtually every aspect of modern astrophysics, exploring the birth and growth of supermassive black holes, the feedback processes that shape galaxies, the life cycles of stars and exoplanet environments, and the nature of compact stellar remnants, supernova remnants, and explosive transients. This book compiles 138 community-contributed science cases developed by five Science Working Groups focused on AGN and supermassive black holes, galaxy evolution and feedback, compact objects and supernova remnants, stellar physics and exoplanets, and time-domain and multi-messenger astrophysics. Together, these studies establish the scientific foundation for next-generation X-ray exploration in the 2030s and highlight strong synergies with facilities of the 2030s, such as JWST, Roman, Rubin/LSST, SKA, ALMA, ngVLA, and next-generation gravitational-wave and neutrino networks.
△ Less
Submitted 6 January, 2026; v1 submitted 31 October, 2025;
originally announced November 2025.
-
ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
Authors:
Christine Ye,
Sihan Yuan,
Suchetha Cooray,
Steven Dillmann,
Ian L. V. Roque,
Dalya Baron,
Philipp Frank,
Sergio Martin-Alvarez,
Nolan Koblischke,
Frank J Qu,
Diyi Yang,
Risa Wechsler,
Ioana Ciuca
Abstract:
Frontier AI agents show increasing promise as scientific research assistants, and may eventually be useful for extended, open-ended research workflows. However, in order to use agents for novel research, we must first assess the underlying faithfulness and correctness of their work. To evaluate agents as research assistants, we introduce ReplicationBench, an evaluation framework that tests whether…
▽ More
Frontier AI agents show increasing promise as scientific research assistants, and may eventually be useful for extended, open-ended research workflows. However, in order to use agents for novel research, we must first assess the underlying faithfulness and correctness of their work. To evaluate agents as research assistants, we introduce ReplicationBench, an evaluation framework that tests whether agents can replicate entire research papers drawn from the astrophysics literature. Astrophysics, where research relies heavily on archival data and computational study while requiring little real-world experimentation, is a particularly useful testbed for AI agents in scientific research. We split each paper into tasks which require agents to replicate the paper's core contributions, including the experimental setup, derivations, data analysis, and codebase. Each task is co-developed with the original paper authors and targets a key scientific result, enabling objective evaluation of both faithfulness (adherence to original methods) and correctness (technical accuracy of results). ReplicationBench is extremely challenging for current frontier language models: even the best-performing language models score under 20%. We analyze ReplicationBench trajectories in collaboration with domain experts and find a rich, diverse set of failure modes for agents in scientific research. ReplicationBench establishes the first benchmark of paper-scale, expert-validated astrophysics research tasks, reveals insights about agent performance generalizable to other domains of data-driven science, and provides a scalable framework for measuring AI agents' reliability in scientific research.
△ Less
Submitted 23 November, 2025; v1 submitted 28 October, 2025;
originally announced October 2025.
-
Lens Model Accuracy in the Expected LSST Lensed AGN Sample
Authors:
Padmavathi Venkatraman,
Sydney Erickson,
Phil Marshall,
Martin Millon,
Philip Holloway,
Simon Birrer,
Steven Dillmann,
Xiangyu Huang,
Sreevani Jaragula,
Ralf Kaehler,
Narayan Khadka,
Grzegorz Madejski,
Ayan Mitra,
Kevin Reil,
Aaron Roodman,
the LSST Dark Energy Science Collaboration
Abstract:
Strong gravitational lensing of active galactic nuclei (AGN) enables measurements of cosmological parameters through time-delay cosmography (TDC). With data from the upcoming LSST survey, we anticipate using a sample of O(1000) lensed AGN for TDC. To prepare for this dataset and enable this measurement, we construct and analyze a realistic mock sample of 1300 systems drawn from the OM10 (Oguri & M…
▽ More
Strong gravitational lensing of active galactic nuclei (AGN) enables measurements of cosmological parameters through time-delay cosmography (TDC). With data from the upcoming LSST survey, we anticipate using a sample of O(1000) lensed AGN for TDC. To prepare for this dataset and enable this measurement, we construct and analyze a realistic mock sample of 1300 systems drawn from the OM10 (Oguri & Marshall 2010) catalog of simulated lenses with AGN sources at $z<3.1$ in order to test a key aspect of the analysis pipeline, that of the lens modeling. We realize the lenses as power law elliptical mass distributions and simulate 5-year LSST i-band coadd images. From every image, we infer the lens mass model parameters using neural posterior estimation (NPE). Focusing on the key model parameters, $θ_E$ (the Einstein Radius) and $γ_{lens}$ (the projected mass density profile slope), with consistent mass-light ellipticity correlations in test and training data, we recover $θ_E$ with less than 1% bias per lens, 6.5% precision per lens and $γ_{lens}$ with less than 3% bias per lens, 8% precision per lens. We find that lens light subtraction prior to modeling is only useful when applied to data sampled from the training prior. If emulated deconvolution is applied to the data prior to modeling, precision improves across all parameters by a factor of 2. Finally, we combine the inferred lens mass models using Bayesian Hierarchical Inference to recover the global properties of the lens sample with less than 1% bias.
△ Less
Submitted 23 October, 2025;
originally announced October 2025.
-
Learning Representations of Event Time Series with Sparse Autoencoders for Anomaly Detection, Similarity Search, and Unsupervised Classification
Authors:
Steven Dillmann,
Juan Rafael Martínez-Galarza
Abstract:
Event time series are sequences of discrete events occurring at irregular time intervals, each associated with a domain-specific observational modality. They are common in domains such as high-energy astrophysics, computational social science, cybersecurity, finance, healthcare, neuroscience, and seismology. Their unstructured and irregular structure poses significant challenges for extracting mea…
▽ More
Event time series are sequences of discrete events occurring at irregular time intervals, each associated with a domain-specific observational modality. They are common in domains such as high-energy astrophysics, computational social science, cybersecurity, finance, healthcare, neuroscience, and seismology. Their unstructured and irregular structure poses significant challenges for extracting meaningful patterns and identifying salient phenomena using conventional techniques. We propose novel two- and three-dimensional tensor representations for event time series, coupled with sparse autoencoders that learn physically meaningful latent representations. These embeddings support a variety of downstream tasks, including anomaly detection, similarity-based retrieval, semantic clustering, and unsupervised classification. We demonstrate our approach on a real-world dataset from X-ray astronomy, showing that these representations successfully capture temporal and spectral signatures and isolate diverse classes of X-ray transients. Our framework offers a flexible, scalable, and generalizable solution for analyzing complex, irregular event time series across scientific and industrial domains.
△ Less
Submitted 10 October, 2025; v1 submitted 15 July, 2025;
originally announced July 2025.
-
Building Machine Learning Challenges for Anomaly Detection in Science
Authors:
Elizabeth G. Campolongo,
Yuan-Tang Chou,
Ekaterina Govorkova,
Wahid Bhimji,
Wei-Lun Chao,
Chris Harris,
Shih-Chieh Hsu,
Hilmar Lapp,
Mark S. Neubauer,
Josephine Namayanja,
Aneesh Subramanian,
Philip Harris,
Advaith Anand,
David E. Carlyn,
Subhankar Ghosh,
Christopher Lawrence,
Eric Moreno,
Ryan Raikman,
Jiaman Wu,
Ziheng Zhang,
Bayu Adhi,
Mohammad Ahmadi Gharehtoragh,
Saúl Alonso Monsalve,
Marta Babicz,
Furqan Baig
, et al. (126 additional authors not shown)
Abstract:
Scientific discoveries are often made by finding a pattern or object that was not predicted by the known rules of science. Oftentimes, these anomalous events or objects that do not conform to the norms are an indication that the rules of science governing the data are incomplete, and something new needs to be present to explain these unexpected outliers. The challenge of finding anomalies can be c…
▽ More
Scientific discoveries are often made by finding a pattern or object that was not predicted by the known rules of science. Oftentimes, these anomalous events or objects that do not conform to the norms are an indication that the rules of science governing the data are incomplete, and something new needs to be present to explain these unexpected outliers. The challenge of finding anomalies can be confounding since it requires codifying a complete knowledge of the known scientific behaviors and then projecting these known behaviors on the data to look for deviations. When utilizing machine learning, this presents a particular challenge since we require that the model not only understands scientific data perfectly but also recognizes when the data is inconsistent and out of the scope of its trained behavior. In this paper, we present three datasets aimed at developing machine learning-based anomaly detection for disparate scientific domains covering astrophysics, genomics, and polar science. We present the different datasets along with a scheme to make machine learning challenges around the three datasets findable, accessible, interoperable, and reusable (FAIR). Furthermore, we present an approach that generalizes to future machine learning challenges, enabling the possibility of large, more compute-intensive challenges that can ultimately lead to scientific discovery.
△ Less
Submitted 29 March, 2025; v1 submitted 3 March, 2025;
originally announced March 2025.
-
A Poisson Process AutoDecoder for X-ray Sources
Authors:
Yanke Song,
Victoria Ashley Villar,
Juan Rafael Martinez-Galarza,
Steven Dillmann
Abstract:
X-ray observing facilities, such as the Chandra X-ray Observatory and the eROSITA, have detected millions of astronomical sources associated with high-energy phenomena. The arrival of photons as a function of time follows a Poisson process and can vary by orders-of-magnitude, presenting obstacles for common tasks such as source classification, physical property derivation, and anomaly detection. P…
▽ More
X-ray observing facilities, such as the Chandra X-ray Observatory and the eROSITA, have detected millions of astronomical sources associated with high-energy phenomena. The arrival of photons as a function of time follows a Poisson process and can vary by orders-of-magnitude, presenting obstacles for common tasks such as source classification, physical property derivation, and anomaly detection. Previous work has either failed to directly capture the Poisson nature of the data or only focuses on Poisson rate function reconstruction. In this work, we present Poisson Process AutoDecoder (PPAD). PPAD is a neural field decoder that maps fixed-length latent features to continuous Poisson rate functions across energy band and time via unsupervised learning. PPAD reconstructs the rate function and yields a representation at the same time. We demonstrate the efficacy of PPAD via reconstruction, regression, classification and anomaly detection experiments using the Chandra Source Catalog.
△ Less
Submitted 4 February, 2025; v1 submitted 3 February, 2025;
originally announced February 2025.
-
Hyperluminous Supersoft X-Ray Sources in the Chandra Catalog
Authors:
Andrea Sacchi,
Kevin Paggeot,
Steven Dillmann,
Juan Rafael Martinez-Galarza,
Peter Kosec
Abstract:
Hyperluminous supersoft X-ray sources, such as bright extragalactic sources characterized by particularly soft X-ray spectra, offer a unique opportunity to study accretion onto supermassive black holes in extreme conditions. Examples of hyperluminous supersoft sources are tidal disruption events, systems exhibiting quasi-periodic eruptions, changing-look AGN, and anomalous nuclear transients. Alth…
▽ More
Hyperluminous supersoft X-ray sources, such as bright extragalactic sources characterized by particularly soft X-ray spectra, offer a unique opportunity to study accretion onto supermassive black holes in extreme conditions. Examples of hyperluminous supersoft sources are tidal disruption events, systems exhibiting quasi-periodic eruptions, changing-look AGN, and anomalous nuclear transients. Although these objects are rare phenomena amongst the population of X-ray sources, we developed an efficient algorithm to identify promising candidates exploiting archival observations. In this work, we present the results of a search for hyperluminous supersoft X-ray sources in the recently released Chandra catalog of serendipitous X-ray sources. This archival search has been performed via both a manual implementation of the algorithm we developed and a novel machine-learning-based approach. This search identified a new tidal disruption event, which might have occurred in an intermediate-mass black hole. This event occurred between 2001 and 2002, making it one of the first tidal disruption events ever observed by Chandra.
△ Less
Submitted 20 April, 2025; v1 submitted 31 January, 2025;
originally announced February 2025.
-
Humanity's Last Exam
Authors:
Long Phan,
Alice Gatti,
Ziwen Han,
Nathaniel Li,
Josephina Hu,
Hugh Zhang,
Chen Bo Calvin Zhang,
Mohamed Shaaban,
John Ling,
Sean Shi,
Michael Choi,
Anish Agrawal,
Arnav Chopra,
Adam Khoja,
Ryan Kim,
Richard Ren,
Jason Hausenloy,
Oliver Zhang,
Mantas Mazeika,
Dmitry Dodonov,
Tung Nguyen,
Jaeho Lee,
Daron Anderson,
Mikhail Doroshenko,
Alun Cennyth Stokes
, et al. (1133 additional authors not shown)
Abstract:
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce Humanity's Last Exam (HLE), a multi-modal benchmark at the frontier of…
▽ More
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce Humanity's Last Exam (HLE), a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. HLE consists of 2,500 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. HLE is developed globally by subject-matter experts and consists of multiple-choice and short-answer questions suitable for automated grading. Each question has a known solution that is unambiguous and easily verifiable, but cannot be quickly answered via internet retrieval. State-of-the-art LLMs demonstrate low accuracy and calibration on HLE, highlighting a significant gap between current LLM capabilities and the expert human frontier on closed-ended academic questions. To inform research and policymaking upon a clear understanding of model capabilities, we publicly release HLE at https://lastexam.ai.
△ Less
Submitted 28 July, 2026; v1 submitted 24 January, 2025;
originally announced January 2025.
-
Representation Learning for Time-Domain High-Energy Astrophysics: Discovery of Extragalactic Fast X-ray Transient XRT 200515
Authors:
Steven Dillmann,
Juan Rafael Martínez-Galarza,
Roberto Soria,
Rosanne Di Stefano,
Vinay L. Kashyap
Abstract:
We present a novel representation learning method for downstream tasks like anomaly detection, unsupervised classification, and similarity searches in high-energy data sets. This enabled the discovery of a new extragalactic fast X-ray transient (FXT) in Chandra archival data, XRT 200515, a needle-in-the-haystack event and the first Chandra FXT of its kind. Recent serendipitous discoveries in X-ray…
▽ More
We present a novel representation learning method for downstream tasks like anomaly detection, unsupervised classification, and similarity searches in high-energy data sets. This enabled the discovery of a new extragalactic fast X-ray transient (FXT) in Chandra archival data, XRT 200515, a needle-in-the-haystack event and the first Chandra FXT of its kind. Recent serendipitous discoveries in X-ray astronomy, including FXTs from binary neutron star mergers and an extragalactic planetary transit candidate, highlight the need for systematic transient searches in X-ray archives. We introduce new event file representations, E-t maps and E-t-dt cubes, that effectively encode both temporal and spectral information, enabling the seamless application of machine learning to variable-length event file time series. Our unsupervised learning approach employs PCA or sparse autoencoders to extract low-dimensional, informative features from these data representations, followed by clustering in the embedding space with DBSCAN. New transients are identified within transient-dominant clusters or through nearest-neighbour searches around known transients, producing a catalogue of 3559 candidates (3447 flares and 112 dips). XRT 200515 exhibits unique temporal and spectral variability, including an intense, hard <10s initial burst, followed by spectral softening in an ~800s oscillating tail. We interpret XRT 200515 as either the first giant magnetar flare observed at low X-ray energies or the first extragalactic Type I X-ray burst from a faint, previously unknown low-mass X-ray binary in the LMC. Our method extends to data sets from other observatories such as XMM-Newton, Swift-XRT, eROSITA, Einstein Probe, and upcoming missions like AXIS.
△ Less
Submitted 3 March, 2025; v1 submitted 2 December, 2024;
originally announced December 2024.