AstaBrief Gives Asta An Open 8B Fast Mode For Scientific Reports
Ai2 released AstaBrief 8B as an open-weights report-generation model for Asta, with a one-pass pipeline that averaged 51.1 seconds per report in Fast mode.

Ai2 has released AstaBrief 8B, an open-weights model designed to turn research questions and retrieved literature excerpts into referenced scientific briefs, in a Fast mode now available inside its Asta research platform.
A Hugging Face blog post from the institute frames the release as a model-and-data package that researchers can examine, reproduce and adapt.
The project addresses a specific weak point in scientific AI workflows: long-form syntheses must stay close to the evidence they cite, not simply attach references to broad claims.
Ai2 built AstaBrief from Qwen3-8B and treated the work as a full pipeline problem, combining model training, data filtering, preference data and a redesigned serving path.
Speed was one of the operating tests.
AstaBrief writes a draft in one pass from a user query and relevant retrieved snippets, bypassing the section-by-section summarization and clustering stages used in Asta's Claude-powered Thinking mode.
Across the full Asta pipeline, Fast mode averaged 51.1 seconds per output, compared with 178.5 seconds for Thinking mode, making it about 3.5 times faster in the tracked setup.
The training data came from real Asta and ScholarQA usage rather than only synthetic benchmark prompts.
Ai2 filtered user logs for quality, privacy and relevance, removing beta-tester and bot traffic, very short prompts, non-English queries, non-scientific requests and prompts containing personal information.
That process left about 90K research-focused queries.
For supervised fine-tuning, the team used ScholarQA's retrieval and generation pipeline to produce full referenced syntheses from the filtered questions.
The source post identifies Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1 among the systems used to generate target outputs.
Ai2's quality-filtered corpus then contained 47K usable training examples.
Preference training contributed a smaller but more selective layer.
A separate query set produced competing outputs from ScholarQA's existing pipeline and from other models, including o3, o4-mini, DeepSeek-V3 and DeepSeek-R1.
GPT-4.1 and DeepSeek-R1 judged each pair, and Ai2 kept only examples where both judges agreed after checking judge alignment with human preferences.
The final direct-preference-optimization dataset contained about 6K examples.
Evaluation focused on output quality, relevance, structure and citation grounding, with SQABench-CS2 as the main benchmark and DeepScholarBench plus pairwise comparisons as secondary checks.
Early supervised runs improved general quality but remained behind the Claude pipeline on answer precision and citation quality.
Filtering low-citation-density synthetic examples produced the strongest gains among the data filters tested, while more complex filter combinations did not add meaningful improvement.
The finished model was competitive with the Claude-powered Asta pipeline and DR Tulu across several development measures, but the post narrows that claim.
Most training and evaluation work was completed in 2025, and Ai2 did not rerun the full comparison against current frontier models.
The result is presented as evidence for the training recipe and system design, not as a current leaderboard claim.
AstaBrief's first production signal is usage inside Asta.
Among 374 Asta users who tried Fast mode, 29.1% used it on two or more days, average usage reached 3.67 report threads, 23% continued with Fast mode without returning to Thinking mode for later threads, and another 18% switched between the two modes depending on the task.
Feedback was sparse, but positive rates were close: 84.2% for Fast mode and 85.2% for Thinking mode.
Open weights change the deployment path as much as the interface.
Institutions can run AstaBrief on their own infrastructure when research questions involve sensitive or unpublished work, and Ai2 is releasing an example workflow for generating briefs from local PDFs.
The remaining research agenda is sharper evaluation of whether referenced claims preserve the scope of the underlying evidence, especially when a synthesis could turn a sample-specific result into a broader population claim or a descriptive finding into a recommendation.




















