AIREITER

AstaBrief 8B vs OpenScholar vs PaperQA2: Workflow Comparison

Last Updated: 2026-10-02 19:15:10

AstaBrief 8B, OpenScholar, and PaperQA2 are often placed in the same scientific-AI shortlist, but they start at different points in the research workflow. OpenScholar is the strongest default for broad, citation-backed literature synthesis; PaperQA2 is the more practical choice for a controlled collection of PDFs; AstaBrief 8B is the speed-focused report writer you add after retrieval is already working. That distinction matters more than the model-size labels.

Three systems, three starting points

SystemStarts withRetrieval scopeMain outputBest first use
AstaBrief 8BA research question plus retrieved excerptsSupplied by the surrounding Asta/ScholarQA workflow or your own pipelineA fast, cited reportGenerate a report quickly after retrieval
OpenScholarA scientific questionOpenScholar DataStore, retrievers, rerankers, and other retrieval sourcesLong-form, citation-backed synthesisSearch and synthesize across broad literature
PaperQA2A question plus a document directory or corpusLocal indexing, metadata search, embeddings, reranking, and agentic searchEvidence-grounded answer with inline citationsAsk repeatable questions over controlled documents

Larger built-in corpora trade preparation time for discovery breadth; user-supplied files trade breadth for evidence control.

AstaBrief 8B: optimized for fast report generation after retrieval

AstaBrief 8B is a compact report-generation model released by Ai2. The official announcement describes its input as a research question and retrieved literature excerpts, rather than an independent paper-search service. The model is available under an Apache 2.0 model license, and Ai2 also released training data and an example workflow for reports from researchers’ own PDFs.

Ai2 reports that the full Asta pipeline averaged 51.1 seconds in Fast mode, compared with 178.5 seconds for its Thinking mode, or about 3.5× faster in that comparison. That is a pipeline measurement, not a universal inference-speed promise for every local deployment.

AstaBrief fits pipelines that already provide retrieved passages, stable citation IDs, and a claim-support review step. It should not be treated as a standalone replacement for OpenScholar’s datastore and retrieval stack: incomplete excerpts mean the report writer cannot recover papers that were never supplied.

For a deeper look at installation and local deployment, see our AstaBrief 8B review and local deployment guide.

OpenScholar: a broad literature-synthesis workflow

OpenScholar uses the OpenScholar DataStore, a corpus of 45 million open-access papers and 236 million passage embeddings, with trained retrievers, rerankers, citation-backed generation, and a self-feedback loop (Nature).

OpenScholar is the best starting point when the question looks like this:

  • “What does the literature say across several subfields?”
  • “Compare methods across a large and changing research area.”
  • “Find the relevant papers before writing the synthesis.”

The Nature paper reports that the combined retrieval configuration reached 49.6 correctness and 47.6 citation F1 in its computer-science ablation, outperforming web-only and Semantic Scholar-only retrieval. A broad datastore, retriever, reranker, and feedback loop require more infrastructure than a small report model; running an 8B checkpoint without those components does not recreate the full system.

PaperQA2: the best fit for a controlled local document set

PaperQA2 is designed around document-grounded research. Its workflow can parse PDFs, Markdown, HTML, Office files, source code, images, and tables; retrieve candidate passages; create contextual summaries; rerank evidence; and produce an answer with citations. The package supports local models and LiteLLM-compatible providers, although its documented defaults use hosted models and embeddings.

PaperQA2 lets you point the system at a directory, build a reusable index, and adjust the model, embedding, evidence count, and source limits. The repository documents a default of 10 evidence passages and 5 maximum answer sources, while its high-quality setting uses more evidence and costs more.

That makes PaperQA2 attractive for:

  • A lab’s private or unpublished PDFs.
  • A project-specific evidence bundle.
  • Repeated questions over the same document collection.
  • Workflows that need PDF figures, tables, metadata, or page references.

The trade-off is provider and configuration dependence. PaperQA2 is open source, but the final quality and cost depend on the selected LLM, embedding model, parser, corpus, and settings. The official repository does not promise a single fixed per-question price.

“The default configuration uses OpenAI models and a NumPy vector database, but the package is designed to support alternative LLMs, embedding models, vector stores, and local servers.” — FutureHouse, PaperQA2 documentation

Same research task, different workflows

For broad discovery across an unfamiliar field: choose OpenScholar

Start with OpenScholar when you do not yet know which papers belong in the answer. Its large open datastore and specialized retrieval pipeline are built for finding and synthesizing evidence across many papers.

For a private project library: choose PaperQA2

Use PaperQA2 when the evidence boundary is the project directory. Index the papers, keep the source set stable, and tune the number of evidence passages and final sources. A reviewer can inspect the exact corpus used for the answer.

For a fast report after retrieval: choose AstaBrief 8B

Use AstaBrief after another component has done the discovery work. It is the cleanest choice when latency, local control, or report throughput matters more than building a complete literature-search system.

For a hybrid stack: retrieve with OpenScholar, write with AstaBrief

Assemble evidence with a scholarly retriever and reranker, pass the selected excerpts to AstaBrief 8B, and run a separate claim-support check before publication. This combines OpenScholar-style discovery with AstaBrief’s compact generation path, but creates an integration responsibility that no single component removes.

What the published benchmark actually proves

The Nature and PMC versions of the OpenScholar paper provide the clearest same-benchmark comparison with PaperQA2. On ScholarQABench, the reported multi-paper Scholar-CS rubric score was 51.1 for OpenScholar-8B and 45.6 for PaperQA2. PaperQA2 had a slight citation-F1 edge in that table, 48.0 versus 47.9 for OpenScholar-8B. The paper reports PaperQA2’s estimated cost at $0.30–$2.30 per question, compared with $0.003 for OpenScholar-8B under the paper’s cost assumptions.

Published ScholarQABench resultOpenScholar-8BPaperQA2
Scholar-CS rubric score51.145.6
Scholar-CS citation F147.948.0
Estimated cost per question$0.003$0.30–$2.30

These numbers are not a universal leaderboard. The paper evaluated PaperQA2 using the OpenScholar DataStore because PaperQA2’s own datastore was not public. The evaluation also used a particular model and configuration. Most importantly, the published comparison does not include AstaBrief 8B on the same ScholarQABench setup. AstaBrief’s reported speed and quality results come from Ai2’s Asta-focused evaluations, so those numbers cannot establish its rank against OpenScholar and PaperQA2.

The paper also notes that PaperQA2’s answers had high citation accuracy but often relied on only one or a few papers, reducing coverage and multi-paper organization. Citation precision and literature coverage are separate objectives.

A reliable selection rule

Your constraintRecommended choiceWhy
Discover unknown literature at broad scaleOpenScholarRetrieval and synthesis are integrated
Keep the evidence set inside a local folderPaperQA2Corpus, index, and settings are user-controlled
Produce a report quickly from supplied evidenceAstaBrief 8BCompact, one-pass report generation
Run behind a firewall with open weightsAstaBrief 8B or OpenScholar-8BOpen model artifacts; the surrounding stack still matters
Audit every claim against a known passagePaperQA2 or a hybridLocal indexing and explicit evidence controls help review
Minimize cost where local inference is viableOpenScholar-8B; test AstaBrief separatelyThe published cost figures are not directly comparable

Before trusting any of the three, verify the retrieval boundary, citation support, literature coverage, claim scope, and reproducibility of the corpus and settings.

FAQ

Is AstaBrief 8B a replacement for OpenScholar?

No. AstaBrief 8B is primarily a report-generation model that consumes retrieved excerpts, while OpenScholar includes a scientific datastore, retrieval, reranking, and self-feedback workflow. AstaBrief can replace part of the generation layer, not automatically the complete discovery stack.

Does AstaBrief 8B retrieve papers itself?

The official AstaBrief description frames the model as taking a research question and retrieved literature excerpts. Retrieval is supplied by the surrounding Asta/ScholarQA workflow or by a separate pipeline you build.

Which has the best citation accuracy?

The published ScholarQABench table gives PaperQA2 a marginal citation-F1 advantage over OpenScholar-8B, 48.0 versus 47.9, while OpenScholar-8B scores higher on the Scholar-CS rubric, 51.1 versus 45.6. There is no same-benchmark AstaBrief result in that comparison.

Can PaperQA2 run fully locally?

PaperQA2 supports local model servers, local embeddings, local indexes, and local document collections. A fully local deployment still depends on the parser, embedding model, LLM quality, hardware, and configuration you choose.

Which is best for a systematic review?

None should replace a registered review protocol, human screening, and citation verification. For discovery, OpenScholar is the better broad-search starting point; for a predefined evidence library, PaperQA2 offers tighter corpus control; AstaBrief is useful for drafting after the evidence set has been assembled.

For the decision in one line: choose OpenScholar when finding the right literature is hardest, PaperQA2 when controlling a known corpus is hardest, and AstaBrief 8B when retrieval is solved and report latency is the bottleneck.