AIREITER

AstaBrief 8B vLLM Local Deployment for Private RAG

Last Updated: 2026-10-02 19:06:21

AstaBrief 8B is not a retriever: it turns a research question plus supplied scientific excerpts into a cited report. That boundary is the key to a private deployment. Run the generator behind your firewall, keep document parsing and search local, and pass only ranked evidence with stable source IDs into the model.

What AstaBrief 8B actually needs

AstaBrief 8B is an 8-billion-parameter text-generation model from Ai2, based on Qwen3-8B and released under Apache 2.0. The official model card defines its input as a research question and retrieved scientific-literature excerpts, not a question that the model searches for by itself. It also warns that changing the fine-tuned prompt or interaction format can produce degraded or inconsistent behavior.

Use the final allenai/AstaBrief_8B checkpoint for this guide. The model card's example contains allenai/AstaBrief_8B_SFT, which is the supervised-fine-tuning predecessor. Treat that name mismatch as a documentation detail to verify against the checkpoint you download, rather than silently assuming the two models are interchangeable.

Ai2's public ScholarQA repository is useful for understanding the reference design: retrieval, optional reranking, paper-level aggregation, quote extraction, and report generation are separate components. A private implementation should preserve that separation even if it replaces Semantic Scholar with an internal index.

A reproducible local architecture

A private pipeline should have six explicit stages:

  1. Ingest: parse PDFs, OCR scanned pages, and retain document ID, title, page, section, and character offsets.
  2. Chunk: split text into moderately sized passages without discarding page boundaries or headings.
  3. Retrieve: combine lexical search with embeddings when terminology, identifiers, or exact phrases matter.
  4. Rerank: score the first-stage candidates against the complete question and keep a small evidence set.
  5. Assemble: assign immutable citation IDs and format the excerpts using AstaBrief's expected reference structure.
  6. Generate: send the assembled prompt to the local vLLM endpoint.

The important design choice is not a particular vector database. It is the evidence contract: every passage sent to the model must carry a stable ID that your application can map back to a document and page.

Keep citation IDs stable

Use IDs such as DOC_014_P07_A rather than array positions. Array positions change when retrieval settings change; a document-page-span ID remains auditable.

Store the mapping outside the prompt:

{
  "DOC_014_P07_A": {
    "document": "internal_protocol.pdf",
    "page": 7,
    "section": "Methods",
    "char_start": 18420,
    "char_end": 19210
  }
}

In the prompt, show the same ID beside the excerpt. After generation, reject or flag citations that are not in the supplied ID set. This does not prove that a cited passage supports every claim, but it prevents the simplest form of citation fabrication.

Choose retrieval depth before context assembly

Do not dump every matching chunk into the context. Retrieve broadly enough for recall, rerank, then pack only passages that fit the model's prompt budget and answer the question directly. Merge adjacent chunks from the same page when that preserves a complete argument, but keep separate IDs if the final report needs page-level traceability.

The Ai2 ScholarQA repository describes a reference configuration that retrieves 256 candidates, reranks them, and retains 50 paper-level results. Those values belong to that public pipeline, not to a universal AstaBrief requirement. Start with smaller private collections, inspect missed evidence, and tune retrieval and context limits against your own questions.

Serve AstaBrief 8B with vLLM

The official materials do not publish a single VRAM requirement for every dtype, context length, and concurrency level. Begin with the unquantized checkpoint on hardware that can load it, then reduce concurrency or context length before evaluating a quantized build; do not assume quantization preserves citation behavior without testing your corpus.

Install vLLM in a clean environment that matches your CUDA and PyTorch stack. Then launch the OpenAI-compatible server with the final checkpoint:

pip install -U vllm openai

vllm serve allenai/AstaBrief_8B \
  --host 127.0.0.1 \
  --port 8000 \
  --dtype auto \
  --max-model-len 16000

The --max-model-len value is an operational ceiling, not a promise that every request should contain 16,000 tokens; the model card reports a 16,000-token training maximum while its example uses max_tokens=4096 for generation.

Verify the local endpoint

curl http://127.0.0.1:8000/v1/models

Then call it with the same prompt style used by the fine-tuning data. The official example uses temperature 0.7, top-p 0.95, a maximum of 4,096 generated tokens, and the tokenizer's EOS token as a stopping condition.

from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:8000/v1",
    api_key="local-only",
)

response = client.chat.completions.create(
    model="allenai/AstaBrief_8B",
    temperature=0.7,
    top_p=0.95,
    max_tokens=4096,
    messages=[
        {"role": "user", "content": assembled_prompt},
    ],
)

print(response.choices[0].message.content)

For a private server, bind to loopback or an internal interface, place authentication and TLS at your gateway, and block public ingress. A local model does not automatically make logs, temporary PDF files, or tracing private.

Build the private retrieval request

The following skeleton deliberately leaves the index implementation open. The critical parts are the ranked evidence list, immutable IDs, and the prompt boundary.

from dataclasses import dataclass
from openai import OpenAI

@dataclass
class Evidence:
    ref_id: str
    text: str
    title: str
    page: int


def build_prompt(question: str, evidence: list[Evidence]) -> str:
    references = "\n\n".join(
        f"[{item.ref_id}] {item.title} (page {item.page})\n{item.text}"
        for item in evidence
    )
    return f"""Research question:
{question}

Retrieved references:
{references}

Write a cited research report answering the question. Use only the supplied
references for factual support. Attach the supplied reference IDs to claims.
If the references do not establish a point, say that the evidence is
insufficient instead of inventing a source.
"""


def answer(question: str, evidence: list[Evidence]) -> str:
    client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="local-only")
    prompt = build_prompt(question, evidence)
    result = client.chat.completions.create(
        model="allenai/AstaBrief_8B",
        temperature=0.7,
        top_p=0.95,
        max_tokens=4096,
        messages=[{"role": "user", "content": prompt}],
    )
    return result.choices[0].message.content

For production, use the official AstaBrief prompt template and insert your retrieved excerpts while preserving their reference IDs. The shortened instruction above illustrates the vLLM integration; it is not a substitute for the fine-tuning format.

Your private index can use BM25, dense retrieval, or a hybrid. Preserve metadata through every stage. A passage without its document ID and page number is not enough for a defensible research report, even if the generated prose sounds correct.

Add a citation and privacy gate before production

AstaBrief's reported ScholarQA-CS2 results are useful context, not a guarantee for your corpus. On the model card's 100-question test set, AstaBrief 8B reports citation precision of 90.5 and citation recall of 78.2; answer precision is 89.0. Citation recall being lower than citation precision is a practical warning: a report can cite supplied material accurately while still omitting relevant evidence.

Use a gate with four checks:

  1. Reference validity: every generated citation ID exists in the request's allow-list.
  2. Metadata resolution: every ID resolves to a document, page, and stored text span.
  3. Evidence support: a reviewer or separate verifier checks whether the passage actually supports the nearby claim.
  4. Retrieval recall: maintain a small labeled question set containing expected documents and pages, then measure misses after changes to chunking, embeddings, or reranking.

Keep the model server, index, object storage, logs, and monitoring inside the same trust boundary unless your policy explicitly permits an external service. Disable request-body logging for sensitive documents, redact queries from traces where possible, and define retention for uploaded PDFs and generated reports.

Do not reuse Ai2's reported 51.1-second Fast-mode figure as your local benchmark. That number describes the end-to-end Asta pipeline, while a self-hosted deployment changes the GPU, retrieval system, batching, prompt length, and network path. Measure retrieval, reranking, time to first token, generation time, and total request time separately.

Troubleshoot the deployment by layer

SymptomLikely layerFirst check
The server loads but output is poorly citedPrompt or evidence contractCompare the prompt with the official format and verify stable reference IDs
Citations point to nothingApplication validationReject IDs absent from the request allow-list
Relevant papers are missingRetrievalEvaluate chunking, hybrid search, and reranker recall before changing the model
Requests run out of memoryServing or context packingLower concurrent requests, output budget, or packed context; then revisit quantization
Local latency is unexpectedly highWhole pipelineTime retrieval, reranking, queueing, and generation independently
Private data appears in logsOperationsInspect gateway, vLLM, tracing, cache, and object-storage retention settings

This layered diagnosis prevents a retrieval miss from being misdiagnosed as a model failure, and prevents a prompt-format mismatch from being “fixed” by adding more documents.

Use AstaBrief 8B when you want a specialized local report generator and are prepared to own retrieval quality, source mapping, and validation. vLLM solves model serving; it does not supply document search, citation provenance, or privacy controls.

Sources