open source · MIT · written in Rust

Context retrieval
for AI agents.

rawq is hybrid semantic + lexical code search in a single offline binary. Your agent asks a question in plain English and gets back only the code that answers it, with file paths, line ranges, scope names and confidence scores.

$ curl -fsSL https://raw.githubusercontent.com/auyelbekov/rawq/main/scripts/install.sh | sh
~/acme-api
The problem

Agents waste most of their context looking for code.

An agent that doesn't know where to look has to grep for keywords and read entire files to find a single function. That burns tokens, slows every task, and fills the context window with noise.

grep + read today

  • Keyword matches miss code that uses different names for the same idea
  • Common words return hundreds of hits across dozens of files
  • Whole files get pulled into context to find one relevant block
  • The agent can't tell how relevant any match actually is

rawq instead

  • Ask by meaning: "how do failed connections get retried"
  • Get ranked, AST-aware chunks instead of whole files
  • Each result carries file, lines, scope and confidence
  • Cap the output with --token-budget so results always fit
Features

Small binary. Serious retrieval.

Everything runs on your machine. No server, no API keys, no code leaving your laptop.

Hybrid search

ONNX embeddings fused with tantivy BM25. Meaning and exact identifiers in one query.

AST-aware chunks

Tree-sitter splits 16 languages into functions, classes and methods, not arbitrary line windows.

Incremental indexing

Per-chunk SHA-256 hashes and git-aware change detection. Only what changed gets re-embedded.

Fully offline

After the one-time model download, rawq makes zero network calls. Your code stays local.

Agent-native output

JSON, NDJSON streaming, token budgets and meaningful exit codes, built for tool calls.

GPU acceleration

CUDA, DirectML and CoreML when available, with automatic CPU fallback.

Warm daemon

Keeps the model loaded between queries so repeated searches stay fast. Exits after 30 minutes idle.

Diff, map, watch

Search only your git diff, print an AST map of the codebase, or re-index on every save.

How it works

From repository to ranked context.

rawq builds a local index once, keeps it fresh incrementally, and answers each query with two retrievers fused into a single ranking.

01

Chunk

Tree-sitter parses each file and splits it at real boundaries such as functions, methods and classes, keeping the scope name.

tree-sitter16 langs
02

Index

Every chunk is embedded locally and written to a full-text index. Unchanged chunks are skipped by hash.

ONNXtantivySHA-256
03

Retrieve

Semantic and BM25 results are fused into one ranking, with optional keyword re-ranking for precision.

hybrid--rerank
04

Return

Top chunks with paths, line ranges, scope, confidence and token counts, trimmed to your token budget.

--json--stream--token-budget
tree-sitter chunking for RustPythonTypeScriptJavaScriptGoJavaCC++RubySwiftDart+ more
Built for agents

A tool call your agent will actually use.

Drop rawq's SKILL.md into Claude Code or any agent that can run shell commands, and it learns when to reach for semantic search and when plain grep is enough.

  1. Orientrawq map . shows the structure of an unfamiliar repo.
  2. Searchrawq search "…" . --json returns the relevant chunks.
  3. ReadOpen only the top results' files for full context.
  4. ActEdit with the right code in context and nothing else.
0results found
1no matches
2error
$ rawq search "how do failed connections get retried" \
    . --json --top 1
{
  "schema_version": 1,
  "model": "snowflake-arctic-embed-s",
  "query_ms": 38,
  "total_tokens": 214,
  "results": [
    {
      "file": "src/db/client.rs",
      "lines": [142, 171],
      "language": "rust",
      "scope": "DatabaseClient.reconnect",
      "confidence": 0.91,
      "token_count": 214,
      "content": "async fn reconnect(&mut self) {\n…"
    }
  ]
}
Models

Pick your embedding model.

Small and fast by default. Switch with rawq model default or the RAWQ_MODEL environment variable.

ModelDimsMax tokensBest for
snowflake-arctic-embed-sdefault384512Fast everyday search on any machine
snowflake-arctic-embed-m-v1.5768512Higher-quality results
jina-embeddings-v2-base-code7688192Code-specialized, long chunks
Install

Install in one line.

Prebuilt binaries for macOS, Linux and Windows, or build from crates.io.

macOS / Linux

$curl -fsSL https://raw.githubusercontent.com/auyelbekov/rawq/main/scripts/install.sh | sh

Windows (PowerShell)

>powershell -ExecutionPolicy Bypass -c "irm https://raw.githubusercontent.com/auyelbekov/rawq/main/scripts/install.ps1 | iex"

Cargo

$cargo install rawq
$cargo install rawq --features cuda

Then run rawq "what you're looking for" . in any repository. The index builds on first search.

in development

rawq for teams.

The CLI stays free and open source. We're building the pieces teams need to run rawq across many repositories and many agents.

Shared indexes

Build once in CI and share across every developer and agent. No cold starts on large monorepos.

Hosted retrieval endpoint

A retrieval API for cloud and background agents that can't run a local binary.

Cross-repo search

One query across services, libraries and docs, with access controls that match your org.