rawq is hybrid semantic + lexical code search in a single offline binary. Your agent asks a question in plain English and gets back only the code that answers it, with file paths, line ranges, scope names and confidence scores.
curl -fsSL https://raw.githubusercontent.com/auyelbekov/rawq/main/scripts/install.sh | sh
An agent that doesn't know where to look has to grep for keywords and read entire files to find a single function. That burns tokens, slows every task, and fills the context window with noise.
"how do failed connections get retried"file, lines, scope and confidence--token-budget so results always fitEverything runs on your machine. No server, no API keys, no code leaving your laptop.
ONNX embeddings fused with tantivy BM25. Meaning and exact identifiers in one query.
Tree-sitter splits 16 languages into functions, classes and methods, not arbitrary line windows.
Per-chunk SHA-256 hashes and git-aware change detection. Only what changed gets re-embedded.
After the one-time model download, rawq makes zero network calls. Your code stays local.
JSON, NDJSON streaming, token budgets and meaningful exit codes, built for tool calls.
CUDA, DirectML and CoreML when available, with automatic CPU fallback.
Keeps the model loaded between queries so repeated searches stay fast. Exits after 30 minutes idle.
Search only your git diff, print an AST map of the codebase, or re-index on every save.
rawq builds a local index once, keeps it fresh incrementally, and answers each query with two retrievers fused into a single ranking.
Tree-sitter parses each file and splits it at real boundaries such as functions, methods and classes, keeping the scope name.
Every chunk is embedded locally and written to a full-text index. Unchanged chunks are skipped by hash.
Semantic and BM25 results are fused into one ranking, with optional keyword re-ranking for precision.
Top chunks with paths, line ranges, scope, confidence and token counts, trimmed to your token budget.
Drop rawq's SKILL.md into Claude Code or any agent that can run shell commands, and it learns when to reach for semantic search and when plain grep is enough.
rawq map . shows the structure of an unfamiliar repo.rawq search "…" . --json returns the relevant chunks.$ rawq search "how do failed connections get retried" \ . --json --top 1 { "schema_version": 1, "model": "snowflake-arctic-embed-s", "query_ms": 38, "total_tokens": 214, "results": [ { "file": "src/db/client.rs", "lines": [142, 171], "language": "rust", "scope": "DatabaseClient.reconnect", "confidence": 0.91, "token_count": 214, "content": "async fn reconnect(&mut self) {\n…" } ] }
Small and fast by default. Switch with rawq model default or the RAWQ_MODEL environment variable.
| Model | Dims | Max tokens | Best for |
|---|---|---|---|
snowflake-arctic-embed-sdefault | 384 | 512 | Fast everyday search on any machine |
snowflake-arctic-embed-m-v1.5 | 768 | 512 | Higher-quality results |
jina-embeddings-v2-base-code | 768 | 8192 | Code-specialized, long chunks |
Prebuilt binaries for macOS, Linux and Windows, or build from crates.io.
curl -fsSL https://raw.githubusercontent.com/auyelbekov/rawq/main/scripts/install.sh | shpowershell -ExecutionPolicy Bypass -c "irm https://raw.githubusercontent.com/auyelbekov/rawq/main/scripts/install.ps1 | iex"cargo install rawqcargo install rawq --features cudaThen run rawq "what you're looking for" . in any repository. The index builds on first search.
The CLI stays free and open source. We're building the pieces teams need to run rawq across many repositories and many agents.
Build once in CI and share across every developer and agent. No cold starts on large monorepos.
A retrieval API for cloud and background agents that can't run a local binary.
One query across services, libraries and docs, with access controls that match your org.