Prune 100 tool schemas to top candidates in 0.4ms offline, or direct-dispatch with calibrated confidence.
Test real TurboQuant vector search in your browser. Type an intent or choose an example query below.
// Loading snippet...
Agents that pass entire MCP catalogs into each prompt burn thousands of redundant tokens every turn. Adjust the sliders below to calculate your savings.
Prompt tokens avoided
Direct LLM invoice savings
Beyond token cost, schema pruning accelerates agent execution: empirical evaluation with Claude Haiku 4.5 confirms 97.5% selection accuracy is preserved while cutting tokens by 72% and slashing latency from 555ms down to 195ms (149ms via direct dispatch).
Evaluated on the Gorilla BFCL v3 multiple-tool dataset across 100 tool schemas and subtle distractor intents.
| Engine | Backend | Distractor Top-1 | 100-Tool Top-5 Prune | Search Latency | Network |
|---|---|---|---|---|---|
| TypeSafe (Jev) | Cloud System One | 100.0% | 100.0% | 189 ms | Cloud API |
| TurboQuant (WASM) | JS WASM (turboquant-search) | 86.7% | 85.0% | 16 ms | Offline |
| TurboQuant (Python) | Rust SIMD (turbovec) | 85.0% | 82.5% | 0.018 ms | Offline |
| TurboQuant (Pure JS) | Pure JS / Python Fallback | 86.7% | 77.5% | 0.14 ms | Offline |
| BM25 Baseline | Lexical Inverted Index | 88.3% | 95.0% | 0.025 ms | Offline |
Empirical evaluation with Claude Haiku 4.5 across 60 production tool schemas and 79 queries, comparing Pruned JSON calling against Code Mode and Full-Context prompts.
| Architecture | Selection Accuracy | Latency (P50) | Tokens / Query | Roundtrips | Execution Model |
|---|---|---|---|---|---|
| Tool-Prune Direct | 97.5% | 149 ms | 385 tokens | 0 turns | LLM Bypassed (91% fast-path) |
| Tool-Prune + LLM | 97.5% | 195 ms | 397 tokens | 1 turn | Top-3 Pruned Schema Prompt |
| Full-Context LLM | 97.5% | 555 ms | 1,409 tokens | 1 turn | All 60 Tool Schemas in Prompt |
| Tool-Search (BM25 + LLM) | 87.3% | 979 ms | 436 tokens | 2 turns | Lexical Retrieval Retrieval Loss |
| Code Mode (search + eval) | 64.6% | 2,209 ms | 1,735 tokens | 1.9 turns | Multi-Turn Discovery Drop |
search_tools → execute_code) causes a 64.6% accuracy collapse due to keyword discovery misses, adds an 11x latency tax (2,209ms vs 195ms), and consumes 4.3x more tokens than calibrated schema pruning.
Three simple stages to eliminate agent prompt bloat and achieve deterministic sub-millisecond execution.
Tool criteria and query tokens are hashed and rotated using a diagonal Rademacher matrix and the Fast Walsh-Hadamard Transform in O(d log d), eliminating data-dependent training overhead.
Vectors are quantized to 2-bit discrete intervals, while residual error vectors are captured with 1-bit Johnson-Lindenstrauss random signs. Preserves high cosine fidelity with tiny memory footprints.
When confidence clears your threshold, dispatch deterministic tool handlers directly in under 160ms. If ambiguous, forward only top-5 candidates to the LLM prompt.
Explore active schemas and add custom tool definitions