Dual-Engine Schema Pruning & Tool Selection

Calibrated tool selection for AI agents.

Prune 100 tool schemas to top candidates in 0.4ms offline, or direct-dispatch with calibrated confidence.

0.4ms offline TurboQuant -92% prompt tokens 100% BFCL Top-1 (TypeSafe) Zero external dependencies
Live Interactive Engine

Playground

Test real TurboQuant vector search in your browser. Type an intent or choose an example query below.

Try:
Candidate Mode:
Direct Dispatch Threshold: 85%
0.4ms zero-dep heuristic

Pruned Candidates (Auto)

Adaptive score drop-off

Real-Time Pruning Impact

Search Latency
0.24 ms
Offline vector search
Prompt Reduction
-91.2% tokens
Compared to 30 schemas
Unpruned Catalog ~4,680 tokens
tool-prune Candidate Set ~450 tokens

// Loading snippet...
            
Financial & Latency ROI

How much does schema pruning save?

Agents that pass entire MCP catalogs into each prompt burn thousands of redundant tokens every turn. Adjust the sliders below to calculate your savings.

5,000
500 25,000 50,000
40 tools
10 tools 60 tools 120 tools
Note: Assumes 160 tokens per tool schema (uncached input), Top-3 candidate pruning, and ~6.3ms LLM prefill latency per pruned schema (measured on 60 tools: 555ms → 195ms).

Estimated Monthly Savings

Monthly Tokens Saved
888M

Prompt tokens avoided

Monthly Cost Saved
$2,664

Direct LLM invoice savings

Total Agent Hours Saved
9.7h
~234ms saved/step (93% schema reduction)

Beyond token cost, schema pruning accelerates agent execution: empirical evaluation with Claude Haiku 4.5 confirms 97.5% selection accuracy is preserved while cutting tokens by 72% and slashing latency from 555ms down to 195ms (149ms via direct dispatch).

Empirical Evaluation

Berkeley Function Calling Benchmark (BFCL v3)

Evaluated on the Gorilla BFCL v3 multiple-tool dataset across 100 tool schemas and subtle distractor intents.

Engine Backend Distractor Top-1 100-Tool Top-5 Prune Search Latency Network
TypeSafe (Jev) Cloud System One 100.0% 100.0% 189 ms Cloud API
TurboQuant (WASM) JS WASM (turboquant-search) 86.7% 85.0% 16 ms Offline
TurboQuant (Python) Rust SIMD (turbovec) 85.0% 82.5% 0.018 ms Offline
TurboQuant (Pure JS) Pure JS / Python Fallback 86.7% 77.5% 0.14 ms Offline
BM25 Baseline Lexical Inverted Index 88.3% 95.0% 0.025 ms Offline
Agent Orchestration Evals

End-to-End Agent Performance (60 Tools)

Empirical evaluation with Claude Haiku 4.5 across 60 production tool schemas and 79 queries, comparing Pruned JSON calling against Code Mode and Full-Context prompts.

Architecture Selection Accuracy Latency (P50) Tokens / Query Roundtrips Execution Model
Tool-Prune Direct 97.5% 149 ms 385 tokens 0 turns LLM Bypassed (91% fast-path)
Tool-Prune + LLM 97.5% 195 ms 397 tokens 1 turn Top-3 Pruned Schema Prompt
Full-Context LLM 97.5% 555 ms 1,409 tokens 1 turn All 60 Tool Schemas in Prompt
Tool-Search (BM25 + LLM) 87.3% 979 ms 436 tokens 2 turns Lexical Retrieval Retrieval Loss
Code Mode (search + eval) 64.6% 2,209 ms 1,735 tokens 1.9 turns Multi-Turn Discovery Drop
Why Code Mode is not enough for atomic tool dispatch: While Code Mode is effective for large-scale data manipulation (loops and aggregations inside isolated micro-runtimes), forcing discrete API calls through two-stage dynamic discovery (search_tools → execute_code) causes a 64.6% accuracy collapse due to keyword discovery misses, adds an 11x latency tax (2,209ms vs 195ms), and consumes 4.3x more tokens than calibrated schema pruning.
Dual-Engine Design

Architecture & Workflow

Three simple stages to eliminate agent prompt bloat and achieve deterministic sub-millisecond execution.

1

FWHT Vector Projection

Tool criteria and query tokens are hashed and rotated using a diagonal Rademacher matrix and the Fast Walsh-Hadamard Transform in O(d log d), eliminating data-dependent training overhead.

Zero network • Pure local math
2

PolarQuant & QJL Residual

Vectors are quantized to 2-bit discrete intervals, while residual error vectors are captured with 1-bit Johnson-Lindenstrauss random signs. Preserves high cosine fidelity with tiny memory footprints.

Calibrated [0, 1] probability range
3

Fast-Path Direct Dispatch

When confidence clears your threshold, dispatch deterministic tool handlers directly in under 160ms. If ambiguous, forward only top-5 candidates to the LLM prompt.

Save 92% tokens and LLM cost