Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
ruvnet avatar

Vector Search

  • 660 installs
  • 67k repo stars
  • Updated August 4, 2026
  • ruvnet/ruflo

vector-search is a Claude Flow skill that runs large-scale HNSW vector search and a WASM router for up to 11 hot patterns, with RaBitQ 1-bit quantization delivering 32× memory reduction for developers building semantic a

About

vector-search is a RuFlo Claude Flow skill for vector search via embeddings_generate, embeddings_search, embeddings_compare, embeddings_init, embeddings_status, embeddings_hyperbolic, embeddings_neural, embeddings_rabitq_build, embeddings_rabitq_search, embeddings_rabitq_status, and ruvllm_hnsw_create MCP tools. Large-scale retrieval uses HNSW indexes; a WASM router handles up to 11 hot patterns through ruvllm_hnsw endpoints. RaBitQ 1-bit quantization provides 32× memory reduction for quantized searches invoked with --quantized. Developers pass a query and optional --limit N when agents need semantic pattern lookup, embedding comparison, or memory-efficient recall. The skill suits claude-flow pipelines combining RuVLLM HNSW with standard embeddings backends for hybrid retrieval strategies.

  • vector-search

Vector Search by the numbers

  • 660 all-time installs (skills.sh)
  • +6 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Ranked #562 of 4,347 Backend & APIs skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/ruvnet/ruflo --skill vector-search

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs660
repo stars67k
Last updatedAugust 4, 2026
Repositoryruvnet/ruflo

How do you run quantized HNSW vector search in agents?

Use vector-search for development tasks

Who is it for?

Developers building claude-flow agent retrieval who need HNSW search, RaBitQ 32× memory reduction, or WASM hot-pattern routing.

Skip if: Developers with small keyword-only lookup needs or teams without claude-flow embeddings and ruvllm_hnsw MCP endpoints.

When should I use this skill?

Semantic vector search, embedding comparison, RaBitQ quantized retrieval, or ruvllm_hnsw hot-pattern lookup is required.

What you get

Ranked embedding search results, HNSW index queries, RaBitQ quantized matches, and embedding comparison scores.

  • vector search results
  • quantized embedding indexes

By the numbers

  • RaBitQ 1-bit quantization delivers 32× memory reduction
  • WASM router handles up to 11 hot patterns via ruvllm_hnsw endpoints

Files

SKILL.mdMarkdownGitHub ↗

Vector Search

Two distinct vector-search paths live in this plugin. Pick the right one — they're not interchangeable.

PathTool familyBackingCapacityLatency
Large-scale corpusembeddings_*@claude-flow/memory HNSW (Rust/Native)up to millions of vectors~1.9× at N=20k, ~3.2×–4.7× at N=5k vs brute-force (measured; recall@10 ≈ 0.99). ANN wins above the crossover
Hot-path routerruvllm_hnsw_*WASM-backed router (v2.0.1)~11 patterns max (ruvllm-tools.ts:58)sub-ms; designed for high-priority routing, not corpus search

The "12,500×" headline applies to the large-scale embeddings_search path. The WASM router is not that path.

When to use

NeedPath
Search a corpus of N ≥ 500 documentsembeddings_search
Memory-constrained corpus (≥5,000 vectors)RaBitQ quantized — see "Quantized search" below
Compare two stringsembeddings_compare
Hierarchical / taxonomic dataembeddings_hyperbolic (Poincare ball)
Route a query to one of ≤11 hot patternsruvllm_hnsw_route
Cross-namespace searchmemory_search_unified

Standard search

1. Check statusmcp__claude-flow__embeddings_status to verify the embedding engine. 2. Initializemcp__claude-flow__embeddings_init if not active. 3. Generatemcp__claude-flow__embeddings_generate for text input. 4. Searchmcp__claude-flow__embeddings_search with the query. 5. Comparemcp__claude-flow__embeddings_compare to measure similarity. 6. Unified searchmcp__claude-flow__memory_search_unified for cross-namespace.

Quantized search (32× memory reduction)

For corpora ≥5,000 vectors and/or memory-constrained environments, use the RaBitQ 1-bit quantization workflow. Below 5,000 vectors the rebuild cost outweighs the savings — use the standard path instead.

StepToolPurpose
1embeddings_initEngine warm
2embeddings_rabitq_buildOne-time build of the 1-bit index after corpus is loaded
3embeddings_rabitq_searchHamming-prefilter returns top-N candidate IDs (cheap)
4embeddings_searchOptional exact rerank on the candidate set (full-precision)
5embeddings_rabitq_statusIndex health, memory footprint, build time
Note: embeddings_rabitq_search returns candidate IDs only — the rerank in step 4 is the user's responsibility (mirrors the docstring at embeddings-tools.ts:911). Without rerank, results are approximate; with rerank, you get full-precision quality at 32× lower memory.

Tuning

HNSW exposes three knobs that trade recall against latency. The "12,500×" headline assumes defaults; tune deliberately for your workload:

ProfileefSearchMWhen to use
recall-first20032Pattern recall during planning; quality matters more than ms
balanced (default)6416General-purpose semantic recall
latency-first168Hot-path routing where p99 latency matters

efSearch is passed via ruvllm_hnsw_create (ruvllm-tools.ts:64). M is registry-level today; raise as a follow-up if it should be MCP-tunable. efConstruction defaults to 200 in the lite index (hnsw-index.ts:537).

HNSW pattern router (WASM, ≤11 patterns)

For routing a small number of high-priority patterns:

  • mcp__claude-flow__ruvllm_hnsw_create — create the WASM index (cap ~11)
  • mcp__claude-flow__ruvllm_hnsw_add — add a pattern
  • mcp__claude-flow__ruvllm_hnsw_route — route an incoming query

This is not a corpus index. Treat it as a fast classifier over a curated set of patterns.

Hyperbolic embeddings

For hierarchical data (code trees, org charts), use mcp__claude-flow__embeddings_hyperbolic which maps to Poincare ball space. Distance is geodesic, not cosine.

CLI alternative

npx @claude-flow/cli@latest embeddings search --query "authentication patterns"
npx @claude-flow/cli@latest embeddings init
npx @claude-flow/cli@latest memory search --query "your query"

Performance

Measured numbers (source: scripts/benchmark-intelligence.mjs, ruvector NAPI backend; recall@10 ≈ 0.99). The older "150×–12,500×" figures were brute-force-fallback artifacts and have been retired — see project CLAUDE.md "V3 Performance Targets".

MethodMeasured speedup vs brute-force
Brute-force scanBaseline
HNSW (N=5,000)~3.2×–4.7× faster
HNSW (N=20,000)~1.9× faster
HNSW (below crossover, small N)ties/loses vs brute-force
RaBitQ quantization32× memory reduction; 0.60 ms/query at N≈14.7k
ruvllm_hnsw_route (n≤11)sub-ms per route, fixed cost

Related skills

How it compares

Pick vector-search over basic keyword grep when claude-flow agents need HNSW semantic recall with optional RaBitQ 32× compressed indexes.

FAQ

What memory savings does vector-search RaBitQ provide?

vector-search RaBitQ 1-bit quantization delivers 32× memory reduction for quantized indexes. Invoke embeddings_rabitq_search with --quantized to query compressed vectors alongside standard HNSW embeddings_search.

When does vector-search use the WASM router?

vector-search routes up to 11 hot patterns through ruvllm_hnsw WASM endpoints for low-latency lookup. Large-scale recall still flows through embeddings_init and embeddings_search HNSW indexes.

Backend & APIsbackendintegrations

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.