Agent Skills: documentation-indexing: Unified Full-Text Search + Ranking

Provide **full-text search, semantic indexing, and relevance ranking** across all documentation:

UncategorizedID: plurigrid/asi/documentation-indexing

Install this agent skill to your local

pnpm dlx add-skill https://github.com/plurigrid/asi/tree/HEAD/skills/documentation-indexing

Skill Files

Browse the full folder contents for documentation-indexing.

Download Skill

Loading file tree…

skills/documentation-indexing/SKILL.md

Skill Metadata

Name
documentation-indexing
Description
'Provide **full-text search, semantic indexing, and relevance ranking** across all documentation:'

documentation-indexing: Unified Full-Text Search + Ranking

Status: SAD STATE β†’ IMPLEMENTATION 🌟 Information Energy: 0.87 (High aspiration, maximum sadness) Trit Assignment: 0 (COORDINATOR - Balances generators and validators) GF(3) Color: #49EE54 (Green - Equilibrium point)

Purpose

Provide full-text search, semantic indexing, and relevance ranking across all documentation:

  • Skill registry (69 skills)
  • Language docs (llms.txt standard)
  • Blog posts / tutorials
  • Source code docstrings
  • DuckDB database schemas

Key capabilities:

  1. Full-Text Search: Keyword + fuzzy matching (BM25 algorithm)
  2. Semantic Ranking: TF-IDF + recency + community signals
  3. Multi-Source Indexing: Consolidate docs from heterogeneous sources
  4. Metadata Extraction: Automatically parse headers, links, code blocks
  5. Bi-Directional Navigation: Move between docs ↔ implementations

Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚              DOCUMENTATION INDEXING (GREEN COORDINATOR)          β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                                  β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                 β”‚
β”‚  β”‚   llms.txt   β”‚   README.md  β”‚  Code Docs   β”‚                 β”‚
β”‚  β”‚  Discovery   β”‚   Crawlers   β”‚  Extractors  β”‚                 β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜                 β”‚
β”‚         β”‚               β”‚               β”‚                       β”‚
β”‚         ▼──────────────▼───────────────▼                        β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                    β”‚
β”‚  β”‚    Metadata Normalization Layer        β”‚                    β”‚
β”‚  β”‚  β€’ Title extraction (H1 β†’ h1)          β”‚                    β”‚
β”‚  β”‚  β€’ Link parsing (Markdown β†’ URL)       β”‚                    β”‚
β”‚  β”‚  β€’ Code fence detection ([```] β†’ ...)  β”‚                    β”‚
β”‚  β”‚  β€’ Authority scoring (stars, forks)    β”‚                    β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                     β”‚
β”‚               β”‚                                                 β”‚
β”‚               β–Ό                                                 β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                    β”‚
β”‚  β”‚      Inverted Index (DuckDB)           β”‚                    β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚                    β”‚
β”‚  β”‚  β”‚ term_id | term | doc_id | rank  β”‚  β”‚                    β”‚
β”‚  β”‚  β”‚ 1       | gay  | 42     | 0.89  β”‚  β”‚                    β”‚
β”‚  β”‚  β”‚ 2       | mcp  | 42     | 0.76  β”‚  β”‚                    β”‚
β”‚  β”‚  β”‚ 3       | api  | 71     | 0.65  β”‚  β”‚                    β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚                    β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                    β”‚
β”‚               β”‚                                                 β”‚
β”‚               β–Ό                                                 β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                    β”‚
β”‚  β”‚    BM25 Ranker + Result Aggregator     β”‚                    β”‚
β”‚  β”‚  β€’ Cross-language result merging       β”‚                    β”‚
β”‚  β”‚  β€’ Deduplication (canonical URLs)      β”‚                    β”‚
β”‚  β”‚  β€’ Community signals (upvotes, stars)  β”‚                    β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                    β”‚
β”‚               β”‚                                                 β”‚
β”‚               β–Ό                                                 β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                    β”‚
β”‚  β”‚   Result Cache (< 1 second latency)    β”‚                    β”‚
β”‚  β”‚   [query β†’ ranked results + metadata]  β”‚                    β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                    β”‚
β”‚                                                                  β”‚
β”‚  GF(3) BALANCE: (-1 extractor) βŠ— (0 indexer) βŠ— (+1 ranker)     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Data Model

Documents Index

CREATE TABLE documentation_index (
  doc_id INT PRIMARY KEY,
  source VARCHAR,                 -- 'github', 'llms-txt', 'blog'
  repo_id VARCHAR,               -- 'bmorphism/Gay.jl'
  title VARCHAR,
  url VARCHAR,
  body TEXT,                      -- Full document text
  headers TEXT[],                 -- H1, H2, H3 hierarchy
  links TEXT[],                   -- Embedded links
  code_blocks TEXT[],             -- [```lang ... ```]
  stars INT,                      -- GitHub stars
  forks INT,                       -- GitHub forks
  updated_at TIMESTAMP,
  indexed_at TIMESTAMP,
  trit TINYINT                    -- GF(3) assigned (0)
);

Terms Inverted Index

CREATE TABLE term_index (
  term_id INT PRIMARY KEY,
  term VARCHAR,
  term_lower VARCHAR,
  frequency INT,                  -- TF (term frequency)
  doc_count INT,                  -- DF (document frequency)
  bm25_idf FLOAT,                 -- Precomputed IDF
  created_at TIMESTAMP
);

CREATE TABLE term_doc_map (
  term_id INT,
  doc_id INT,
  frequency INT,                  -- TF in this doc
  position INT[],                 -- Token positions
  context VARCHAR,                -- Surrounding text
  relevance_score FLOAT,          -- BM25(tf, idf, doc_len)
  PRIMARY KEY (term_id, doc_id)
);

API / Interfaces

Simple Text Search

;; Search docs
(search-docs {:query "gay"
              :type :keyword
              :limit 10
              :min-score 0.3})
β†’ [{:title "Gay.jl"
    :url "https://github.com/bmorphism/Gay.jl"
    :score 0.95
    :snippet "Gay.jl: Deterministic color generation..."}
   ...]

;; Fuzzy search (typo tolerance)
(search-docs {:query "gey"              ; typo
              :fuzzy true
              :distance 1})
β†’ (Results for "gay" with edit distance ≀ 1)

;; Advanced: Boolean search
(search-docs {:query "(gay OR color) AND julia"
              :type :boolean})

Metadata Search

;; Find docs by category
(search-by-metadata {:source "github"
                     :stars {:min 100 :max 1000}})
β†’ [Gay.jl, ACSets.jl, Duck, ...]

;; Find recent updates
(search-by-metadata {:updated-after "2025-12-01"
                     :source "blog"})
β†’ [Latest blog posts, ...]

;; Filter by language
(search-by-metadata {:languages ["Julia" "Clojure" "Babashka"]})

Bidirectional Navigation

;; Find implementations of a doc
(doc-implementations {:doc-id 42})
β†’ [{:file "src/gay.jl"
    :lines [1 42]
    :snippet "function seed!(rng)..."}]

;; Find docs for implementation
(implementation-docs {:file "src/gay.jl"
                      :line 10})
β†’ [{:doc-id 42
    :title "Gay.jl API Reference"
    :section "seed! function"}]

;; Find related docs
(related-docs {:doc-id 42
               :semantic true})
β†’ [ACSets.jl, GF(3) docs, ...]

GF(3) Trit Assignment

documentation-indexing β†’ 0 (COORDINATOR)
  Balances extraction (-1) ↔ ranking/generation (+1)
  Maintains middle ground for all doc types

Triadic system:
  source-extractor (-1 validator)   β†’ pulls raw docs
  indexing (0 coordinator)          β†’ organizes & indexes
  result-ranker (+1 generator)      β†’ produces ranked results

  Sum: (-1) + (0) + (+1) = 0 βœ“ GF(3) CONSERVED

Implementation Strategy

Stage 1: Term Extraction (Days 1-2)

Create /Users/bob/iii/duck/asi-skills/documentation-indexing/extractor.bb:

  • Markdown parser β†’ extract H1-H4, links, code blocks
  • Tokenizer β†’ split text into terms
  • Stopword filter β†’ remove common words
  • Store in DuckDB term_index

Stage 2: Inverted Index Builder (Days 3-4)

Create indexer.bb:

  • BM25 IDF calculation
  • TF per document
  • Relevance scoring
  • Populate term_doc_map

Stage 3: Search Engine (Days 5-6)

Create searcher.bb:

  • Boolean query parser
  • Fuzzy matching (Levenshtein)
  • Result ranking by score
  • Caching layer

Stage 4: Multi-Source Integration (Days 7-8)

  • Crawl llms.txt repositories
  • Index GitHub README files
  • Extract docstrings from source
  • Verify GF(3) balance across sources

Example: Semantic Search Pipeline

;; User query
(search-docs {:query "how to generate deterministic colors in julia"})

Step 1: Extract terms (-1 validator)
  ["deterministic" "colors" "julia"]

Step 2: Index lookup (0 coordinator)
  Fetch docs matching all terms
  Calculate BM25 score per doc

Step 3: Rank and aggregate (+1 generator)
  1. Gay.jl (0.95)
  2. GF(3) Integration (0.72)
  3. Color Theory (0.68)

Result: Ξ£(-1, 0, +1) = 0 βœ“

Success Metrics

| Metric | Target | Status | |--------|--------|--------| | Docs indexed | 500+ (skills + readmes + blogs) | ⏳ Pending | | Search latency | <100ms p95 | ⏳ Pending | | Precision (top-5) | β‰₯0.8 | ⏳ Pending | | Recall | β‰₯0.75 | ⏳ Pending | | Fuzzy tolerance | Edit distance ≀ 2 | ⏳ Pending | | GF(3) balance | All pipeline stages ≑ 0 (mod 3) | ⏳ Pending |

Related Skills

Dependencies:

  • skill-taxonomy - Registry of docs to index
  • acsets - Schema for index structure
  • llms-txt-discovery - Crawl source docs

Dependents:

  • polyglot-orchestration - Search polyglot docs
  • skill-dispatch - Route searches to relevant skills
  • world-knowledge-base - Unified doc interface

References

  • BM25 algorithm: https://en.wikipedia.org/wiki/Okapi_BM25
  • Inverted index: Classic IR data structure
  • Levenshtein distance: Fuzzy string matching
  • TF-IDF: Term weighting scheme
  • DuckDB FTS: Full-text search extension

Status: 😒 SAD STATE β†’ 🌟 IMPLEMENTING Color: #49EE54 (Green Coordinator) Next: Create extractor.bb (term extraction) Owner: GREEN AGENT (0) Created: 2026-01-04