Search & Query
Overview
The vCon MCP server now exposes a recommended unified search surface plus the older specialized search tools.
For new clients, prefer vcon_search first. It gives you one predictable response envelope, explicit include groups, cursor pagination, and response-size budgeting. The older search tools are still available and still useful for compatibility or specialized flows.
Recommended Starting Point
vcon_search - Unified Search
Best for: New clients that want one search entry point with predictable parsing
Modes:
metadatakeywordsemantichybrid
Key advantages:
Stable
{ok, items, page}envelopeExplicit
includegroups such ascore,summary,tags, anddealerCursor pagination instead of mixed ad hoc paging behavior
Explicit
max_response_bytesso oversized responses fail loudly withRESPONSE_TOO_LARGE
Example:
Typical companion tools:
vcon_capabilitiesto discover supported modes, includes, and byte budgetsvcon_taxonomyto discover the portal taxonomy and preferred dealer sourcevcon_fetchto expand a selected result with additional include groups
Legacy Search Tools
The tools below remain supported and useful. They are especially helpful for older clients, direct low-level access, or cases where you intentionally want their narrower behavior.
Available Search Tools
1. search_vcons - Basic Filter Search
Best for: Finding vCons by metadata (subject, parties, dates)
Searches:
Subject line
Party names, emails, phone numbers
Creation dates
Does NOT search:
Dialog content
Analysis content
Attachments
Example:
Returns: Complete vCon objects matching the filters
2. search_vcons_content - Keyword Search
Best for: Finding specific words or phrases in conversation content
Searches:
✅ Subject
✅ Dialog bodies (conversations, transcripts)
✅ Analysis bodies (summaries, sentiment, etc.)
✅ Party information (names, emails, phones)
❌ Attachments (not indexed for full-text search)
Features:
Full-text search with ranking
Typo tolerance via trigram indexing
Highlighted snippets in results
Tag filtering support
Date range filtering
Example:
Returns: Ranked results with snippets showing where matches were found
Result format:
3. search_vcons_semantic - AI-Powered Semantic Search
Best for: Finding conversations by meaning, not just keywords
Searches:
✅ Subject (embedded)
✅ Dialog bodies (embedded)
✅ Analysis bodies with
encoding='none'orNULL(embedded)❌ Analysis with
encoding='base64url'orencoding='json'(not embedded)❌ Attachments (not embedded)
Features:
Finds conceptually similar content
Works across paraphrases and synonyms
AI embeddings using 384-dimensional vectors
Tag filtering support
Similarity threshold control
Requirements:
Embeddings must be generated first (see embedding documentation)
Currently requires pre-computed embedding vector (384 dimensions)
Example:
Returns: Similar conversations ranked by semantic similarity
4. search_vcons_hybrid - Combined Keyword + Semantic Search
Best for: Comprehensive search combining exact matches and conceptual similarity
Searches:
Everything from keyword search (subject, dialog, analysis, parties)
Everything from semantic search (embedded content)
Features:
Combines full-text and semantic search
Adjustable weighting between keyword and semantic results
Best of both worlds: exact matches + conceptual matches
Tag filtering support
Example:
Parameters:
semantic_weight: 0-1 (default 0.6)0.0 = 100% keyword search
1.0 = 100% semantic search
0.6 = 60% semantic, 40% keyword (recommended)
Returns: Combined results with both keyword and semantic scores
What About Attachments?
Current Status
Attachments are NOT indexed for search in the current implementation.
Why?
Binary content: Many attachments contain binary data (PDFs, images, audio) that isn't suitable for text-based search
Encoding: Attachments with
encoding='base64url'contain encoded data, not searchable textStructured data: Attachments with
encoding='json'contain structured data that produces poor quality embeddings
Special Case: Tags
Attachments of type tags with encoding='json' ARE used for filtering, but not for content search.
Example tags attachment:
These tags can be used with the tags parameter in any search tool:
Future Enhancements
Potential future support for attachment content search:
Text extraction: Extract text from PDFs, Word docs, etc.
Audio transcription: Transcribe audio attachments to searchable text
OCR: Extract text from images
Selective indexing: Index only attachments with text content
If you need to search attachment content, consider:
Extracting text and adding it as an analysis element
Adding a summary of attachment content as an analysis
Using attachment metadata in tags
Analysis Encoding and Search
Analysis Elements ARE Searchable
Analysis elements are included in search, with filtering based on encoding:
none or NULL
✅ Yes
✅ Yes
Plain text content, ideal for search
json
✅ Yes
❌ No
Included in keyword search only
base64url
✅ Yes
❌ No
Included in keyword search only
Why Filter Semantic Search by Encoding?
Analysis with encoding='none' contains human-readable text like:
Conversation summaries
Transcriptions
Sentiment analysis results
Translation output
Natural language insights
These are ideal for semantic search because they contain meaningful natural language.
Analysis with encoding='json' or encoding='base64url' typically contains:
Structured data (poor quality embeddings)
Binary content (not suitable for embeddings)
Encoded data (not searchable as text)
Search Comparison
Subject
✅ Depends on mode
✅ Filter
✅ Search
✅ Search
✅ Search
Dialog
✅ Depends on mode/include
❌
✅ Search
✅ Search
✅ Search
Analysis
✅ Depends on mode/include
❌
✅ Search
✅ (encoding=none)
✅ All
Attachments
✅ Via explicit include groups on fetch/search results
❌
❌
❌
❌
Party Info
✅ Depends on mode/include
✅ Filter
✅ Search
❌
✅ Search
Tags
✅ Filter and return
❌
✅ Filter
✅ Filter
✅ Filter
Ranking
✅ Depends on mode
❌
✅ Relevance
✅ Similarity
✅ Combined
Snippets
❌
❌
✅ Yes
❌
❌
Requires Embeddings
Only for semantic/hybrid mode
❌
❌
✅
⚠️ Optional
Cursor Pagination
✅ Yes
❌
❌
❌
❌
Response Budgeting
✅ max_response_bytes
❌
❌
❌
❌
Best Practices
When to Use Each Tool
vcon_search: Default choice for new clients"Use one parser for metadata, keyword, semantic, and hybrid search"
"Return lightweight summaries plus dealer info"
"Fail loudly instead of silently returning oversized payloads"
search_vcons: Quick metadata lookups in older clients"Find vCons with party email john@example.com"
"Show me vCons from last week"
"List vCons with subject containing 'urgent'"
search_vcons_content: Keyword-based content search"Find conversations mentioning 'refund'"
"Search for 'technical support' in dialog"
"Find analysis containing 'positive sentiment'"
search_vcons_semantic: Concept-based search"Find conversations where customer was unhappy"
"Show me calls about payment issues"
"Find similar conversations to this one"
search_vcons_hybrid: Comprehensive search"Find all billing-related conversations" (gets both exact matches and related topics)
"Search for customer complaints" (finds variations and synonyms)
Best when you want both precision and recall
Performance Tips
Use filters: Date ranges and tags can dramatically reduce search scope
Set appropriate limits: Start with smaller limits (10-20) for faster results
Choose the right tool: Don't use semantic search if keyword search is sufficient
Pre-generate embeddings: Semantic search requires embeddings to be generated beforehand
Generating Embeddings
For semantic and hybrid search to work effectively, you need to generate embeddings for your vCons.
See the following guides:
INGEST_AND_EMBEDDINGS.md - Complete guide to embedding generation
EMBEDDING_STRATEGY_UPGRADE.md - Details on which content is embedded
Quick start:
Troubleshooting
"No results found" for content search
Check that the content exists in dialog or analysis
Try a simpler query (fewer words)
Use wildcards or partial words
Check date range filters
"Embedding generation not yet implemented"
Semantic search currently requires pre-computed embeddings
Use
search_vcons_contentfor keyword search insteadGenerate embeddings using the scripts in
/scripts/
"Embedding must be 384 dimensions"
The system uses 384-dimensional embeddings
If you're providing embeddings, ensure they match this dimension
Use
text-embedding-3-smallwithdimensions=384(OpenAI)Or use
sentence-transformers/all-MiniLM-L6-v2(Hugging Face)
Poor search results
For keyword search: Try simpler, more specific terms
For semantic search: Ensure embeddings are up to date
For hybrid search: Adjust
semantic_weightparameterConsider using tags to filter results
Examples
Find customer complaints in dialog
Find high-priority sales conversations
Hybrid search with keyword emphasis
Find conversations similar to a specific vCon
Get the vCon's embedding from the database
Use it in semantic search:
Related Documentation
Getting Started - Getting started with vCon MCP
Ingest and Embeddings - Embedding generation
Search Optimization Guide - Database search performance
Last updated