
Sf Datacloud Retrieve
- 941 installs
- 423 repo stars
- Updated April 27, 2026
- jaganpro/sf-skills
sf-datacloud-retrieve is an agent skill that generates and manages hybrid search indexes on Salesforce Data Cloud DMOs for developers who need RAG-style retrieval inside Salesforce agents and CRM-integrated applications.
About
sf-datacloud-retrieve in jaganpro/sf-skills is part of the sf-datacloud-* family for Salesforce Data Cloud retrieval setup. The skill guides creation of hybrid search indexes on structured Data Cloud DMOs, including chunk and vector DMO naming patterns such as INDEX_NAME_chunk and INDEX_NAME_index linked to a source DMO developer name ending in __dlm. Developers reach for sf-datacloud-retrieve when building agent retrieval over CRM objects that need both keyword and vector search rather than a single-mode index. Configuration covers index labels, developer names, descriptions, source DMO selection, and search type settings required for hybrid RAG pipelines inside Salesforce. The skill fits teams wiring Einstein or custom agents to grounded enterprise data without exporting DMOs to an external vector database. Triggers include Data Cloud hybrid index, RAG on DMOs, Salesforce retrieval setup, and chunk or vector DMO provisioning.
- Creates hybrid search indexes combining keyword and vector search on structured Data Cloud DMOs
- Configures automatic chunking with passage_extraction (max_tokens: 512, strip_html: true)
- Uses e5_large_v2 embedding model with 1024 dimensions and HNSW indexing
- Generates required chunk and vector DMOs with proper naming conventions
- Part of the reusable sf-datacloud-* skill family for Salesforce AI agents
Sf Datacloud Retrieve by the numbers
- 941 all-time installs (skills.sh)
- +4 installs in the week ending Aug 5, 2026 (Skillselion tracking)
- Ranked #1,169 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Security screen: LOW risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/jaganpro/sf-skills --skill sf-datacloud-retrieveAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 941 |
|---|---|
| repo stars | ★ 423 |
| Security audit | 3 / 3 scanners passed |
| Last updated | April 27, 2026 |
| Repository | jaganpro/sf-skills ↗ |
How do you build hybrid search on Data Cloud DMOs?
Generate and manage hybrid search indexes on Salesforce Data Cloud DMOs for RAG-style retrieval inside agents.
Who is it for?
Salesforce developers implementing RAG retrieval over Data Cloud structured DMOs for agents and CRM-integrated apps.
Skip if: Non-Salesforce vector stores or simple SOQL queries without Data Cloud hybrid index provisioning.
When should I use this skill?
User asks to create Data Cloud hybrid search indexes, RAG on DMOs, chunk or vector DMO setup, or sf-datacloud retrieval configuration.
What you get
Configured hybrid search index, chunk DMO, vector index DMO, and retrieval metadata wired to a source __dlm DMO.
- Hybrid search index configuration
- Chunk and vector DMO definitions
By the numbers
- Part of the sf-datacloud-* skill family for Data Cloud retrieval
Files
sf-datacloud-retrieve: Data Cloud Retrieve Phase
Use this skill when the user needs query, search, and metadata introspection for Data Cloud: sync SQL, paginated SQL, async query workflows, table describe, vector search, hybrid search, or search index operations.
When This Skill Owns the Task
Use sf-datacloud-retrieve when the work involves:
sf data360 query *sf data360 search-index *sf data360 metadata *sf data360 profile *orsf data360 insight *inspection- understanding Data Cloud SQL results or query shape
Delegate elsewhere when the user is:
- writing standard CRM SOQL only → sf-soql
- designing segment or calculated insight assets → sf-datacloud-segment
- analyzing STDM/session tracing/parquet telemetry → sf-ai-agentforce-observability
---
Required Context to Gather First
Ask for or infer:
- target org alias
- whether the user needs quick count, medium result set, large export, schema inspection, or semantic search
- table/index name if known
- whether the task is read-only SQL or search-index lifecycle management
---
Core Operating Rules
- Treat Data Cloud SQL as its own query language, not SOQL.
- Run the shared readiness classifier before relying on query/search surfaces:
node ~/.claude/skills/sf-datacloud/scripts/diagnose-org.mjs -o <org> --phase retrieve --json. - Use describe before guessing columns.
- Prefer
sqlv2or async query flows for larger result sets. - Use vector search or hybrid search only when the search index lifecycle is healthy.
- Keep STDM/parquet/session-tracing workflows out of this skill family.
---
Recommended Workflow
1. Classify readiness for retrieve work
node ~/.claude/skills/sf-datacloud/scripts/diagnose-org.mjs -o <org> --phase retrieve --json
# optional query-plane probe, only with a real table name
node ~/.claude/skills/sf-datacloud/scripts/diagnose-org.mjs -o <org> --phase retrieve --describe-table MyDMO__dlm --json2. Choose the smallest correct query shape
sf data360 query sql -o <org> --sql 'SELECT COUNT(*) FROM "ssot__Individual__dlm"' 2>/dev/null
sf data360 query sqlv2 -o <org> --sql 'SELECT * FROM "ssot__Individual__dlm"' 2>/dev/null
sf data360 query async-create -o <org> --sql 'SELECT * FROM "ssot__Individual__dlm"' 2>/dev/null3. Use describe before guessing fields
sf data360 query describe -o <org> --table ssot__Individual__dlm 2>/dev/null4. Use vector or hybrid search only when an index exists
sf data360 search-index list -o <org> 2>/dev/null
sf data360 query vector -o <org> --index Knowledge_Index --query "reset password" --limit 5 2>/dev/null
sf data360 query hybrid -o <org> --index Knowledge_Index --query "reset password" --limit 5 2>/dev/null
sf data360 query hybrid -o <org> --index Insurance_Index --query "weather damage coverage" --prefilter "Type_of_Insurance__c='Home'" --limit 10 2>/dev/null5. Reuse curated search-index examples when creating indexes
Use the phase-owned examples instead of inventing JSON from scratch:
examples/search-indexes/vector-knowledge.jsonexamples/search-indexes/hybrid-structured.json
---
High-Signal Gotchas
- Data Cloud SQL is not SOQL.
- Table names should be double-quoted in SQL.
sqlv2is better than ad hoc OFFSET paging for medium result sets.- async query is preferable for large results.
- search-index operations and vector/hybrid queries depend on the index lifecycle being healthy.
- Hybrid search can use
--prefilter, but only on fields configured as prefilter-capable when the search index was created. - HNSW index parameters are typically read-only on create; leave
userValues: []unless the platform explicitly documents otherwise. query describeis not a universal tenant probe; only run it with a known DMO or DLO table after broader readiness has been confirmed.
---
Output Format
Retrieve task: <sql / sqlv2 / async / describe / vector / search-index>
Target org: <alias>
Target object: <table or index>
Commands: <key commands run>
Verification: <query rows / schema / status>
Next step: <segment / harmonize / follow-up>---
References
- README.md
- examples/search-indexes/vector-knowledge.json
- examples/search-indexes/hybrid-structured.json
- ../sf-datacloud/assets/definitions/search-index.template.json
- ../sf-datacloud/references/plugin-setup.md
- ../sf-datacloud/references/feature-readiness.md
Credits & Acknowledgments
Primary contributor: Gnanasekaran Thoppae
This skill is part of the sf-datacloud-* family. Shared attribution, upstream source mapping, and maintenance notes live in:
- ../sf-datacloud/CREDITS.md
- ../sf-datacloud/UPSTREAM.md
{
"label": "<INDEX_NAME>",
"developerName": "<INDEX_NAME>",
"description": "Hybrid search index on a structured Data Cloud DMO",
"sourceDmoDeveloperName": "<SOURCE_DMO>__dlm",
"chunkDmoName": "<INDEX_NAME> chunk",
"chunkDmoDeveloperName": "<INDEX_NAME>_chunk",
"vectorDmoName": "<INDEX_NAME> index",
"vectorDmoDeveloperName": "<INDEX_NAME>_index",
"searchType": "HYBRID",
"vectorEmbedding": {
"vectorEmbeddingRelatedFields": []
},
"rankingConfigurations": [],
"chunkingConfiguration": {
"fieldLevelConfigurations": [
{
"sourceDmoDeveloperName": "<SOURCE_DMO>__dlm",
"sourceDmoFieldDeveloperName": "<TEXT_FIELD>__c",
"config": {
"id": "passage_extraction",
"userValues": [
{ "id": "max_tokens", "value": "512" },
{ "id": "strip_html", "value": "true" }
]
}
}
]
},
"vectorEmbeddingConfiguration": {
"embeddingModel": {
"id": "e5_large_v2",
"userValues": [
{ "id": "dimension", "value": "1024" },
{ "id": "max_token_limit", "value": "512" }
]
},
"index": {
"id": "HNSW",
"userValues": []
},
"similarityMetric": "COSINE"
}
}
{
"label": "My_kav",
"developerName": "My_kav",
"sourceDmoDeveloperName": "ssot__KnowledgeArticleVersion__dlm",
"chunkDmoName": "My_kav chunk",
"chunkDmoDeveloperName": "My_kav_chunk",
"vectorDmoName": "My_kav index",
"vectorDmoDeveloperName": "My_kav_index",
"searchType": "VECTOR",
"vectorEmbedding": {
"vectorEmbeddingRelatedFields": []
},
"chunkingConfiguration": {
"fieldLevelConfigurations": [
{
"sourceDmoDeveloperName": "ssot__KnowledgeArticleVersion__dlm",
"sourceDmoFieldDeveloperName": "ssot__Name__c",
"config": {
"id": "passage_extraction",
"userValues": [
{ "id": "strip_html", "value": "true" },
{ "id": "max_tokens", "value": "512" }
]
}
}
]
},
"vectorEmbeddingConfiguration": {
"embeddingModel": {
"id": "e5_large_v2",
"userValues": [
{ "id": "dimension", "value": "1024" },
{ "id": "max_token_limit", "value": "512" }
]
},
"index": {
"id": "HNSW",
"userValues": []
},
"similarityMetric": "COSINE"
},
"rankingConfigurations": []
}
MIT License
Copyright (c) 2024-2025 Jag Valaiyapathy
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
sf-datacloud-retrieve
Query and search workflows for Salesforce Data Cloud.
Use this skill for
- quick SQL counts
- paginated SQL (
sqlv2) - async query lifecycles
- table describe
- vector search
- hybrid search with optional prefilter
- search index inspection and lifecycle work
Example requests
"Run a Data Cloud SQL query against unified profiles"
"Describe this Data Cloud table before I write SQL"
"Help me troubleshoot vector search in Data Cloud"
"Run a hybrid search with a prefilter in Data Cloud"
"Create and inspect a search index"Common commands
sf data360 query sql -o myorg --sql 'SELECT COUNT(*) FROM "ssot__Individual__dlm"' 2>/dev/null
sf data360 query describe -o myorg --table ssot__Individual__dlm 2>/dev/null
sf data360 search-index list -o myorg 2>/dev/null
sf data360 query vector -o myorg --index Knowledge_Index --query "reset password" --limit 5 2>/dev/null
sf data360 query hybrid -o myorg --index Knowledge_Index --query "reset password" --limit 5 2>/dev/nullExample payloads
- examples/search-indexes/vector-knowledge.json
- examples/search-indexes/hybrid-structured.json
References
- SKILL.md
- ../sf-datacloud/assets/definitions/search-index.template.json
- CREDITS.md
License
MIT License - See LICENSE.
Related skills
How it compares
Pick sf-datacloud-retrieve for hybrid indexes on Data Cloud DMOs; pick external vector DB skills when CRM data is exported outside Salesforce.
FAQ
What objects does sf-datacloud-retrieve configure?
sf-datacloud-retrieve configures hybrid search indexes on Data Cloud DMOs, including chunk DMOs and vector index DMOs named from the index developer name and linked to a source __dlm DMO.
What search mode does sf-datacloud-retrieve target?
sf-datacloud-retrieve targets hybrid search indexes that combine keyword and vector retrieval on structured Data Cloud DMOs for RAG-style grounding inside Salesforce agents.
Is Sf Datacloud Retrieve safe to install?
skills.sh reports 3 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.