Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
jaganpro avatar

Sf Datacloud Prepare

  • 938 installs
  • 423 repo stars
  • Updated April 27, 2026
  • jaganpro/sf-skills

sf-datacloud-prepare is a Salesforce agent skill that formats and sends structured records into Salesforce Data Cloud via the Ingestion API for developers automating data loads from Python or agent workflows.

About

sf-datacloud-prepare belongs to the jaganpro sf-datacloud skill family and helps developers prepare structured records for Salesforce Data Cloud using the Ingestion API. Configuration spans consumer key and secret from an External Client App, Salesforce username, login URL, tenant URL, private key file, connector name, and object name variables. The folder includes minimal Python Ingestion API examples agents can adapt for badge scans or similar object loads. Engineers reach for this skill when application-generated events or batch records must land in Data Cloud without manual CSV uploads. It complements connector-based ingress skills by handling programmatic record preparation and authenticated API submission from scripts or agent-driven workflows inside Salesforce-centric stacks.

  • Minimal public-safe example for Ingestion API record submission
  • Handles JWT authentication with private key and consumer credentials
  • Includes schema upsert and connector setup instructions
  • Works with PyJWT and cryptography libraries
  • Part of the reusable sf-datacloud skill family

Sf Datacloud Prepare by the numbers

  • 938 all-time installs (skills.sh)
  • +4 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Ranked #415 of 4,347 Backend & APIs skills by installs in the Skillselion catalog
  • Security screen: LOW risk (skills.sh audit)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/jaganpro/sf-skills --skill sf-datacloud-prepare

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs938
repo stars423
Security audit2 / 3 scanners passed
Last updatedApril 27, 2026
Repositoryjaganpro/sf-skills

How do you ingest records into Salesforce Data Cloud?

Prepare and send structured records into Salesforce Data Cloud via the Ingestion API from Python scripts or agent workflows.

Who is it for?

Salesforce developers pushing structured application records into Data Cloud through the Ingestion API from Python or agent workflows.

Skip if: Teams only needing database ingress connectors without programmatic record submission or those avoiding External Client App OAuth setup.

When should I use this skill?

The user needs to prepare, format, or send records to Salesforce Data Cloud via the Ingestion API or Python examples.

What you get

Python Ingestion API payloads, OAuth environment configuration, and structured records loaded into named Data Cloud objects.

  • Ingestion API Python scripts
  • OAuth environment configuration
  • structured record payloads

By the numbers

  • Configures 7 environment variables including CONSUMER_KEY, TENANT_URL, and PRIVATE_KEY_FILE
  • Part of the sf-datacloud skill family with shared CREDITS.md and UPSTREAM.md

Files

SKILL.mdMarkdownGitHub ↗

sf-datacloud-prepare: Data Cloud Prepare Phase

Use this skill when the user needs ingestion and lake preparation work: data streams, Data Lake Objects (DLOs), transforms, Document AI, unstructured ingestion, or the handoff from connector setup into a live stream.

When This Skill Owns the Task

Use sf-datacloud-prepare when the work involves:

  • sf data360 data-stream *
  • sf data360 dlo *
  • sf data360 transform *
  • sf data360 docai *
  • choosing how data should enter Data Cloud
  • rerunning or rescanning ingestion after a source update
  • preparing Ingestion API-backed streams after connector setup is complete

Delegate elsewhere when the user is:

  • still creating/testing source connections → sf-datacloud-connect
  • mapping to DMOs or designing IR/data graphs → sf-datacloud-harmonize
  • querying ingested data → sf-datacloud-retrieve

---

Required Context to Gather First

Ask for or infer:

  • target org alias
  • source connection name
  • source object / dataset / document source
  • desired stream type
  • DLO naming expectations
  • whether the user is creating, updating, running, or deleting a stream
  • whether the source is CRM, a database connector, an unstructured file source, or an Ingestion API feed

---

Core Operating Rules

  • Verify the external plugin runtime before running Data Cloud commands.
  • Run the shared readiness classifier before mutating ingestion assets: node ~/.claude/skills/sf-datacloud/scripts/diagnose-org.mjs -o <org> --phase prepare --json.
  • Prefer inspecting existing streams and DLOs before creating new ingestion assets.
  • Suppress linked-plugin warning noise with 2>/dev/null for normal usage.
  • Treat DLO naming and field naming as Data Cloud-specific, not CRM-native.
  • Confirm whether each dataset should be treated as Profile, Engagement, or Other before creating the stream.
  • Distinguish stream-level refresh from connection-level reruns when working with unstructured sources.
  • Use UI setup intentionally when initial stream or unstructured asset creation is platform-gated.
  • Hand off to Harmonize only after ingestion assets are clearly healthy.

---

Recommended Workflow

1. Classify readiness for prepare work

node ~/.claude/skills/sf-datacloud/scripts/diagnose-org.mjs -o <org> --phase prepare --json

2. Inspect existing ingestion assets

sf data360 data-stream list -o <org> 2>/dev/null
sf data360 dlo list -o <org> 2>/dev/null

3. Confirm the stream category before creation

Use these rules when suggesting categories:

CategoryUse forTypical requirement
Profileperson/entity recordsprimary key
Engagementtime-based events or interactionsprimary key + event time field
Otherreference/configuration/supporting datasetsprimary key

When the source is ambiguous, ask the user explicitly whether the dataset should be treated as Profile, Engagement, or Other.

4. Create or inspect streams intentionally

sf data360 data-stream get -o <org> --name <stream> 2>/dev/null
sf data360 data-stream create-from-object -o <org> --object Contact --connection SalesforceDotCom_Home 2>/dev/null
sf data360 data-stream create -o <org> -f stream.json 2>/dev/null
sf data360 data-stream run -o <org> --name <stream> 2>/dev/null

5. Check DLO shape

sf data360 dlo get -o <org> --name Contact_Home__dll 2>/dev/null

6. Choose the right refresh mechanism

Use the smaller refresh scope that matches the user goal:

sf data360 data-stream run -o <org> --name <stream> 2>/dev/null
sf data360 connection run-existing -o <org> --name <connection-id> 2>/dev/null
  • data-stream run is the closest match to a stream-level refresh or re-scan.
  • connection run-existing runs at the connection level and can be useful for some connector workflows, but it is not a reliable replacement for stream refresh on unstructured sources.
  • For unstructured document connectors, prefer data-stream run when the goal is to re-scan newly added or changed files.

7. Handle unstructured sources deliberately

For SharePoint-style document ingestion, a minimal unstructured DLO payload can look like:

{
  "name": "my_udlo",
  "label": "My UDLO",
  "category": "Directory_Table",
  "dataSource": {
    "sourceType": "SF_DRIVE",
    "directoryAndFilesDetails": [
      {
        "dirName": "SPUnstructuredDocument/<CONNECTION_ID>/<SITE_ID>",
        "fileName": "*"
      }
    ],
    "sourceConfig": {
      "reservedPrefix": "$dcf_content$"
    }
  }
}

Use the UI for the first-time unstructured setup when the user needs the richer end-to-end pipeline. The UI path can seed additional document metadata fields and downstream assets that a bare CLI DLO create flow may not provision automatically.

8. Use the local Ingestion API example for send-data workflows

For external systems pushing records into Data Cloud:

1. create the connector in sf-datacloud-connect 2. upload the schema with sf data360 connection schema-upsert 3. create the stream in the UI when required 4. send records with the local example in examples/ingestion-api/

cd examples/ingestion-api
cp .env.example .env
python3 send-data.py

Key details:

  • auth is a staged flow: JWT → Salesforce token → Data Cloud token
  • the ingestion endpoint uses the tenant URL, not the Salesforce instance URL
  • 202 means the payload was accepted for processing, not that records are queryable immediately
  • validation failures often surface in the Problem Records DLO family

9. Only then move into harmonization

Once the stream and DLO are healthy, hand off to sf-datacloud-harmonize.

---

High-Signal Gotchas

  • CRM-backed stream behavior is not the same as fully custom connector-framework ingestion.
  • sf data360 data-stream run and sf data360 connection run-existing are not interchangeable; prefer stream-level refresh for unstructured rescans.
  • SFDC streams sync on a platform-managed schedule; data-stream run is not the general control path for CRM connector refresh.
  • Some external database connectors can be created via API while stream creation still requires UI flow or org-specific browser automation. Do not promise a pure CLI stream-creation path for every connector type.
  • Initial SharePoint-style unstructured setup can be richer in the UI than in a minimal CLI DLO create flow.
  • Stream deletion can also delete the associated DLO unless the delete mode says otherwise.
  • DLO field naming differs from CRM field naming, including __c_c transformations.
  • Query DLO record counts with Data Cloud SQL instead of assuming list output is sufficient.
  • CdpDataStreams means the stream module is gated for the current org/user; guide the user to provisioning/permissions review instead of retrying blindly.

---

Output Format

Prepare task: <stream / dlo / transform / docai>
Source: <connection + object>
Target org: <alias>
Artifacts: <stream names / dlo names / json definitions>
Verification: <passed / partial / blocked>
Next step: <harmonize or retrieve>

---

References

  • README.md
  • examples/ingestion-api/README.md
  • ../sf-datacloud/assets/definitions/data-stream.template.json
  • ../sf-datacloud/references/plugin-setup.md
  • ../sf-datacloud/references/feature-readiness.md

Related skills

How it compares

Use sf-datacloud-prepare for programmatic Ingestion API loads; use sf-datacloud-connect when wiring Heroku Postgres through ingress connectors.

FAQ

What authentication does sf-datacloud-prepare use?

sf-datacloud-prepare configures OAuth with a consumer key and consumer secret from a Salesforce External Client App, plus username, login URL, tenant URL, and a private key file for Ingestion API access.

What does sf-datacloud-prepare output?

sf-datacloud-prepare produces environment configuration and Python Ingestion API examples that send structured records to a named Data Cloud object such as a connector-specific entity like Badge_Scan.

Is Sf Datacloud Prepare safe to install?

skills.sh reports 2 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

Backend & APIsautomationagents

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.