
Product Management Human Data Platform
- 28 installs
- 7 repo stars
- Updated May 20, 2026
- daemon-blockint-tech/agentic-enteprises-skill
Guides product management for human data platforms: annotation/labeling products, workforce workflows, task design, quality systems, and privacy-safe training-data handling.
About
Guides product management for human data platforms covering annotation and labeling products, workforce workflows, task and quality design, customer ML-team delivery, and privacy-safe data handling. A PM uses it when prioritizing roadmap for labeling/RLHF/eval platforms or writing annotation-feature PRDs.
- Specifies quality programs: gold tasks, consensus, adjudication, and IAA
- Sets metrics for throughput, quality, cost per label, and contributor retention
Product Management Human Data Platform by the numbers
- 28 all-time installs (skills.sh)
- Ranked #9,505 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Jul 29, 2026 (Skillselion catalog sync)
npx skills add https://github.com/daemon-blockint-tech/agentic-enteprises-skill --skill product-management-human-data-platformAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 28 |
|---|---|
| repo stars | ★ 7 |
| Last updated | May 20, 2026 |
| Repository | daemon-blockint-tech/agentic-enteprises-skill ↗ |
What it does
Guides product management for human data platforms: annotation/labeling products, workforce workflows, task design, quality systems, and privacy-safe training-data handling.
Files
Product Management — Human Data Platform
When to Use
- Define vision, roadmap, and prioritization for labeling, RLHF, or human-eval products
- Write PRDs for annotation UI, project setup, QA, workforce, or export/API features
- Design annotation tasks (taxonomy, instructions, rubrics, edge cases)
- Specify quality programs: gold tasks, consensus, adjudication, rejection reasons
- Scope customer workflows (ML teams): projects, batches, SLAs, delivery formats
- Improve contributor/annotator productivity, fairness, and trust/safety product surfaces
- Set metrics: throughput, quality, cost per label, time-to-delivery, contributor retention
- Partner on privacy and ethics requirements for human-submitted data (PII, consent, locale)
When NOT to Use
- Facilitate generic process maps and BRDs without product ownership →
business-analyst - Wireframes and visual design only →
product-designer - RAG/copilot enterprise architecture →
applied-ai-architect-commercial-enterprise - Build eval harnesses and judges in code →
prompt-engineer-agent-prompts-evals - SOC/ISO evidence automation →
compliance-engineer - Data warehouse modeling →
data-warehouse-engineer - Cross-team delivery RAID without product discovery →
technical-program-manager
Related skills
| Need | Skill |
|---|---|
| BRD/user story format | business-analyst |
| Annotator and customer UX | product-designer |
| How labels feed model programs | applied-ai-architect-commercial-enterprise |
| Golden sets and regression evals | prompt-engineer-agent-prompts-evals |
| Privacy controls and audit evidence | compliance-engineer |
| Taxonomy/ontology for labels | ontology-engineer |
| Analytics for product teams | analytics-data-engineering-manager-product |
Core Workflows
1. Vision, roadmap, and prioritization
Outcomes, segments, themes, RICE/ICE.
See `references/roadmap_prioritization.md`.
2. Annotation task and taxonomy design
Instructions, rubrics, schema, edge cases.
See `references/annotation_task_design.md`.
3. Quality systems
Gold sets, IAA, adjudication, rejection taxonomy.
See `references/quality_systems.md`.
4. Customer (ML team) delivery
Projects, pipelines, exports, SLAs.
See `references/customer_ml_workflows.md`.
5. Contributor and workforce product
Task UX, payments, trust, locale.
See `references/contributor_workforce_product.md`.
6. Privacy, ethics, and policy
PII, consent, retention, labor.
See `references/privacy_ethics_policy.md`.
Output standards
- PRDs state persona, problem, success metrics, non-goals, and launch tier
- Task specs include worked examples (gold, borderline, reject)
- Quality bar defined as measurable thresholds, not "high quality"
- Every feature maps to cost, quality, or speed lever
- Escalate legal/labor questions; do not ship policy in product copy alone
When to load references
- Roadmap →
references/roadmap_prioritization.md - Tasks →
references/annotation_task_design.md - Quality →
references/quality_systems.md - Customers →
references/customer_ml_workflows.md - Contributors →
references/contributor_workforce_product.md - Privacy →
references/privacy_ethics_policy.md
Annotation Task Design
Task design deliverables
| Artifact | Purpose |
|---|---|
| Taxonomy | Labels, hierarchies, mutually exclusive rules |
| Instructions | Step-by-step for contributors |
| Rubric | Borderline examples and decision rules |
| UI spec | Controls, required fields, media handling |
| Edge-case catalog | Ambiguous items with gold answer |
| Calibration set | Items for onboarding and drift checks |
Instruction quality checklist
- [ ] Objective defined in one sentence
- [ ] Definitions for every label with positive and negative examples
- [ ] Order of operations (e.g. "check safety before sentiment")
- [ ] What to do when unsure (skip, flag, best guess policy)
- [ ] Media-specific rules (audio clipping, occlusion, PII blur)
- [ ] Change log when taxonomy updates
Taxonomy rules
| Rule type | Example |
|---|---|
| Mutually exclusive | Single primary intent per utterance |
| Hierarchical | Vehicle → car → sedan |
| Multi-label | Tags where independent |
| Conditional | If "toxic" then severity required |
Pair with ontology-engineer when enterprise customers need shared ontologies across programs.
Modality considerations
| Modality | PM must specify |
|---|---|
| Text | Language, encoding, span vs document level |
| Image | Bounding box vs polygon vs keypoint |
| Audio | Transcription vs diarization vs emotion |
| Video | Frame sampling rate, temporal segments |
| Multi-modal | Which field is source of truth |
Pre-labeling and automation
| Strategy | When |
|---|---|
| Model pre-label + human edit | High volume, stable taxonomy |
| Active learning queue | Edge cases for next batch |
| Rules engine | Hard constraints (regex, blocklists) |
Document expected edit rate so ops can staff reviewers.
Versioning
- Task spec version tied to exported dataset metadata
- Breaking taxonomy changes require relabel or migrate plan
- Customer notification for downstream training impact
Review with design and eng
product-designer: task UI affordances, error prevention- Eng: latency, autosave, conflict on concurrent edit
- Ops: estimated handle time per item for pricing/SLA
Contributor and Workforce Product
Contributor jobs to be done
- Understand task quickly
- Complete work accurately with fair pay
- Know why work was rejected and how to improve
- Trust platform on payments, safety, and privacy
Core product surfaces
| Surface | PM focus |
|---|---|
| Task feed | Matching skills, fair distribution, no starvation |
| Task UI | Clarity, keyboard shortcuts, accessibility |
| Training | Calibration modules tied to tier promotion |
| Feedback | Reject reasons, examples, appeal flow |
| Earnings | Transparent rates, payout status, tax docs (locale) |
| Trust & safety | Report abuse, content warnings, wellbeing breaks |
Fairness and gaming resistance
| Risk | Product mitigation |
|---|---|
| Speed over quality | Minimum time thresholds + gold monitoring |
| Gold memorization | Rotate gold, per-user seeds |
| Collusion | Random assignment, anomaly detection alerts |
| Bias in assignment | Skill-based routing, audit on distribution |
Workforce ops product hooks
- Capacity forecasts by locale and skill
- Surge pricing or bonus campaigns (policy-reviewed)
- Contributor blocking and reinstatement workflows
- Language and timezone coverage maps
Localization
- Instructions and UI in contributor language
- Locale-specific payment and labor disclosures
- Cultural nuance in rubrics (not literal translation only)
Metrics
| Metric | Signal |
|---|---|
| Task abandonment | Instruction or UX friction |
| Active contributors / week | Supply health |
| Avg handle time | Pricing and SLA feasibility |
| Appeal overturn rate | Rubric or reviewer quality |
| Contributor churn | Trust, pay, or task boredom |
RLHF / preference-specific UX
- Side-by-side comparison clarity
- Tie and "both bad" handling
- Rationale capture optional vs required (trade speed vs quality)
- Leakage prevention (hide model identity if blind eval required)
Ethics
- Avoid deceptive task framing about purpose of data use where regulated
- Limit exposure to traumatic content; rotation and wellness features
- Clear opt-out for sensitive categories
Legal review for labor classification and regional workforce law—not PM alone.
Customer ML Team Workflows
End-to-end journey
1. Scope — modality, volume, deadline, quality bar 2. Configure — project, taxonomy, workforce tier, instructions 3. Pilot — small batch, calibrate IAA and handle time 4. Scale — production batches, monitoring, SLA tracking 5. Accept — QA sign-off, export, lineage metadata 6. Iterate — taxonomy v2, active learning loops
Project setup (product requirements)
| Field | Why it matters |
|---|---|
| Delivery date | Drives staffing and feature flags |
| Quality tier | Gold %, consensus, adjudication depth |
| Data classification | Privacy features, region, retention |
| Export schema | JSONL, COCO, Parquet, custom |
| Versioning | Model training reproducibility |
Self-serve vs managed
| Mode | Customer | Platform |
|---|---|---|
| Self-serve | Configures tasks, monitors dashboards | Templates, guardrails, billing |
| Managed | Solution engineer / PM runs project | Playbooks, dedicated QA, custom SLA |
Product should not blur modes without clear pricing.
API and integration
- Batch create/upload, status webhooks, export download
- Idempotent job IDs for pipeline orchestration
- Rate limits and pagination documented
- Sandbox project for integration testing
Dashboards (customer)
- Progress: completed / in review / rejected
- Quality: gold accuracy, IAA, rework trend
- Cost: spend vs estimate (if usage-based)
- Blockers: queue starvation, instruction tickets
Delivery acceptance
## Acceptance checklist
- [ ] Volume and format match SOW
- [ ] Quality metrics meet appendix thresholds
- [ ] Label version and task spec hash in metadata
- [ ] Known limitations documented (classes with low support)
- [ ] PII attestation if applicableExpansion and retention
- Track time-to-first-export for new accounts
- Instrument repeat project rate by vertical
- Capture quality incident root cause for product backlog
Handoffs
- Technical integration issues → support/engineering paths in org
- Legal/DPA →
commercial-counsel+compliance-engineer - Custom ontology →
ontology-engineer
Privacy, Ethics, and Policy
Human data risks
| Risk | Product response |
|---|---|
| PII in source media | Detection, blur, block, escalate |
| Sensitive attributes | Restricted tasks, expert tier only |
| Re-identification | Aggregation limits on exports |
| Consent gaps | Customer attestation fields, block upload |
| Cross-tenant leakage | Strict project isolation, audit logs |
| Misuse of labels | Acceptable use policy, export controls |
Privacy-by-design features
- Role-based access (customer, contributor, reviewer, admin)
- Region pinning for storage and processing
- Retention TTL per project with legal hold override
- Download logging and watermarking on exports
- Contributor view: minimum data needed for task (masking)
Customer responsibilities (product enforces)
Before production labeling:
- [ ] Customer confirms rights to use data for labeling
- [ ] Classification selected (public / confidential / regulated)
- [ ] Retention and deletion requirements captured
- [ ] Banned content categories acknowledgedEthics review triggers
Escalate product/legal review when:
- Biometric, medical, or children's data
- Deception studies or undisclosed recording
- Political, religious, or violence-heavy content at scale
- Surveillance or law-enforcement sensitive use cases
- Export to jurisdictions with conflicting rules
Alignment with compliance engineering
Map features to controls with compliance-engineer:
- Access reviews for admin roles
- Encryption at rest/transit
- Audit logs for label changes and exports
- DPIA inputs for new modalities or regions
Transparency
- Contributor: what data is used for (within legal bounds)
- Customer: how quality and workforce tiers work
- Document known limitations in release notes
Incident response (product angle)
| Event | PM actions |
|---|---|
| PII spill in task | Pause project, purge policy, notify customer |
| Rubric causes systematic harm | Halt task type, revise instructions |
| Workforce protest / pay dispute | Coordinate with ops and legal |
Pair with community-executive-escalations-program-manager if public reputational escalation.
AI training downstream
- Metadata: label provenance, spec version, workforce tier
- Avoid implying labels are "ground truth" in API docs—document error bars
- Support customer deletion requests propagating to derived exports where feasible
Anti-patterns
- Shipping geo features without residency analysis
- Using contributor-generated content to train platform models without disclosed consent
- Hiding quality problems from customer dashboards
Quality Systems
Quality model
Raw labels → QA gates → Accepted dataset → Customer export
↑
Gold / consensus / adjudicationMechanisms
| Mechanism | Use when |
|---|---|
| Gold tasks | Known-answer items mixed into production |
| Consensus | Multiple annotators; majority or unanimous rule |
| Adjudication | Expert resolves disagreement |
| Reviewer sampling | % audit of accepted work |
| Automated checks | Schema, bounds, regex, model confidence |
Metrics (define targets per project type)
| Metric | Definition |
|---|---|
| Accuracy | % match to gold (on gold items) |
| IAA | Cohen's kappa / Fleiss / Krippendorff as appropriate |
| Rework rate | % items sent back to contributor |
| Throughput | Labels per contributor hour |
| SLA adherence | % batches delivered on time |
| Dispute rate | Appeals per 1k labels |
Report confidence intervals on small samples.
Gold set program
- Size: enough power for weekly monitoring (often 1–5% injection)
- Refresh: rotate gold to prevent memorization
- Stratify: cover rare classes and edge cases
- Never use customer production data as gold without rights
Rejection taxonomy
Standardize reject reasons for product analytics:
- Instruction misunderstanding
- Taxonomy error
- Incomplete annotation
- Tooling bug
- Policy violation (PII, unsafe)
- Borderline / needs adjudicationTiered workforce
| Tier | Access |
|---|---|
| Trainee | Calibration only |
| Standard | Production with gold monitoring |
| Expert | Adjudication, rubric updates |
| Customer reviewer | Final accept on private queues |
Promotion rules must be transparent to contributors.
Customer-facing quality SLAs
Document in contract appendix:
- Minimum accuracy on gold (if customer supplies gold)
- Consensus configuration
- Escalation when IAA drops below threshold
- Rework turnaround time
Product features roadmap signals
- Spike in rework → instruction or UI issue
- IAA collapse on new class → taxonomy or training gap
- Gold accuracy drift → contributor gaming or spec ambiguity
Pair with eval teams
Export formats should support prompt-engineer-agent-prompts-evals golden sets and regression harnesses without manual reformatting.
Roadmap and Prioritization
Platform value levers
| Lever | Product moves |
|---|---|
| Quality | Gold tasks, adjudication, reviewer tiers, auto-QA |
| Speed | Task UX, prefetch, batching, SLA dashboards |
| Cost | Pre-labeling, active learning, tiered workforce |
| Scale | Multi-language, 24/7 queues, API throughput |
| Trust | PII handling, audit trails, customer isolation |
Personas (typical)
| Persona | Jobs to be done |
|---|---|
| ML lead / PM (customer) | Ship model with labeled data on deadline |
| Annotation manager (customer) | Run projects, hit quality bar, report up |
| Contributor | Earn fairly, clear tasks, minimal friction |
| Reviewer / QA | Catch errors fast, consistent rubric |
| Ops / workforce | Staff queues, handle disputes, compliance |
| Platform admin | Tenancy, billing, access, integrations |
Roadmap themes (examples)
- Core labeling — task types, media support, shortcuts
- Quality — consensus, gold injection, analytics
- Delivery — exports, API, versioning, lineage metadata
- Workforce — skills tests, tiers, incentives, appeals
- Enterprise — SSO, VPC, private workforce, DPA features
- GenAI-era — RLHF pairs, rubric ranking, red-team data collection
Prioritization framework (RICE-lite)
| Factor | Question |
|---|---|
| Reach | # projects or labels affected per quarter |
| Impact | Effect on quality, speed, or revenue retention |
| Confidence | Evidence from customers, pilots, metrics |
| Effort | Eng + ops + policy review |
Force-rank one primary lever per epic to avoid "everything is P0."
PRD skeleton
## Problem
## Personas
## Success metrics (baseline → target)
## Scope (in / out)
## User stories
## Task/quality implications
## Dependencies (eng, legal, ops)
## Launch plan (alpha → GA)
## RisksDiscovery inputs
- Customer QBR themes and churn reasons
- Contributor NPS and task-abandon funnels
- Quality regressions (IAA drops, rework rate)
- Competitive gaps (modalities, pricing model)
- Cost per label trend by project type
Anti-patterns
- Roadmap driven only by largest customer's custom ask without platform abstraction
- Features without quality metric definition
- Ignoring contributor experience when optimizing customer dashboards only