
Sentiment Forecasting Engineer
- 20 installs
- 7 repo stars
- Updated May 20, 2026
- daemon-blockint-tech/agentic-enteprises-skill
Forecast aggregate sentiment over time: sentiment indices from text streams, temporal rollups, time-series/sequence models (ARIMA, Prophet, state-space), and walk-forward backtests.
About
Guides forecasting of aggregate sentiment and opinion dynamics over time covering sentiment indices from text streams, temporal rollups, leading/lagging KPI links, time-series and sequence models, nowcasting, and walk-forward backtests. A developer uses it when building sentiment indices or forecasting opinion trajectories.
- Selects ARIMA, Prophet, state-space, or TFT models with prediction intervals
- Backtests with walk-forward validation and handles spikes, bot noise, and regime shifts
Sentiment Forecasting Engineer by the numbers
- 20 all-time installs (skills.sh)
- Ranked #1,268 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Data as of Jul 27, 2026 (Skillselion catalog sync)
npx skills add https://github.com/daemon-blockint-tech/agentic-enteprises-skill --skill sentiment-forecasting-engineerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 20 |
|---|---|
| repo stars | ★ 7 |
| Last updated | May 20, 2026 |
| Repository | daemon-blockint-tech/agentic-enteprises-skill ↗ |
What it does
Forecast aggregate sentiment over time: sentiment indices from text streams, temporal rollups, time-series/sequence models (ARIMA, Prophet, state-space), and walk-forward backtests.
Files
Sentiment Forecasting Engineer
When to Use
- Build aggregate sentiment indices from high-volume text streams (social, news, reviews, surveys)
- Design temporal rollups — hourly, daily, weekly aggregation with consistent weighting rules
- Forecast opinion trajectories — point forecasts, prediction intervals, and scenario bands
- Model leading/lagging relationships between sentiment and sales, traffic, volatility, or brand KPIs
- Select and implement time-series and sequence models — ARIMA, Prophet, state-space, TFT, etc.
- Run nowcasts and choose forecast horizons aligned to decision cadence
- Engineer features from volume, velocity, topic mix, and engagement quality
- Backtest with walk-forward validation and report calibration of uncertainty
- Handle spikes, bot noise, sample bias, and regime shifts in language or product mix
- Integrate outputs with BI dashboards, brand monitoring, or research workflows (methodology only)
When NOT to Use
- Per-document or per-span polarity labeling, annotation, or classifier training →
sentiment-analysis-engineer - Generic demand, inventory, or logistics forecasting without sentiment inputs →
predictive-logistics-developer,data-scientist - Investment advice, trade recommendations, or actionable trading signals → provide forecasting methodology and uncertainty only
- Marketing copy, campaigns, or brand voice →
content-creator,brand-voice-enforcement - Broad macro econometrics or financial modeling without text-derived sentiment →
financial-analyst(partial overlap only) - Exploratory NLP or single-shot sentiment scores on a static corpus →
sentiment-analysis-engineer - LLM product features, agents, or RAG (unless sentiment forecasting is one pipeline component) →
ai-engineer
Related skills
| Need | Skill |
|---|---|
| Document-level polarity, ABSA, annotation, classifier eval | sentiment-analysis-engineer |
| General ML, experimentation, non-time-series predictive modeling | data-scientist |
| Warehouse metrics, dbt, analytics pipelines (if present in repo) | analytics-engineer |
| Demand/inventory forecasting without opinion indices | predictive-logistics-developer |
| Campaign ROI and channel performance (if present in repo) | marketing-analyst |
| Ratios, valuation, macro series without text sentiment (if present) | financial-analyst |
| LLM apps, feature stores for agent products | ai-engineer |
Core Workflows
1. Scope and index design
Clarify population (brand, product, geo), text sources, aggregation grain, target horizon, and downstream KPIs.
See `references/sentiment_forecasting_engineer_scope.md`.
2. Indices, aggregation, and features
Define index formulas, rollups, topic/strata splits, and covariates (volume, velocity, mix).
See `references/indices_aggregation_and_features.md`.
3. Time-series and forecast models
Choose baselines and advanced models; align seasonality, holidays, and exogenous drivers.
See `references/time_series_and_forecast_models.md`.
4. Backtesting, validation, and metrics
Walk-forward evaluation, interval calibration, and spike-event holdouts.
See `references/backtesting_validation_and_metrics.md`.
5. Data quality, bias, and events
Bot filtering, sample bias, language drift, and shock labeling for scenario analysis.
See `references/references_data_quality_bias_and_events.md`.
6. Production monitoring and stakeholders
Serving cadence, drift monitors, dashboard contracts, and stakeholder-ready narratives.
See `references/production_monitoring_and_stakeholders.md`.
Outputs
- Index specification — formula, universe, weights, strata, and revision policy
- Feature catalog — engineered signals with definitions and lag structure
- Forecast spec — horizon, frequency, model family, and exogenous inputs
- Backtest report — walk-forward metrics, interval coverage, and failure slices
- Nowcast playbook — latency budget, refresh rules, and stale-data handling
- Monitoring plan — drift, spike alerts, and human review triggers
- Stakeholder brief — trajectory narrative with explicit uncertainty (no trade advice)
Principles
- Forecast aggregates, not individual opinions — index stability and definitional clarity come first
- Treat index construction as part of the model — changing weights invalidates historical comparability
- Prefer walk-forward evaluation over single holdout splits for time-ordered data
- Report intervals and scenarios, not point estimates alone; disclose coverage on backtests
- Separate methodology from decisions — do not present forecasts as buy/sell or guaranteed outcomes
- Document known biases (platform mix, bot share, demographic skew) beside every published index
Backtesting, validation, and metrics
Table of contents
1. Walk-forward protocol 2. Point forecast metrics 3. Interval and scenario metrics 4. Event and spike holdouts 5. Reporting template
Walk-forward protocol
Timeline: |---- train ----|-- test --|-- train' --|-- test' --| ...
origin expands or rolls forward each step| Parameter | Guidance |
|---|---|
| Origin | First date with stable index definition |
| Step | Match production retrain cadence (e.g., weekly) |
| Test window | ≥ one full season for seasonal series |
| Embargo | Gap between train end and test start if labels lag |
| Retrain | Freeze hyperparameters per protocol; document tuning set |
Forbidden: random train/test splits on time series; tuning on the full history; peeking at future spikes when engineering features.
Point forecast metrics
| Metric | Use | Caveat |
|---|---|---|
| MAE / RMSE | Level accuracy | Scale-dependent across indices |
| MAPE | Stakeholder-friendly | Undefined near zero; biased low volumes |
| sMAPE | Symmetric % errors | Still unstable at zero |
| MASE | Compare to naive | Good for model selection |
| Directional hit rate | Trend sign | Ignores magnitude |
Report metrics per horizon and per stratum (channel, region). Aggregate with volume weights when combining strata.
Interval and scenario metrics
| Metric | Definition | Target |
|---|---|---|
| Coverage | % actuals inside PI | Near nominal (e.g., 80% for 80% PI) |
| Width | Mean interval width | Minimize subject to coverage |
| CRPS | Probabilistic score | Lower is better |
| Calibration plot | Nominal vs empirical | Diagnose under/over-confidence |
Scenarios: stress bands (optimistic/base/pessimistic) must map to documented shocks (volume −50%, bot spike, classifier drift), not arbitrary ±2σ.
Event and spike holdouts
Reserve labeled crisis windows (product recall, viral backlash, outage) for out-of-sample spike tests:
- Did intervals widen before the peak?
- Did nowcast lag exceed SLA?
- Did bot-filter failure inflate the false spike?
Do not remove spikes from training without documenting that production forecasts must handle spikes too.
Reporting template
1. Data window and index version 2. Walk-forward diagram and fold count 3. Baseline vs challenger metrics table (by horizon) 4. Interval coverage and calibration figure 5. Worst misses — dates, drivers, data gaps 6. Sensitivity — weighting, neutral policy, bot filter on/off 7. Production readiness — latency, refresh, monitoring hooks
Indices, aggregation, and features
Table of contents
1. Index families 2. Aggregation grains 3. Weighting and stratification 4. Feature engineering 5. Revision and comparability
Index families
| Index type | Definition | Typical use |
|---|---|---|
| Net polarity | (positive − negative) / volume or weighted mean of signed scores | Headline brand trajectory |
| Positive share | positive / (positive + negative + neutral policy) | Simple stakeholder dashboards |
| Intensity-weighted mean | mean score × engagement weight | Social where likes/shares matter |
| Aspect/strata indices | Same formulas per topic, product, geo | Diagnostic decomposition |
| Dispersion | variance or entropy of scores within window | Polarization early warning |
Document neutral handling explicitly (exclude, separate series, or map to 0). Changing neutral policy breaks historical continuity.
Aggregation grains
| Grain | Pros | Cons |
|---|---|---|
| Hourly | Fast nowcasts, event detection | Noisy; needs stronger smoothing |
| Daily | Standard for brand and markets | Misses intraday shocks |
| Weekly | Stable for executive reporting | Lags fast-moving crises |
Rollup rules:
- Use consistent timezone (UTC vs market local) and document market-hours filters for finance use cases
- For partial periods (today), emit provisional flags until window closes
- Apply minimum volume thresholds; suppress or widen intervals when N is low
Weighting and stratification
| Scheme | When to use | Pitfall |
|---|---|---|
| Equal post | Democratic voice | Dominated by spam/bots |
| Engagement-weighted | Influence-aware | Celebrities swamp signal |
| Author-deduped | Reduce bot farms | Needs stable author IDs |
| Sample-weighted | Correct platform mix | Requires population estimates |
Stratify by language, channel, product, and region when mix shifts drive headline moves more than opinion shifts.
Feature engineering
Build forecast-ready covariates alongside the target index:
| Feature | Construction | Notes |
|---|---|---|
| Volume | post count, unique authors | Log-transform; watch outages |
| Velocity | Δvolume, acceleration | Spike precursor |
| Topic mix | share of top-K topics | Drift when product launches |
| Engagement rate | interactions per post | Separates “loud” from “many” |
| New-author share | first-time posters / total | Bot/inorganic indicator |
| Score dispersion | std dev within window | Polarization |
| Exogenous KPIs | sales, search, prices | Align lags in modeling stage |
Lag structure: pre-specify max lag (e.g., 0–14 days) for outcome linkage studies; avoid fishing every lag without multiple-testing control.
Revision and comparability
- Version indices when classifier, lexicon, or weighting changes
- Publish overlap period where v1 and v2 run in parallel
- Never restate backtests using formulas not available at historical decision times
- Store raw inputs (hashed IDs, counts) to reproduce aggregates without re-scraping
Production monitoring and stakeholders
Table of contents
1. Serving architecture 2. SLAs and refresh cadence 3. Drift and quality monitors 4. Dashboard and API contracts 5. Stakeholder communications 6. Governance
Serving architecture
Typical pipeline:
Ingest → score (batch/stream) → aggregate index → feature store → forecast job → publish| Mode | Latency | Notes |
|---|---|---|
| Batch daily | Hours acceptable | Reconcile partial days next run |
| Intraday nowcast | Minutes–hours | Strong staleness alerts |
| On-demand API | Seconds | Cache latest forecast + metadata |
Version every publish with index_version, model_id, and data_through timestamp.
SLAs and refresh cadence
Document:
- Data through time shown on every chart
- Maximum acceptable ingest lag
- Retrain schedule vs forecast-only refresh
- Fallback when upstream classifier unavailable (last good scores vs halt)
Drift and quality monitors
| Monitor | Threshold idea | Response |
|---|---|---|
| Volume Z-score | \ | z\ |
| Score distribution PSI | > 0.2 | Audit classifier; consider retrain |
| Bot rate | Week-over-week +10pp | Tighten filters |
| Coverage | Authors or posts −20% | Widen intervals; flag |
| Forecast error | Rolling MAE +30% | Model review |
| Interval coverage live | Below nominal −10pp | Recalibrate uncertainty |
Run human audit samples weekly on high-visibility brands.
Dashboard and API contracts
Minimum fields for consumers:
{
"index_id": "brand_x_daily_net",
"as_of": "2026-05-20T14:00:00Z",
"data_through": "2026-05-20T12:00:00Z",
"point": 0.12,
"pi80": [0.05, 0.19],
"horizon_days": 7,
"provisional": false,
"index_version": "v3",
"model_id": "prophet_v2_20260501"
}Charts should show history, forecast, intervals, and annotated events. Provide strata drill-down when headline moves are mix-driven.
Stakeholder communications
| Audience | Emphasize | Avoid |
|---|---|---|
| Brand / CX | Trajectory, drivers, uncertainty | False precision |
| Product | Topic strata, launch windows | Single-number hype |
| Research / markets | Methodology, lags, backtest | Trade instructions |
| Executives | Scenarios, risks, data limits | Causal claims without evidence |
Use plain language: “directionally positive with wide uncertainty” when coverage is poor.
Governance
- No automated trading from this skill’s outputs without separate compliance review
- Log who approved index formula changes
- Retain reproducibility artifacts (configs, hashes) per retention policy
- Escalate when forecasts could affect public communications or regulated disclosures
Partner with sentiment-analysis-engineer for classifier changes; with data-scientist for generic experiment design when A/B testing forecast-informed decisions.
Data quality, bias, and events
Table of contents
1. Coverage and completeness 2. Bot and inorganic noise 3. Sample and platform bias 4. Language and product drift 5. Shock and event taxonomy 6. Mitigation playbook
Coverage and completeness
| Check | Signal | Action |
|---|---|---|
| Ingest lag | Timestamp gap vs wall clock | Flag nowcast staleness |
| API quota drops | Volume cliff | Pause forecasts; widen intervals |
| Source outage | Single channel → 0 | Impute vs suppress stratum |
| Dedup failures | Inflated volume | Fix pipeline before re-forecast |
Track effective sample size per window; suppress headline index when below threshold.
Bot and inorganic noise
Detection signals (combine, do not rely on one):
- Repetitive text / template hashes
- Young accounts with burst posting
- Coordinated timing clusters
- Engagement mismatch (views ≫ plausible reach)
Policy choices:
- Hard drop suspected bots from index
- Down-weight by bot probability
- Separate organic vs all-in series for transparency
Document bot rules; changing rules requires index version bump.
Sample and platform bias
| Bias type | Manifestation | Mitigation |
|---|---|---|
| Platform mix | Twitter vs Reddit shift | Mix-adjust or fixed basket |
| Demographic skew | Age/geo not representative | Weight to census or customer base |
| Selection bias | Only angry customers review | Pair with surveys; strata |
| Survivorship | Deleted posts | Archive at ingest |
| Amplification | Press picks up social | Tag earned media separately |
State who is not heard in the index footnote for stakeholder briefs.
Language and product drift
Monitor:
- Language share shifts (new locale launch)
- Topic model drift (new product names)
- Classifier score distribution shift (prior calibration)
Triggers for retrain or remap:
- Population stability index on scores
- Sudden neutral→positive ratio jump without volume story
- Human audit sample failure rate ↑
Shock and event taxonomy
| Class | Examples | Forecast behavior |
|---|---|---|
| Exogenous shock | Recall, CEO scandal, outage | Widen bands; scenario mode |
| Endogenous spike | Viral campaign (planned) | Annotate; optional exclude from train |
| Data artifact | Crawler bug, double count | Halt publish; fix upstream |
| Regime change | Rebrand, M&A, platform ban | New index version |
Maintain an event register with start/end, affected strata, and whether historical restatement is required.
Mitigation playbook
1. Detect — automated QA + human spot checks on volume/velocity anomalies 2. Classify — artifact vs organic vs planned 3. Respond — suppress, annotate, or scenario-forecast; never hide uncertainty 4. Review — post-mortem on bot filter and index formula 5. Version — document changes for reproducibility
Sentiment forecasting engineer scope
Table of contents
1. Role boundary 2. In-scope deliverables 3. Out of scope 4. Engagement checklist 5. Handoff map
Role boundary
Own aggregate opinion dynamics over time—how sentiment indices move, correlate with outcomes, and forecast under uncertainty.
| Own | Partner skill |
|---|---|
| Index design, rollups, and revision policy | sentiment-analysis-engineer — per-text labeling and classifiers |
| Time-series and sequence forecasting | data-scientist — general predictive modeling without opinion indices |
| Demand/inventory forecasts without sentiment | predictive-logistics-developer |
| LLM product surfaces beyond forecasting pipelines | ai-engineer |
| Marketing copy and campaign creative | content-creator |
| Trade recommendations or investment advice | Methodology only; no actionable signals |
Unit of analysis: time-indexed population aggregates (brand, product line, geo, channel), not individual posts unless explicitly modeling distributions.
In-scope deliverables
| Artifact | Contents |
|---|---|
| Index specification | Universe, sampling, polarity→score mapping, weights, strata |
| Aggregation playbook | Grain (H/D/W), missing-data rules, revision and restatement policy |
| Feature catalog | Volume, velocity, topic mix, engagement quality, exogenous KPIs |
| Forecast design | Horizon, frequency, model ladder, exogenous structure |
| Backtest report | Walk-forward metrics, interval coverage, spike holdouts |
| Nowcast SOP | Latency budget, partial-period handling, stale-source fallbacks |
| Monitoring spec | Drift, bot share, coverage, alert thresholds |
| Stakeholder brief | Trajectory + uncertainty; explicit limitations |
Out of scope
- Training or evaluating document-level sentiment classifiers (see
sentiment-analysis-engineer) - Buy/sell, position sizing, or guaranteed alpha claims from sentiment forecasts
- Marketing messaging or creative based on forecast direction
- Legal/compliance determinations from sentiment alone
- Pure macro econometric models with no text-derived sentiment component
Engagement checklist
- [ ] Business decision named (brand health, product launch, risk desk research, ops staffing)
- [ ] Target population and text sources listed with access and retention constraints
- [ ] Index formula frozen (or versioned) before backtest claims
- [ ] Forecast horizon and decision cadence aligned (e.g., daily index → 7/14/28-day horizons)
- [ ] Outcome variables for lead/lag study defined (sales, NPS proxy, volatility, search)
- [ ] Bot/spam and sample-bias mitigation documented
- [ ] Walk-forward protocol pre-registered (origin, step, embargo if needed)
- [ ] Interval/scenario reporting required for all external forecasts
- [ ] Dashboard or API consumer identified; SLA for refresh documented
Handoff map
Text streams → Score/label (partner: sentiment-analysis-engineer)
→ Aggregate index (this skill)
→ Features + forecast (this skill)
→ BI / research dashboard (consumer)
→ Monitor drift + restate (this skill)When upstream classifiers change, rebuild or re-link historical indices with a version bump; do not silently splice incompatible scores.
Time-series and forecast models
Table of contents
1. Model ladder 2. Classical time-series 3. State-space and Bayesian 4. ML sequence models 5. Exogenous and multivariate 6. Horizon and nowcasting
Model ladder
Always fit a simple baseline before complex models:
1. Seasonal naive / last-value 2. Moving average with seasonality 3. ARIMA / SARIMA or ETS 4. Prophet or structural time-series (holidays, changepoints) 5. State-space (DLM, Kalman) with exogenous regressors 6. ML sequence (gradient boosting on lags, TFT, N-BEATS, etc.)
Report lift over baseline on walk-forward metrics, not in-sample R² alone.
Classical time-series
| Method | Strengths | Watchouts |
|---|---|---|
| ARIMA/SARIMA | Interpretable; good short horizons | Stationarity; regime breaks |
| ETS | Trend/season components | Sparse or zero-inflated volumes |
| Prophet | Holidays, changepoints | Can overfit changepoints on short history |
Seasonality: align to business cycle (weekly for social, monthly for surveys). Include known events (launches, earnings, crises) as regressors or dummy spikes.
State-space and Bayesian
Use when you need online updates or explicit uncertainty:
- Dynamic linear models for slowly drifting level
- Kalman filters for nowcasting partial days
- Bayesian structural time-series for sparse exogenous series
Document prior sensitivity if stakeholders interpret credible intervals as guarantees.
ML sequence models
| Approach | When | Risk |
|---|---|---|
| Lag features + GBM | Medium history; heterogeneous features | Leakage if features use future data |
| TFT / deep seq | Rich covariates; long history | Overfit; opaque |
| Direct multi-horizon | Fixed horizon grid | Needs retrain per horizon change |
Enforce causal feature windows: features at time t may only use information ≤ t.
Exogenous and multivariate
Joint modeling patterns:
- Sentiment → outcome (leading indicator studies): Granger-style tests with pre-specified lags
- Outcome → sentiment (feedback): avoid claiming lead without out-of-sample tests
- Vector models (VAR): when both series forecast together; check stability
Include control series (ad spend, pricing, seasonality) to reduce spurious correlation.
Horizon and nowcasting
| Mode | Definition | Implementation |
|---|---|---|
| Nowcast | Estimate current incomplete period | Intraday partial aggregates + DLM |
| Short horizon | 1–7 steps | ARIMA, GBM lags |
| Medium horizon | 2–12 weeks | Seasonal models + exogenous plans |
| Long horizon | Months+ | Scenario bands; widen intervals |
Match horizon to decision latency; do not report 90-day precision intervals for a daily trading desk unless backtest coverage supports it.