
Linkfox Mpstats Ozon Product Trend
- 249 installs
- 64 repo stars
- Updated August 3, 2026
- linkfox-ai/linkfox-skills
Analyze Ozon product trend and sales dynamics via MPStats to benchmark rivals, spot category momentum, and prioritize SKUs for the Russian marketplace.
About
Connects agents to MPStats Ozon analytics for product trend lines, estimated sales velocity, and competitive category context. Useful for cross-border sellers evaluating Russian marketplace entry, pricing bands, and which competitor listings to emulate or avoid.
- Ozon SKU trend curves
- MPStats sales estimates
- Category share shifts
- Rival listing benchmarks
- Russia-focused assortment signals
Linkfox Mpstats Ozon Product Trend by the numbers
- 249 all-time installs (skills.sh)
- +37 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #897 of 1,879 Marketing & SEO skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/linkfox-ai/linkfox-skills --skill linkfox-mpstats-ozon-product-trendAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 249 |
|---|---|
| repo stars | ★ 64 |
| Last updated | August 3, 2026 |
| Repository | linkfox-ai/linkfox-skills ↗ |
What it does
Analyze Ozon product trend and sales dynamics via MPStats to benchmark rivals, spot category momentum, and prioritize SKUs for the Russian marketplace.
Files
MPSTATS Ozon Product Trend (Daily Time-Series)
This skill returns a daily time-series of a single Ozon (Russia) SKU — sales units, price, stock, rating, and optionally search-position / visibility metrics. It is the go-to for validating growth, seasonality, or anomalies for a specific product.
Core Concepts
Single-SKU scope: Each call analyzes exactly one productId. For batch per-SKU snapshots (period aggregates), use mpstats-ozon-product-detail instead.
Daily granularity: The response is an array of daily points (top-level field data) across the [startDate, endDate] window. Each point carries a hasData boolean — if hasData=false, the day has no observation (distinct from sales=0 with hasData=true).
T-1 delay: MPSTATS trend data is delayed by one day; the latest selectable end date is yesterday. Today or future dates are rejected.
Search-visibility add-on: Set includeSearchStats: true to append search-position / visibility signals. Some niches (especially small categories) may not have search-stats coverage — expect partial or empty fields in those cases.
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
| productId | integer | yes | Ozon SKU (numeric) |
| startDate | string | no | Window start, YYYY-MM-DD; latest = yesterday |
| endDate | string | no | Window end, YYYY-MM-DD; latest = yesterday |
| includeFbs | boolean | no | Include FBS data alongside FBO |
| includeSearchStats | boolean | no | Attach search position / visibility signals |
API Usage
This tool calls the LinkFox tool gateway API. See references/api.md for calling conventions, request parameters, response structure, and error codes. You can also execute scripts/mpstats_ozon_product_trend.py directly for ad-hoc queries.
Usage Examples
1. Monthly trend for a SKU
{
"productId": 1786874757,
"startDate": "2025-03-01",
"endDate": "2025-03-31"
}2. Trend with search visibility
{
"productId": 1786874757,
"startDate": "2025-02-01",
"endDate": "2025-02-28",
"includeSearchStats": true
}3. Combined FBO+FBS trend
{
"productId": 151623766,
"startDate": "2025-01-01",
"endDate": "2025-01-31",
"includeFbs": true
}How to Chain with Other Ozon Skills
1. Discovery → trend: Use mpstats-ozon-product-search to find a SKU, then check growth / volatility here before committing. 2. Aggregate vs time-series: mpstats-ozon-product-detail gives a one-number-per-metric period view; this skill shows the day-by-day shape behind those numbers. 3. Drill-down → trend: After brand-products / category-products / seller-products surfaces a hot SKU, use this skill to validate whether the hotness is recent, seasonal, or sustained.
Display Rules
1. Prefer a simple table or sparkline-friendly output — one row per date with date, price, sales, balance, rating, comments; do not overfit a 90-point series into a single paragraph. 2. Use `hasData` to distinguish gaps from zero sales — hasData=false means the day has no observation; don't report it as a zero-sale day. 3. Call out anomalies — large single-day spikes or stockouts (balance=0 runs where hasData=true) should be flagged factually, not as buying advice. 4. Currency is RUB unless upstream layer is already converting (the currency field per point carries the symbol, e.g. ₽); state the currency when showing price movement. 5. Revenue is not returned per day — if the user asks for daily revenue, estimate via sales * price and note it's an estimate. 6. `includeSearchStats` gaps — when no search-visibility fields come back, note "搜索位次数据在该赛道暂不可用" rather than silently omitting. 7. No business advice — present the shape; leave "should we buy this listing?" to the user.
Important Limitations
- Single SKU per call — cannot pass a list of
productIds; loop at the Agent layer if needed. - T-1 data —
endDatecannot be today or a future date. - Search stats optional —
includeSearchStats=truedoesn't guarantee coverage for all niches. - Ozon-only — Wildberries and other Russian marketplaces are not covered.
- Missing days — the series may have nulls / gaps where no data was captured; do not treat nulls as zero sales.
User Expression & Scenario Quick Reference
Applicable — Single-SKU temporal analysis:
| User Says | Scenario |
|---|---|
| "What's the sales trend of Ozon SKU 1786874757 last month" | Monthly time-series |
| "Is this Ozon listing seasonal or stable" | Seasonality check |
| "Did this Ozon product have stockouts recently" | Stock anomaly detection |
| "Price walk for this Ozon product over Q1" | Price movement |
| "Did this listing's search position improve" | Search visibility (requires includeSearchStats) |
Not applicable — Needs beyond single-SKU time-series:
- Batch snapshot of many SKUs →
mpstats-ozon-product-detail - Brand / category / seller drill-down → matching
*-productsskill - Pre-IDed discovery →
mpstats-ozon-product-search
Boundary judgment: Use this skill when the question starts with "how did this ONE product change over time". For multi-SKU comparisons or dimension-level filtering, go elsewhere.
Feedback:
Auto-detect and report feedback via the Feedback API when any of the following apply: 1. The functionality or purpose described in this skill does not match actual behavior 2. The skill's results do not match the user's intent 3. The user expresses dissatisfaction or praise about this skill 4. Anything you believe could be improved
Call the feedback API as specified in references/api.md. Do not interrupt the user's flow.
<!-- LF_LARGE_RESPONSE_BLOCK -->
Handling Large Responses
To avoid overflowing the agent context, persist the response to disk and extract only the fields you need:
python scripts/response_io.py run --script scripts/mpstats_ozon_product_trend.py --out-dir <DIR> '<params>'
python scripts/response_io.py read <file> --fields "<paths>" # or --path "<JMESPath>"Pick--out-diroutside any git working tree (e.g./tmp/...on Unix,%TEMP%/...on Windows). Persisted responses may contain PII, pricing, or auth-sensitive data — do not commit them. Files are not auto-deleted; clean up when the task is done.
run writes the full response to a file and emits only a schema preview + file path. read projects specific fields, with --limit/--offset for slicing and --format json|jsonl|csv|table for output.
When to prefer this pattern — apply your judgment based on the response characteristics, e.g.:
- High field count per record, or fields you don't need
- Batch/paginated results (multiple items per call)
- Long-text fields (descriptions, reviews, HTML, time series)
- Output reused across later steps rather than consumed immediately
For small, single-use responses, calling the main script directly is fine.
⚠️ The preview is a truncated schema + sample, not the full data. Any field-level decision must read from the persisted file via read. <!-- /LF_LARGE_RESPONSE_BLOCK -->
--- For more high-quality, professional cross-border e-commerce skills, set [LinkFox Skills](https://skill.linkfox.com/).
MPSTATS Ozon 商品趋势(分日)API 参考
调用规范
- 请求地址:
https://tool-gateway.linkfox.com/mpstats/ozon/productTrend - 请求方式:POST,Content-Type: application/json
- 认证方式:Header
Authorization: <api_key>,api_key 从环境变量LINKFOXAGENT_API_KEY读取(如未配置,提示用户前往 https://skill.linkfox.com/linkfoxskills/guide.htm 申请)
请求参数
POST Body(JSON)。以下字段与工具网关当前登记的「MPSTATS-Ozon-商品趋势」入参 schema 一致(同步日期 2026-04-30)。
| 参数 | 类型 | 必填 | 说明 |
|---|---|---|---|
| productId | integer | 是 | Ozon 商品 SKU |
| startDate | string | 否 | 统计起始日,YYYY-MM-DD;数据延迟 T-1,最晚可选昨日 |
| endDate | string | 否 | 统计结束日,YYYY-MM-DD;数据延迟 T-1,最晚可选昨日 |
| includeFbs | boolean | 否 | 是否包含 FBS 数据 |
| includeSearchStats | boolean | 否 | 是否附带搜索位次 / 可见性;部分赛道不支持 |
响应结构
| 字段 | 类型 | 说明 |
|---|---|---|
| code | string | 返回码(字符串),"200" 表示成功 |
| errcode | integer | 返回码(整数),200 表示成功 |
| msg / errmsg | string | 消息;成功为 ok |
| total | integer | 分日数据点数量(窗口天数) |
| data | array | 分日数据点列表(详见下方) |
| columns | array | 渲染列定义 |
| costTime | integer | 接口耗时(毫秒) |
| costToken | integer | 消耗 Token 数量 |
| type | string | 响应类型 |
注意:分日序列字段名为 `data`,不是trend;响应体不包含独立的productId回显。
data 数据点字段
按官方 outputSchema 定义(_mpstats_ozon_productTrend,同步日期 2026-05-06)。共 13 个字段:
| 字段 | 类型 | 说明 |
|---|---|---|
| date | string | 日期,YYYY-MM-DD |
| hasData | boolean | 当日是否有数据(false 表示缺失日,区别于 sales=0) |
| price | number | 当日售价 |
| oldPrice | number | 当日折扣前原价 |
| ozonCardPrice | number | Ozon Card 价(Ozon 官方银行卡优惠价) |
| discount | integer | 折扣,百分比整数 0-100 |
| currency | string | 币种符号(如 ₽ / $ / €) |
| sales | integer | 当日销量(件) |
| balance | integer | 当日 FBO 仓库库存(件) |
| rating | number | 评分,取值 0-5 |
| comments | integer | 评论数 |
| isBestseller | boolean | 当日是否带"畅销"标识 |
| isNew | boolean | 当日是否带"新品"标识 |
Schema 未声明的字段不会返回:端点不独立返回revenue、reviewCount、balanceFbs、isInStock等;如需销售额,用sales × price估算。
includeSearchStats 说明
入参层 includeSearchStats=true 仅作为服务端可选能力开关;官方 outputSchema 未声明任何额外顶层数组或 per-point 字段。若未来 schema 扩展,请以工具网关 listEnabledTool 返回的最新 outputSchema 为准。
错误码
| errcode | 含义 | 处理建议 |
|---|---|---|
| 200 | 成功 | 解析 data |
| 401 | 认证失败 | 检查 Authorization API Key |
| 其他非 200 值 | 业务异常 | 查看 errmsg / msg;常见为 productId 无效、日期越过昨日等 |
curl 示例
curl -X POST https://tool-gateway.linkfox.com/mpstats/ozon/productTrend \
-H "Authorization: $LINKFOXAGENT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"productId": 1786874757,
"startDate": "2025-03-01",
"endDate": "2025-03-31",
"includeSearchStats": true
}'---
Feedback API
该接口与上方工具接口不同,请勿混用两个基础 URL。
- POST
https://skill-api.linkfox.com/api/v1/public/feedback - Content-Type:
application/json
{
"skillName": "linkfox-mpstats-ozon-product-trend",
"sentiment": "POSITIVE",
"category": "OTHER",
"content": "Spotted a clean seasonal peak for the SKU."
}字段说明:
skillName:使用本 skill 的 YAMLnamesentiment:POSITIVE/NEUTRAL/NEGATIVEcategory:BUG/COMPLAINT/SUGGESTION/OTHERcontent:用户表达、实际现象、为什么算问题或好评
#!/usr/bin/env python3
"""
MPSTATS Ozon Product Trend (Daily) - LinkFox Skill
Returns daily time-series for a single Ozon SKU.
Usage:
python mpstats_ozon_product_trend.py '{"productId": 1786874757, "startDate": "2025-03-01", "endDate": "2025-03-31"}'
"""
import json
import os
import sys
if sys.stdout.encoding and sys.stdout.encoding.lower() != "utf-8":
try: sys.stdout.reconfigure(encoding="utf-8")
except Exception: pass
from urllib.request import urlopen, Request
from urllib.error import HTTPError, URLError
API_URL = "https://tool-gateway.linkfox.com/mpstats/ozon/productTrend"
def get_api_key():
key = os.environ.get("LINKFOXAGENT_API_KEY")
if not key:
print(
"API Key not configured. Please complete authorization first:\n"
"1. Visit https://skill.linkfox.com/linkfoxskills/guide.htm to obtain your Key\n"
"2. Set the environment variable: export LINKFOXAGENT_API_KEY=your-key-here",
file=sys.stderr,
)
sys.exit(1)
return key
def call_api(params: dict) -> dict:
api_key = get_api_key()
data = json.dumps(params).encode("utf-8")
req = Request(
API_URL,
data=data,
headers={
"Authorization": api_key,
"Content-Type": "application/json",
"User-Agent": "LinkFox-Skill/1.0",
},
method="POST",
)
try:
with urlopen(req, timeout=60) as response:
return json.loads(response.read().decode("utf-8"))
except HTTPError as e:
body = e.read().decode("utf-8") if e.fp else ""
return {"error": f"HTTP {e.code}: {e.reason}", "details": body}
except URLError as e:
return {"error": f"Connection failed: {e.reason}"}
def print_summary(result: dict):
if "error" in result:
print(f"Error: {result['error']}", file=sys.stderr)
if "details" in result:
print(f"Details: {result['details']}", file=sys.stderr)
return
points = result.get("data", []) or []
total = result.get("total", len(points))
print(f"Trend points: {len(points)} | total: {total}")
print("-" * 90)
print(f"{'date':<12} {'price':>10} {'oldPrice':>10} {'sales':>8} {'balance':>8} {'rating':>6} {'comments':>9} {'hasData':>8}")
print("-" * 90)
for p in points:
print(
f"{(p.get('date') or ''):<12} "
f"{(p.get('price') or 0):>10.2f} "
f"{(p.get('oldPrice') or 0):>10.2f} "
f"{(p.get('sales') or 0):>8} "
f"{(p.get('balance') or 0):>8} "
f"{(p.get('rating') or 0):>6.2f} "
f"{(p.get('comments') or 0):>9} "
f"{str(p.get('hasData')):>8}"
)
def main():
if len(sys.argv) < 2:
print("Usage: mpstats_ozon_product_trend.py '<JSON parameters>'", file=sys.stderr)
print(
"Example: mpstats_ozon_product_trend.py "
"'{\"productId\": 1786874757, \"startDate\": \"2025-03-01\", \"endDate\": \"2025-03-31\"}'",
file=sys.stderr,
)
sys.exit(1)
try:
params = json.loads(sys.argv[1])
except json.JSONDecodeError as e:
print(f"Invalid parameter format: {e}", file=sys.stderr)
sys.exit(1)
result = call_api(params)
if sys.stdout.isatty():
print_summary(result)
else:
print(json.dumps(result, indent=2, ensure_ascii=False))
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
Skill response I/O helper — wraps any main script to persist large API
responses to disk, then offers a `read` subcommand to extract specific fields
from those persisted files. Generic, business-agnostic.
This script is bundled into each skill's scripts/ directory by tools/response_io/sync.py.
The agent must pass --script <path> to identify which main script to execute.
Usage:
python scripts/response_io.py run --script <PATH> --out-dir <DIR> '<json_params>' [--label NAME] [--timeout SEC]
python scripts/response_io.py read <file> (--path "<JMESPath>" | --fields "f1,f2,...") [--limit N] [--offset M] [--format json|jsonl|csv|table]
"""
from __future__ import annotations
import sys
if sys.version_info < (3, 10):
sys.exit(
"Error: Python 3.10+ required (current: "
f"{sys.version_info.major}.{sys.version_info.minor}). "
"Please upgrade Python."
)
import argparse
import csv
import io
import json
import os
import re
import secrets
import subprocess
from datetime import datetime
from pathlib import Path
from typing import Any
# Force UTF-8 stdout/stderr so non-ASCII chars in previews and API responses
# print correctly on Windows (default cp936 / gbk).
for stream in (sys.stdout, sys.stderr):
try:
stream.reconfigure(encoding="utf-8") # type: ignore[attr-defined]
except (AttributeError, OSError):
pass
try:
import jmespath # type: ignore
HAS_JMESPATH = True
except ImportError:
HAS_JMESPATH = False
MAX_STRING_LEN = 120
MAX_DEPTH = 3
SAMPLE_KEY_CAP = 15
RAW_TEXT_PEEK = 500
DEFAULT_TIMEOUT_SEC = 300
# ---------------------------------------------------------------------------
# Shared helpers
# ---------------------------------------------------------------------------
def _err(msg: str, code: int = 1) -> None:
print(msg, file=sys.stderr)
sys.exit(code)
def _resolve_script(script_arg: str) -> Path:
p = Path(script_arg).expanduser()
if not p.is_absolute():
# Resolve relative to the current working directory the agent invoked from.
p = (Path.cwd() / p).resolve()
else:
p = p.resolve()
if not p.is_file():
_err(f"--script path not found: {p}")
return p
def _resolve_skill_name(main_script: Path) -> str:
"""Best-effort skill name extraction for filename prefixing.
main_script lives at <skill_dir>/scripts/<name>.py — return <skill_dir>'s
folder name. Fall back to the script's stem if structure differs.
"""
try:
if main_script.parent.name == "scripts":
return main_script.parents[1].name
except IndexError:
pass
return main_script.stem
def _sanitize_label(label: str) -> str:
"""Allow only safe filename chars in --label to prevent path traversal."""
cleaned = re.sub(r"[^\w\-]", "_", label)
return cleaned[:64] # cap length
def _truncate_string(s: str) -> str:
if len(s) <= MAX_STRING_LEN:
return s
return s[:MAX_STRING_LEN] + f"...(truncated, total {len(s)} chars)"
def _truncate_value(value: Any, depth: int = 0) -> Any:
"""Recursively truncate strings, deep nesting, and large arrays for preview."""
if depth >= MAX_DEPTH:
if isinstance(value, dict):
return f"<truncated nested object, keys: {list(value.keys())[:10]}>"
if isinstance(value, list):
return f"<truncated nested array, length: {len(value)}>"
if isinstance(value, str):
return _truncate_string(value)
return value
if isinstance(value, str):
return _truncate_string(value)
if isinstance(value, dict):
out = {k: _truncate_value(v, depth + 1) for k, v in value.items()}
return out
if isinstance(value, list):
if not value:
return []
truncated = [_truncate_value(value[0], depth + 1)]
if len(value) > 1:
# Note total length on the parent — keep the array type-homogeneous
# so downstream consumers can iterate without special-casing strings.
truncated.append({"_omitted_items": len(value) - 1})
return truncated
return value
def _shape_of(value: Any, top: bool = False) -> Any:
"""Lightweight schema description for the preview block."""
if isinstance(value, dict):
keys = list(value.keys())
out: dict[str, Any] = {"type": "object", "top_keys" if top else "keys": keys}
if top:
for k in keys[:8]:
out[k] = _shape_of(value[k])
return out
if isinstance(value, list):
out = {"type": "array", "length": len(value)}
if value and isinstance(value[0], dict):
out["item_keys"] = list(value[0].keys())
elif value:
out["item_type"] = type(value[0]).__name__
return out
return {"type": type(value).__name__}
def _build_sample(value: Any) -> Any:
"""First-record sample with explicit truncation marker."""
if isinstance(value, list):
if not value:
return {"_truncated_record": True, "_note": "array is empty"}
first = value[0]
if isinstance(first, dict):
sample = {"_truncated_record": True, "_note": f"first of {len(value)} items"}
sample.update(_truncate_value(first, depth=1))
return sample
return {"_truncated_record": True, "_note": f"first of {len(value)} items", "value": _truncate_value(first, depth=1)}
if isinstance(value, dict):
sample = {"_truncated_record": True, "_note": "top-level object (truncated)"}
sample.update(_truncate_value(value, depth=1))
return sample
return {"_truncated_record": True, "value": _truncate_value(value, depth=1)}
def _shrink_preview(preview: dict) -> dict:
"""Cap the sample's value fields when it has many keys.
`shape.*.item_keys` is the single source of truth for the full key list
(always complete, no truncation). The sample only ever shows up to
SAMPLE_KEY_CAP fields with their concrete values, since the agent only
needs a feel for value shapes — for the full menu of available fields,
they read `shape`.
"""
sample = preview.get("sample")
if isinstance(sample, dict):
meta_keys = {"_truncated_record", "_note"}
data_keys = [k for k in sample.keys() if k not in meta_keys]
if len(data_keys) > SAMPLE_KEY_CAP:
kept = data_keys[:SAMPLE_KEY_CAP]
new_sample = {k: v for k, v in sample.items() if k in meta_keys or k in kept}
base_note = sample.get("_note", "")
extra = (
f"showing first {SAMPLE_KEY_CAP} of {len(data_keys)} fields "
f"(see `shape` for the complete key list)"
)
new_sample["_note"] = f"{base_note}; {extra}" if base_note else extra
preview["sample"] = new_sample
return preview
# ---------------------------------------------------------------------------
# `run` subcommand
# ---------------------------------------------------------------------------
def cmd_run(args: argparse.Namespace) -> int:
main_script = _resolve_script(args.script)
skill_name = _resolve_skill_name(main_script)
out_dir = Path(args.out_dir).expanduser().resolve()
try:
out_dir.mkdir(parents=True, exist_ok=True)
except OSError as e:
_err(f"Failed to create --out-dir {out_dir}: {e}")
if not os.access(out_dir, os.W_OK):
_err(f"--out-dir is not writable: {out_dir}")
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
rand = secrets.token_hex(3)
safe_label = _sanitize_label(args.label) if args.label else ""
label_part = f"__{safe_label}" if safe_label else ""
out_file = out_dir / f"{skill_name}__{timestamp}_{rand}{label_part}.json"
# Force the child process to emit UTF-8 regardless of the host console
# encoding (Windows defaults to cp936 / gbk and would otherwise corrupt
# non-ASCII bytes when we read them back).
child_env = os.environ.copy()
child_env["PYTHONIOENCODING"] = "utf-8"
timed_out = False
try:
proc = subprocess.run(
[sys.executable, str(main_script), args.params],
capture_output=True,
text=True,
encoding="utf-8",
errors="replace",
env=child_env,
timeout=args.timeout,
)
stdout_text = proc.stdout or ""
stderr_text = proc.stderr or ""
returncode = proc.returncode
except subprocess.TimeoutExpired as e:
timed_out = True
stdout_text = (e.stdout.decode("utf-8", errors="replace") if isinstance(e.stdout, bytes) else (e.stdout or "")) or ""
stderr_text = (e.stderr.decode("utf-8", errors="replace") if isinstance(e.stderr, bytes) else (e.stderr or "")) or ""
returncode = 124 # convention for timeout
# Always write the captured stdout to disk, even if not JSON.
try:
out_file.write_text(stdout_text, encoding="utf-8")
except OSError as e:
_err(f"Failed to write output file {out_file}: {e}")
if stderr_text:
sys.stderr.write(stderr_text)
# Try to parse the captured stdout as JSON for the preview.
try:
parsed = json.loads(stdout_text) if stdout_text.strip() else None
format_kind = "json"
except json.JSONDecodeError:
parsed = None
format_kind = "raw_text"
preview: dict[str, Any] = {
"_preview": {
"is_preview": True,
"warning": (
"PREVIEW ONLY — NOT FULL DATA. The full response is saved to `file`. "
"Use `python scripts/response_io.py read <file> --fields '...'` to extract "
"specific fields, or `--path '<JMESPath>'` for complex projections."
),
},
}
# Surface failures prominently so agents don't mistake a stub preview for success.
if returncode != 0 or timed_out:
stderr_snippet = stderr_text[-500:] if stderr_text else ""
preview["_error"] = {
"exit_code": returncode,
"timed_out": timed_out,
"stderr_snippet": stderr_snippet,
"hint": "The wrapped script failed or timed out. The output file may be empty or partial.",
}
preview.update({
"file": str(out_file),
"size_bytes": out_file.stat().st_size,
"skill": skill_name,
"exit_code": returncode,
"format": format_kind,
"label": safe_label or None,
"next_steps_hint": (
"use: python scripts/response_io.py read <file> --fields '...' | --path '...'"
),
})
if format_kind == "json":
preview["shape"] = _shape_of(parsed, top=True)
preview["sample"] = _build_sample(parsed)
else:
peek = stdout_text[:RAW_TEXT_PEEK]
preview["raw_text_peek"] = peek
preview["raw_text_total_chars"] = len(stdout_text)
preview["sample"] = {
"_truncated_record": True,
"_note": f"stdout was not valid JSON; first {RAW_TEXT_PEEK} chars shown above in raw_text_peek",
}
preview = _shrink_preview(preview)
print(json.dumps(preview, ensure_ascii=False, indent=2))
return returncode
# ---------------------------------------------------------------------------
# `read` subcommand
# ---------------------------------------------------------------------------
def _load_json(path: Path) -> Any:
try:
text = path.read_text(encoding="utf-8")
except OSError as e:
_err(f"Failed to read file {path}: {e}")
try:
return json.loads(text)
except json.JSONDecodeError as e:
_err(f"File is not valid JSON: {path}\n{e}")
def _basic_dot_path(data: Any, path: str) -> Any:
"""Pure-stdlib dot-path resolver. No [*] support — callers fall back here only when jmespath is unavailable AND the path has no [*]."""
cur = data
for part in path.split("."):
if isinstance(cur, dict):
cur = cur.get(part)
else:
return None
return cur
def _resolve_field(data: Any, expr: str) -> Any:
if HAS_JMESPATH:
return jmespath.search(expr, data)
if "[" in expr or "*" in expr:
_err(
f"jmespath is required for expression '{expr}'. "
f"Install with: pip install jmespath"
)
return _basic_dot_path(data, expr)
def _project_fields(data: Any, fields: list[str]) -> Any:
"""Run each field expr; if any returns a list, zip them into list-of-dicts."""
resolved: dict[str, Any] = {f: _resolve_field(data, f) for f in fields}
list_lengths = [len(v) for v in resolved.values() if isinstance(v, list)]
if not list_lengths:
return resolved
# All list values must be same length to zip cleanly.
if len(set(list_lengths)) > 1:
# Fallback: return the dict as-is so caller can inspect mismatches.
return resolved
n = list_lengths[0]
rows = []
for i in range(n):
row = {}
for f, v in resolved.items():
row[f] = v[i] if isinstance(v, list) else v
rows.append(row)
return rows
def _apply_slice(value: Any, limit: int | None, offset: int | None) -> Any:
if not isinstance(value, list):
return value
start = offset or 0
end = (start + limit) if limit is not None else None
return value[start:end]
def _format_output(value: Any, fmt: str) -> str:
if fmt == "json":
return json.dumps(value, ensure_ascii=False, indent=2)
if fmt == "jsonl":
if isinstance(value, list):
return "\n".join(json.dumps(item, ensure_ascii=False) for item in value)
return json.dumps(value, ensure_ascii=False)
if fmt in ("csv", "table"):
if not isinstance(value, list) or not value:
_err(f"--format {fmt} requires a non-empty list result")
if not all(isinstance(item, dict) for item in value):
_err(f"--format {fmt} requires list-of-objects, got list of {type(value[0]).__name__}")
keys: list[str] = []
for item in value:
for k in item.keys():
if k not in keys:
keys.append(k)
if fmt == "csv":
buf = io.StringIO()
writer = csv.DictWriter(buf, fieldnames=keys, extrasaction="ignore")
writer.writeheader()
for item in value:
writer.writerow({k: _stringify(item.get(k)) for k in keys})
return buf.getvalue().rstrip("\n")
# table: simple aligned columns
rows = [[_stringify(item.get(k)) for k in keys] for item in value]
widths = [len(k) for k in keys]
for row in rows:
for i, cell in enumerate(row):
widths[i] = max(widths[i], len(cell))
lines = [
" ".join(k.ljust(widths[i]) for i, k in enumerate(keys)),
" ".join("-" * widths[i] for i in range(len(keys))),
]
for row in rows:
lines.append(" ".join(row[i].ljust(widths[i]) for i in range(len(keys))))
return "\n".join(lines)
_err(f"Unknown --format: {fmt}")
return "" # unreachable
def _stringify(v: Any) -> str:
if v is None:
return ""
if isinstance(v, (dict, list)):
return json.dumps(v, ensure_ascii=False)
return str(v)
def cmd_read(args: argparse.Namespace) -> int:
if not args.path and not args.fields:
_err("read: either --path or --fields is required")
if args.path and args.fields:
_err("read: --path and --fields are mutually exclusive")
file_path = Path(args.file).expanduser().resolve()
data = _load_json(file_path)
if args.path:
result = _resolve_field(data, args.path)
else:
fields = [f.strip() for f in args.fields.split(",") if f.strip()]
if not fields:
_err("--fields parsed to empty list")
result = _project_fields(data, fields)
result = _apply_slice(result, args.limit, args.offset)
print(_format_output(result, args.format))
return 0
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def main() -> int:
parser = argparse.ArgumentParser(
prog="response_io.py",
description="Persist large skill API responses to disk and read fields on demand.",
)
sub = parser.add_subparsers(dest="cmd", required=True)
p_run = sub.add_parser(
"run",
help="Execute a main script and persist its stdout to a file; "
"print only a lightweight preview to stdout.",
)
p_run.add_argument("params", help="JSON params string passed verbatim to the main script (argv[1]).")
p_run.add_argument("--script", required=True, help="Path to the main script to execute, e.g. scripts/my_api.py")
p_run.add_argument("--out-dir", required=True, help="Directory to write the response file into (created if missing).")
p_run.add_argument("--label", default=None, help="Optional filename suffix; sanitized to safe filename characters.")
p_run.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT_SEC, help=f"Subprocess timeout in seconds (default: {DEFAULT_TIMEOUT_SEC}).")
p_run.set_defaults(func=cmd_run)
p_read = sub.add_parser(
"read",
help="Extract specific fields from a previously persisted response file.",
)
p_read.add_argument("file", help="Path to the persisted JSON response file.")
g = p_read.add_mutually_exclusive_group()
g.add_argument("--path", default=None, help="JMESPath expression, e.g. 'data[*].{asin: asin, title: title}'.")
g.add_argument("--fields", default=None, help="Comma-separated field paths, e.g. 'data[*].asin,data[*].title'.")
p_read.add_argument("--limit", type=int, default=None, help="Take at most N items (when result is a list).")
p_read.add_argument("--offset", type=int, default=None, help="Skip the first M items (when result is a list).")
p_read.add_argument("--format", choices=["json", "jsonl", "csv", "table"], default="json", help="Output format (default: json).")
p_read.set_defaults(func=cmd_read)
args = parser.parse_args()
return args.func(args)
if __name__ == "__main__":
sys.exit(main())