
Linkfox Junglescout Sales Estimates
- 236 installs
- 64 repo stars
- Updated August 3, 2026
- linkfox-ai/linkfox-skills
Estimate monthly Amazon unit sales and revenue for candidate ASINs via Jungle Scout to decide if demand, velocity, and margin assumptions are worth pursuing.
About
Uses Jungle Scout sales estimates to project Amazon unit volume and revenue for one or more ASINs. Agents can compare candidate products, stress-test margin scenarios, and reject weak opportunities before investing in listings, ads, or supplier MOQs.
- Jungle Scout sales velocity estimates
- Revenue potential modeling by ASIN
- Demand validation before sourcing
- Competitive benchmark comparisons
- Agent-driven market sizing
Linkfox Junglescout Sales Estimates by the numbers
- 236 all-time installs (skills.sh)
- +36 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #281 of 853 Sales & Marketing skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/linkfox-ai/linkfox-skills --skill linkfox-junglescout-sales-estimatesAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 236 |
|---|---|
| repo stars | ★ 64 |
| Last updated | August 3, 2026 |
| Repository | linkfox-ai/linkfox-skills ↗ |
What it does
Estimate monthly Amazon unit sales and revenue for candidate ASINs via Jungle Scout to decide if demand, velocity, and margin assumptions are worth pursuing.
Files
Jungle Scout — ASIN 销售估算
This skill queries daily sales estimates and last known price for a given Amazon ASIN via the Jungle Scout data source, returning day-level data points over a specified date range across 10 Amazon marketplaces.
Core Concepts
Jungle Scout ASIN 销售估算工具提供亚马逊各站点单个 ASIN 的日维度预估销量及最近已知价格。卖家可以通过查询指定时间范围内的销量变化来:
- 监控竞品销量:了解竞品每日出单量,评估其市场份额
- 验证选品机会:用实际销量数据验证产品需求是否足够大
- 追踪季节性规律:观察产品在不同月份的销量波动,判断旺淡季
- 评估定价影响:结合价格与销量的变化关系,辅助定价决策
- 新品表现跟踪:追踪新品上架后的销量爬升曲线
数据粒度:每条记录代表 1 天,包含该日的预估售出件数和最近已知价格(美元)。
Data Fields
Output Fields
| Field | API Name | Description | Example |
|---|---|---|---|
| ASIN | asin | 查询的 ASIN | B0CXXX1234 |
| 数据标识 | id | 数据点标识 | sales_estimate_B0CXXX1234_20260301 |
| 资源类型 | type | 固定值 | sales_estimate_result |
| 父 ASIN | parentAsin | 父体 ASIN(变体场景) | B0CXXX0000 |
| 是否父体 | isParent | 是否为父体商品 | true / false |
| 是否变体 | isVariant | 是否为变体商品 | true / false |
| 是否独立 | isStandalone | 是否为独立商品(非变体) | true / false |
| 变体列表 | variants | 该父体下的变体 ASIN 数组 | ["B0CX1", "B0CX2"] |
| 每日估算 | dailyEstimates | 每日数据数组 | 见下方 |
| 消耗 Token | costToken | 本次调用消耗的 token 数 | 1 |
dailyEstimates 数组中每个对象
| Field | API Name | Description | Example |
|---|---|---|---|
| 日期 | date | 数据日期(YYYY-MM-DD) | 2026-03-15 |
| 预估日销量 | estimatedUnitsSold | 当日预估售出件数 | 42 |
| 最近已知价格 | lastKnownPrice | 最近已知价格(USD) | 29.99 |
Supported Marketplaces
10 个亚马逊站点:us(美国)、uk(英国)、de(德国)、in(印度)、ca(加拿大)、fr(法国)、it(意大利)、es(西班牙)、mx(墨西哥)、jp(日本)。默认站点为 us。当用户未指定站点时,使用 us。
API Usage
This tool calls the LinkFox tool gateway API. See references/api.md for calling conventions, request parameters, and response structure. You can also execute scripts/junglescout_sales_estimates.py directly to run queries.
How to Build Queries
所有四个参数均为必填:marketplace、asin、startDate、endDate。
Principles for Building API Calls
1. 站点映射:用户说"美国站"→ us,"日本站"→ jp,"德国站"→ de;未指定时默认 us 2. 日期格式:必须为 YYYY-MM-DD,如 2026-03-01 3. endDate 限制:endDate 必须早于当前日期(不能包含今天及未来日期) 4. ASIN 格式:标准亚马逊 ASIN,通常以 B0 开头,共 10 位 5. 常用时间推算:
- "过去30天" → endDate 取昨天,startDate 取30天前
- "上个月" → 上月1日到上月末日
- "Q3 vs Q4" → 分两次调用,分别查 7-9 月和 10-12 月
Common Query Scenarios
1. 查看竞品最近30天的销量
{
"marketplace": "us",
"asin": "B0CXXX1234",
"startDate": "2026-03-18",
"endDate": "2026-04-16"
}2. 对比 Q3 与 Q4 销量表现
分两次调用:
- Q3:
startDate=2025-07-01,endDate=2025-09-30 - Q4:
startDate=2025-10-01,endDate=2025-12-31
3. 验证选品机会——查看产品全年销量
{
"marketplace": "us",
"asin": "B0CXXX5678",
"startDate": "2025-04-01",
"endDate": "2026-03-31"
}4. 追踪新品上架表现
{
"marketplace": "de",
"asin": "B0DYYY9999",
"startDate": "2026-01-15",
"endDate": "2026-04-15"
}5. 监控大促期间销量变化(如 Prime Day)
{
"marketplace": "us",
"asin": "B0CXXX1234",
"startDate": "2025-07-01",
"endDate": "2025-07-21"
}Display Rules
1. 折线图优先:建议以折线图展示每日销量变化,横轴为日期,纵轴为预估日销量;如有价格数据可叠加第二 Y 轴显示价格走势 2. 表格辅助:同时提供数据表格供精确查阅,列包括:日期、预估销量、最近已知价格 3. 汇总统计:在数据之后汇总关键指标——总销量、日均销量、预估总收入(总销量 × 均价) 4. 趋势总结:简要总结趋势方向(上升/下降/平稳/周期性波动),标注销量峰值和谷值日期 5. Error handling: When a query fails, explain the reason based on the error response and suggest adjusting parameters(如 endDate 不能包含今天或未来日期)
Important Limitations
- endDate 不可包含今天:
endDate必须早于当前日期,不能查询今天及未来的销量 - 单次单 ASIN:每次调用只能查询一个 ASIN;对比多个 ASIN 需分多次调用
- 所有参数必填:
marketplace、asin、startDate、endDate缺一不可 - 价格为美元:
lastKnownPrice单位为 USD,非本地货币
User Expression & Scenario Quick Reference
Applicable - ASIN 销售估算与销量趋势分析:
| User Says | Scenario |
|---|---|
| "这个ASIN一天能卖多少" | 查询近期日销量估算 |
| "竞品最近卖得怎么样" | 监控竞品近30天销量 |
| "这个产品有没有季节性" | 全年销量数据判断季节规律 |
| "帮我看看这个品的销量趋势" | 指定时间段的销量走势 |
| "Q4旺季销量如何" | 特定季度销量查询 |
| "这个产品值不值得做" | 通过历史销量验证选品机会 |
| "大促期间卖了多少" | 活动期间销量监控 |
Not applicable - 超出 ASIN 销售估算范围:
- 关键词搜索量(需要关键词历史搜索量工具)
- BSR 排名历史(需要 BSR 追踪工具)
- 类目整体销量/市场规模
- 非亚马逊平台的销量数据
- 实时/当前时刻的销量(数据有滞后,不含今天)
Boundary judgment: When users say "销量", "日销", or "卖了多少", if they want to see a specific ASIN's daily estimated sales over a time range, this skill applies. If they want keyword search volume, category rankings, or real-time live sales, it does not apply.
Feedback:
Auto-detect and report feedback via the Feedback API when any of the following apply: 1. The functionality or purpose described in this skill does not match actual behavior 2. The skill's results do not match the user's intent 3. The user expresses dissatisfaction or praise about this skill 4. Anything you believe could be improved
Call the feedback API as specified in references/api.md. Do not interrupt the user's flow.
<!-- LF_LARGE_RESPONSE_BLOCK -->
Handling Large Responses
To avoid overflowing the agent context, persist the response to disk and extract only the fields you need:
python scripts/response_io.py run --script scripts/junglescout_sales_estimates.py --out-dir <DIR> '<params>'
python scripts/response_io.py read <file> --fields "<paths>" # or --path "<JMESPath>"Pick--out-diroutside any git working tree (e.g./tmp/...on Unix,%TEMP%/...on Windows). Persisted responses may contain PII, pricing, or auth-sensitive data — do not commit them. Files are not auto-deleted; clean up when the task is done.
run writes the full response to a file and emits only a schema preview + file path. read projects specific fields, with --limit/--offset for slicing and --format json|jsonl|csv|table for output.
When to prefer this pattern — apply your judgment based on the response characteristics, e.g.:
- High field count per record, or fields you don't need
- Batch/paginated results (multiple items per call)
- Long-text fields (descriptions, reviews, HTML, time series)
- Output reused across later steps rather than consumed immediately
For small, single-use responses, calling the main script directly is fine.
⚠️ The preview is a truncated schema + sample, not the full data. Any field-level decision must read from the persisted file via read. <!-- /LF_LARGE_RESPONSE_BLOCK -->
--- For more high-quality, professional cross-border e-commerce skills, visit [LinkFox Skills](https://skill.linkfox.com/).
Jungle Scout ASIN 销售估算 API 参考
调用规范
- 请求地址:
https://tool-gateway.linkfox.com/tool-jungle-scout/sales-estimates/query - 请求方式:POST,Content-Type: application/json
- 认证方式:Header
Authorization: <api_key>,api_key 从环境变量LINKFOXAGENT_API_KEY读取(如未配置,提示用户前往 https://skill.linkfox.com/linkfoxskills/guide.htm 申请)
请求参数
POST Body(JSON):
| 参数 | 类型 | 必填 | 说明 |
|---|---|---|---|
| marketplace | string | 是 | 目标市场代码。可选值:us、uk、de、in、ca、fr、it、es、mx、jp |
| asin | string | 是 | 要查询的亚马逊 ASIN |
| startDate | string | 是 | 开始日期(格式:YYYY-MM-DD) |
| endDate | string | 是 | 结束日期(格式:YYYY-MM-DD);必须早于当前日期 |
站点映射
| 站点 | marketplace 值 |
|---|---|
| 美国 | us |
| 英国 | uk |
| 德国 | de |
| 印度 | in |
| 加拿大 | ca |
| 法国 | fr |
| 意大利 | it |
| 西班牙 | es |
| 墨西哥 | mx |
| 日本 | jp |
响应结构
| 字段 | 类型 | 说明 |
|---|---|---|
| costToken | integer | 消耗 token 数 |
| salesEstimateList | array | 销售估算结果列表 |
salesEstimateList 数组中每个对象
| 字段 | 类型 | 说明 |
|---|---|---|
| asin | string | 查询的 ASIN |
| id | string | 数据点标识 |
| type | string | 资源类型,固定值 sales_estimate_result |
| parentAsin | string | 父体 ASIN(变体场景下返回) |
| isParent | boolean | 是否为父体商品 |
| isVariant | boolean | 是否为变体商品 |
| isStandalone | boolean | 是否为独立商品(非变体) |
| variants | array | 该父体下的变体 ASIN 数组 |
| dailyEstimates | array | 每日估算数据数组 |
dailyEstimates 数组中每个对象
| 字段 | 类型 | 说明 |
|---|---|---|
| date | string | 数据日期(YYYY-MM-DD) |
| estimatedUnitsSold | integer | 当日预估售出件数 |
| lastKnownPrice | number | 最近已知价格(USD) |
错误码
正常情况下,接口的 HTTP 状态码均为 200,业务的成功与否通过响应体中的 errorCode 字段区分(errorCode = 200 表示成功,其他值表示业务错误)。当遇到未授权等情况时,HTTP 状态码为 401,且对应的 errorCode 也是 401。
| errcode | 含义 | 处理建议 |
|---|---|---|
| 200 | 成功 | 正常解析 salesEstimateList |
| 401 | 认证失败 | 检查请求头 Authorization 是否正确携带 API Key |
| 其他非200值 | 业务异常 | 参考 errmsg 字段获取具体错误原因 |
错误响应示例:
{
"errcode": 401,
"errmsg": "authorized error"
}curl 示例
curl -X POST https://tool-gateway.linkfox.com/tool-jungle-scout/sales-estimates/query \
-H "Authorization: $LINKFOXAGENT_API_KEY" \
-H "Content-Type: application/json" \
-d '{"marketplace": "us", "asin": "B0CXXX1234", "startDate": "2026-03-01", "endDate": "2026-03-31"}'响应示例
{
"costToken": 1,
"salesEstimateList": [
{
"asin": "B0CXXX1234",
"id": "sales_estimate_B0CXXX1234_20260301",
"type": "sales_estimate_result",
"parentAsin": "B0CXXX0000",
"isParent": false,
"isVariant": true,
"isStandalone": false,
"variants": [],
"dailyEstimates": [
{
"date": "2026-03-01",
"estimatedUnitsSold": 35,
"lastKnownPrice": 29.99
},
{
"date": "2026-03-02",
"estimatedUnitsSold": 42,
"lastKnownPrice": 29.99
},
{
"date": "2026-03-03",
"estimatedUnitsSold": 38,
"lastKnownPrice": 27.99
}
]
}
]
}---
Feedback API
This endpoint is separate from the tool API above. Do not mix the two base URLs.
- POST
https://skill-api.linkfox.com/api/v1/public/feedback - Content-Type:
application/json
{
"skillName": "linkfox-junglescout-sales-estimates",
"sentiment": "POSITIVE",
"category": "OTHER",
"content": "Results were accurate, user was satisfied."
}Field rules:
skillName: Use this skill'snamefrom the YAML frontmattersentiment: Choose ONE —POSITIVE(praise),NEUTRAL(suggestion without emotion),NEGATIVE(complaint or error)category: Choose ONE —BUG(malfunction or wrong data),COMPLAINT(user dissatisfaction),SUGGESTION(improvement idea),OTHERcontent: Include what the user said or intended, what actually happened, and why it is a problem or praise
#!/usr/bin/env python3
"""
Jungle Scout — ASIN 销售估算 - LinkFox Skill
Calls the tool-jungle-scout/sales-estimates/query API endpoint
Usage:
python junglescout_sales_estimates.py '{"marketplace": "us", "asin": "B0CXXX1234", "startDate": "2026-03-01", "endDate": "2026-03-31"}'
"""
import json
import os
import sys
from urllib.request import urlopen, Request
from urllib.error import HTTPError, URLError
API_URL = "https://tool-gateway.linkfox.com/tool-jungle-scout/sales-estimates/query"
REQUIRED_PARAMS = ["marketplace", "asin", "startDate", "endDate"]
VALID_MARKETPLACES = {"us", "uk", "de", "in", "ca", "fr", "it", "es", "mx", "jp"}
def get_api_key():
"""Retrieve the API key from environment, with a friendly prompt if missing."""
key = os.environ.get("LINKFOXAGENT_API_KEY")
if not key:
print(
"API Key not configured. Please complete authorization first:\n"
"1. Visit https://skill.linkfox.com/linkfoxskills/guide.htm to obtain your Key\n"
"2. Set the environment variable: export LINKFOXAGENT_API_KEY=your-key-here",
file=sys.stderr,
)
sys.exit(1)
return key
def validate_params(params: dict):
"""Validate all required parameters."""
missing = [p for p in REQUIRED_PARAMS if p not in params]
if missing:
print(f"Error: missing required parameters: {', '.join(missing)}", file=sys.stderr)
sys.exit(1)
mp = params["marketplace"]
if mp not in VALID_MARKETPLACES:
print(
f"Error: invalid marketplace '{mp}'. "
f"Valid values: {', '.join(sorted(VALID_MARKETPLACES))}",
file=sys.stderr,
)
sys.exit(1)
for date_field in ("startDate", "endDate"):
val = params[date_field]
if len(val) != 10 or val[4] != "-" or val[7] != "-":
print(f"Error: {date_field} must be in YYYY-MM-DD format, got '{val}'", file=sys.stderr)
sys.exit(1)
if params["endDate"] >= _today():
print("Error: endDate must be before the current date", file=sys.stderr)
sys.exit(1)
if params["startDate"] > params["endDate"]:
print("Error: startDate must be before or equal to endDate", file=sys.stderr)
sys.exit(1)
def _today() -> str:
"""Return today's date as YYYY-MM-DD without importing datetime."""
import time
t = time.localtime()
return f"{t.tm_year:04d}-{t.tm_mon:02d}-{t.tm_mday:02d}"
def call_api(params: dict) -> dict:
"""Call the tool gateway API."""
api_key = get_api_key()
data = json.dumps(params).encode("utf-8")
req = Request(
API_URL,
data=data,
headers={
"Authorization": api_key,
"Content-Type": "application/json",
"User-Agent": "LinkFox-Skill/1.0",
},
method="POST",
)
try:
with urlopen(req, timeout=60) as response:
return json.loads(response.read().decode("utf-8"))
except HTTPError as e:
body = e.read().decode("utf-8") if e.fp else ""
return {"error": f"HTTP {e.code}: {e.reason}", "details": body}
except URLError as e:
return {"error": f"Connection failed: {e.reason}"}
def main():
if len(sys.argv) < 2:
print("Usage: junglescout_sales_estimates.py '<JSON parameters>'", file=sys.stderr)
print(
'Example: junglescout_sales_estimates.py \'{"marketplace": "us", "asin": "B0CXXX1234", '
'"startDate": "2026-03-01", "endDate": "2026-03-31"}\'',
file=sys.stderr,
)
sys.exit(1)
try:
params = json.loads(sys.argv[1])
except json.JSONDecodeError as e:
print(f"Invalid parameter format: {e}", file=sys.stderr)
sys.exit(1)
validate_params(params)
result = call_api(params)
print(json.dumps(result, indent=2, ensure_ascii=False))
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
Skill response I/O helper — wraps any main script to persist large API
responses to disk, then offers a `read` subcommand to extract specific fields
from those persisted files. Generic, business-agnostic.
This script is bundled into each skill's scripts/ directory by tools/response_io/sync.py.
The agent must pass --script <path> to identify which main script to execute.
Usage:
python scripts/response_io.py run --script <PATH> --out-dir <DIR> '<json_params>' [--label NAME] [--timeout SEC]
python scripts/response_io.py read <file> (--path "<JMESPath>" | --fields "f1,f2,...") [--limit N] [--offset M] [--format json|jsonl|csv|table]
"""
from __future__ import annotations
import sys
if sys.version_info < (3, 10):
sys.exit(
"Error: Python 3.10+ required (current: "
f"{sys.version_info.major}.{sys.version_info.minor}). "
"Please upgrade Python."
)
import argparse
import csv
import io
import json
import os
import re
import secrets
import subprocess
from datetime import datetime
from pathlib import Path
from typing import Any
# Force UTF-8 stdout/stderr so non-ASCII chars in previews and API responses
# print correctly on Windows (default cp936 / gbk).
for stream in (sys.stdout, sys.stderr):
try:
stream.reconfigure(encoding="utf-8") # type: ignore[attr-defined]
except (AttributeError, OSError):
pass
try:
import jmespath # type: ignore
HAS_JMESPATH = True
except ImportError:
HAS_JMESPATH = False
MAX_STRING_LEN = 120
MAX_DEPTH = 3
SAMPLE_KEY_CAP = 15
RAW_TEXT_PEEK = 500
DEFAULT_TIMEOUT_SEC = 300
# ---------------------------------------------------------------------------
# Shared helpers
# ---------------------------------------------------------------------------
def _err(msg: str, code: int = 1) -> None:
print(msg, file=sys.stderr)
sys.exit(code)
def _resolve_script(script_arg: str) -> Path:
p = Path(script_arg).expanduser()
if not p.is_absolute():
# Resolve relative to the current working directory the agent invoked from.
p = (Path.cwd() / p).resolve()
else:
p = p.resolve()
if not p.is_file():
_err(f"--script path not found: {p}")
return p
def _resolve_skill_name(main_script: Path) -> str:
"""Best-effort skill name extraction for filename prefixing.
main_script lives at <skill_dir>/scripts/<name>.py — return <skill_dir>'s
folder name. Fall back to the script's stem if structure differs.
"""
try:
if main_script.parent.name == "scripts":
return main_script.parents[1].name
except IndexError:
pass
return main_script.stem
def _sanitize_label(label: str) -> str:
"""Allow only safe filename chars in --label to prevent path traversal."""
cleaned = re.sub(r"[^\w\-]", "_", label)
return cleaned[:64] # cap length
def _truncate_string(s: str) -> str:
if len(s) <= MAX_STRING_LEN:
return s
return s[:MAX_STRING_LEN] + f"...(truncated, total {len(s)} chars)"
def _truncate_value(value: Any, depth: int = 0) -> Any:
"""Recursively truncate strings, deep nesting, and large arrays for preview."""
if depth >= MAX_DEPTH:
if isinstance(value, dict):
return f"<truncated nested object, keys: {list(value.keys())[:10]}>"
if isinstance(value, list):
return f"<truncated nested array, length: {len(value)}>"
if isinstance(value, str):
return _truncate_string(value)
return value
if isinstance(value, str):
return _truncate_string(value)
if isinstance(value, dict):
out = {k: _truncate_value(v, depth + 1) for k, v in value.items()}
return out
if isinstance(value, list):
if not value:
return []
truncated = [_truncate_value(value[0], depth + 1)]
if len(value) > 1:
# Note total length on the parent — keep the array type-homogeneous
# so downstream consumers can iterate without special-casing strings.
truncated.append({"_omitted_items": len(value) - 1})
return truncated
return value
def _shape_of(value: Any, top: bool = False) -> Any:
"""Lightweight schema description for the preview block."""
if isinstance(value, dict):
keys = list(value.keys())
out: dict[str, Any] = {"type": "object", "top_keys" if top else "keys": keys}
if top:
for k in keys[:8]:
out[k] = _shape_of(value[k])
return out
if isinstance(value, list):
out = {"type": "array", "length": len(value)}
if value and isinstance(value[0], dict):
out["item_keys"] = list(value[0].keys())
elif value:
out["item_type"] = type(value[0]).__name__
return out
return {"type": type(value).__name__}
def _build_sample(value: Any) -> Any:
"""First-record sample with explicit truncation marker."""
if isinstance(value, list):
if not value:
return {"_truncated_record": True, "_note": "array is empty"}
first = value[0]
if isinstance(first, dict):
sample = {"_truncated_record": True, "_note": f"first of {len(value)} items"}
sample.update(_truncate_value(first, depth=1))
return sample
return {"_truncated_record": True, "_note": f"first of {len(value)} items", "value": _truncate_value(first, depth=1)}
if isinstance(value, dict):
sample = {"_truncated_record": True, "_note": "top-level object (truncated)"}
sample.update(_truncate_value(value, depth=1))
return sample
return {"_truncated_record": True, "value": _truncate_value(value, depth=1)}
def _shrink_preview(preview: dict) -> dict:
"""Cap the sample's value fields when it has many keys.
`shape.*.item_keys` is the single source of truth for the full key list
(always complete, no truncation). The sample only ever shows up to
SAMPLE_KEY_CAP fields with their concrete values, since the agent only
needs a feel for value shapes — for the full menu of available fields,
they read `shape`.
"""
sample = preview.get("sample")
if isinstance(sample, dict):
meta_keys = {"_truncated_record", "_note"}
data_keys = [k for k in sample.keys() if k not in meta_keys]
if len(data_keys) > SAMPLE_KEY_CAP:
kept = data_keys[:SAMPLE_KEY_CAP]
new_sample = {k: v for k, v in sample.items() if k in meta_keys or k in kept}
base_note = sample.get("_note", "")
extra = (
f"showing first {SAMPLE_KEY_CAP} of {len(data_keys)} fields "
f"(see `shape` for the complete key list)"
)
new_sample["_note"] = f"{base_note}; {extra}" if base_note else extra
preview["sample"] = new_sample
return preview
# ---------------------------------------------------------------------------
# `run` subcommand
# ---------------------------------------------------------------------------
def cmd_run(args: argparse.Namespace) -> int:
main_script = _resolve_script(args.script)
skill_name = _resolve_skill_name(main_script)
out_dir = Path(args.out_dir).expanduser().resolve()
try:
out_dir.mkdir(parents=True, exist_ok=True)
except OSError as e:
_err(f"Failed to create --out-dir {out_dir}: {e}")
if not os.access(out_dir, os.W_OK):
_err(f"--out-dir is not writable: {out_dir}")
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
rand = secrets.token_hex(3)
safe_label = _sanitize_label(args.label) if args.label else ""
label_part = f"__{safe_label}" if safe_label else ""
out_file = out_dir / f"{skill_name}__{timestamp}_{rand}{label_part}.json"
# Force the child process to emit UTF-8 regardless of the host console
# encoding (Windows defaults to cp936 / gbk and would otherwise corrupt
# non-ASCII bytes when we read them back).
child_env = os.environ.copy()
child_env["PYTHONIOENCODING"] = "utf-8"
timed_out = False
try:
proc = subprocess.run(
[sys.executable, str(main_script), args.params],
capture_output=True,
text=True,
encoding="utf-8",
errors="replace",
env=child_env,
timeout=args.timeout,
)
stdout_text = proc.stdout or ""
stderr_text = proc.stderr or ""
returncode = proc.returncode
except subprocess.TimeoutExpired as e:
timed_out = True
stdout_text = (e.stdout.decode("utf-8", errors="replace") if isinstance(e.stdout, bytes) else (e.stdout or "")) or ""
stderr_text = (e.stderr.decode("utf-8", errors="replace") if isinstance(e.stderr, bytes) else (e.stderr or "")) or ""
returncode = 124 # convention for timeout
# Always write the captured stdout to disk, even if not JSON.
try:
out_file.write_text(stdout_text, encoding="utf-8")
except OSError as e:
_err(f"Failed to write output file {out_file}: {e}")
if stderr_text:
sys.stderr.write(stderr_text)
# Try to parse the captured stdout as JSON for the preview.
try:
parsed = json.loads(stdout_text) if stdout_text.strip() else None
format_kind = "json"
except json.JSONDecodeError:
parsed = None
format_kind = "raw_text"
preview: dict[str, Any] = {
"_preview": {
"is_preview": True,
"warning": (
"PREVIEW ONLY — NOT FULL DATA. The full response is saved to `file`. "
"Use `python scripts/response_io.py read <file> --fields '...'` to extract "
"specific fields, or `--path '<JMESPath>'` for complex projections."
),
},
}
# Surface failures prominently so agents don't mistake a stub preview for success.
if returncode != 0 or timed_out:
stderr_snippet = stderr_text[-500:] if stderr_text else ""
preview["_error"] = {
"exit_code": returncode,
"timed_out": timed_out,
"stderr_snippet": stderr_snippet,
"hint": "The wrapped script failed or timed out. The output file may be empty or partial.",
}
preview.update({
"file": str(out_file),
"size_bytes": out_file.stat().st_size,
"skill": skill_name,
"exit_code": returncode,
"format": format_kind,
"label": safe_label or None,
"next_steps_hint": (
"use: python scripts/response_io.py read <file> --fields '...' | --path '...'"
),
})
if format_kind == "json":
preview["shape"] = _shape_of(parsed, top=True)
preview["sample"] = _build_sample(parsed)
else:
peek = stdout_text[:RAW_TEXT_PEEK]
preview["raw_text_peek"] = peek
preview["raw_text_total_chars"] = len(stdout_text)
preview["sample"] = {
"_truncated_record": True,
"_note": f"stdout was not valid JSON; first {RAW_TEXT_PEEK} chars shown above in raw_text_peek",
}
preview = _shrink_preview(preview)
print(json.dumps(preview, ensure_ascii=False, indent=2))
return returncode
# ---------------------------------------------------------------------------
# `read` subcommand
# ---------------------------------------------------------------------------
def _load_json(path: Path) -> Any:
try:
text = path.read_text(encoding="utf-8")
except OSError as e:
_err(f"Failed to read file {path}: {e}")
try:
return json.loads(text)
except json.JSONDecodeError as e:
_err(f"File is not valid JSON: {path}\n{e}")
def _basic_dot_path(data: Any, path: str) -> Any:
"""Pure-stdlib dot-path resolver. No [*] support — callers fall back here only when jmespath is unavailable AND the path has no [*]."""
cur = data
for part in path.split("."):
if isinstance(cur, dict):
cur = cur.get(part)
else:
return None
return cur
def _resolve_field(data: Any, expr: str) -> Any:
if HAS_JMESPATH:
return jmespath.search(expr, data)
if "[" in expr or "*" in expr:
_err(
f"jmespath is required for expression '{expr}'. "
f"Install with: pip install jmespath"
)
return _basic_dot_path(data, expr)
def _project_fields(data: Any, fields: list[str]) -> Any:
"""Run each field expr; if any returns a list, zip them into list-of-dicts."""
resolved: dict[str, Any] = {f: _resolve_field(data, f) for f in fields}
list_lengths = [len(v) for v in resolved.values() if isinstance(v, list)]
if not list_lengths:
return resolved
# All list values must be same length to zip cleanly.
if len(set(list_lengths)) > 1:
# Fallback: return the dict as-is so caller can inspect mismatches.
return resolved
n = list_lengths[0]
rows = []
for i in range(n):
row = {}
for f, v in resolved.items():
row[f] = v[i] if isinstance(v, list) else v
rows.append(row)
return rows
def _apply_slice(value: Any, limit: int | None, offset: int | None) -> Any:
if not isinstance(value, list):
return value
start = offset or 0
end = (start + limit) if limit is not None else None
return value[start:end]
def _format_output(value: Any, fmt: str) -> str:
if fmt == "json":
return json.dumps(value, ensure_ascii=False, indent=2)
if fmt == "jsonl":
if isinstance(value, list):
return "\n".join(json.dumps(item, ensure_ascii=False) for item in value)
return json.dumps(value, ensure_ascii=False)
if fmt in ("csv", "table"):
if not isinstance(value, list) or not value:
_err(f"--format {fmt} requires a non-empty list result")
if not all(isinstance(item, dict) for item in value):
_err(f"--format {fmt} requires list-of-objects, got list of {type(value[0]).__name__}")
keys: list[str] = []
for item in value:
for k in item.keys():
if k not in keys:
keys.append(k)
if fmt == "csv":
buf = io.StringIO()
writer = csv.DictWriter(buf, fieldnames=keys, extrasaction="ignore")
writer.writeheader()
for item in value:
writer.writerow({k: _stringify(item.get(k)) for k in keys})
return buf.getvalue().rstrip("\n")
# table: simple aligned columns
rows = [[_stringify(item.get(k)) for k in keys] for item in value]
widths = [len(k) for k in keys]
for row in rows:
for i, cell in enumerate(row):
widths[i] = max(widths[i], len(cell))
lines = [
" ".join(k.ljust(widths[i]) for i, k in enumerate(keys)),
" ".join("-" * widths[i] for i in range(len(keys))),
]
for row in rows:
lines.append(" ".join(row[i].ljust(widths[i]) for i in range(len(keys))))
return "\n".join(lines)
_err(f"Unknown --format: {fmt}")
return "" # unreachable
def _stringify(v: Any) -> str:
if v is None:
return ""
if isinstance(v, (dict, list)):
return json.dumps(v, ensure_ascii=False)
return str(v)
def cmd_read(args: argparse.Namespace) -> int:
if not args.path and not args.fields:
_err("read: either --path or --fields is required")
if args.path and args.fields:
_err("read: --path and --fields are mutually exclusive")
file_path = Path(args.file).expanduser().resolve()
data = _load_json(file_path)
if args.path:
result = _resolve_field(data, args.path)
else:
fields = [f.strip() for f in args.fields.split(",") if f.strip()]
if not fields:
_err("--fields parsed to empty list")
result = _project_fields(data, fields)
result = _apply_slice(result, args.limit, args.offset)
print(_format_output(result, args.format))
return 0
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def main() -> int:
parser = argparse.ArgumentParser(
prog="response_io.py",
description="Persist large skill API responses to disk and read fields on demand.",
)
sub = parser.add_subparsers(dest="cmd", required=True)
p_run = sub.add_parser(
"run",
help="Execute a main script and persist its stdout to a file; "
"print only a lightweight preview to stdout.",
)
p_run.add_argument("params", help="JSON params string passed verbatim to the main script (argv[1]).")
p_run.add_argument("--script", required=True, help="Path to the main script to execute, e.g. scripts/my_api.py")
p_run.add_argument("--out-dir", required=True, help="Directory to write the response file into (created if missing).")
p_run.add_argument("--label", default=None, help="Optional filename suffix; sanitized to safe filename characters.")
p_run.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT_SEC, help=f"Subprocess timeout in seconds (default: {DEFAULT_TIMEOUT_SEC}).")
p_run.set_defaults(func=cmd_run)
p_read = sub.add_parser(
"read",
help="Extract specific fields from a previously persisted response file.",
)
p_read.add_argument("file", help="Path to the persisted JSON response file.")
g = p_read.add_mutually_exclusive_group()
g.add_argument("--path", default=None, help="JMESPath expression, e.g. 'data[*].{asin: asin, title: title}'.")
g.add_argument("--fields", default=None, help="Comma-separated field paths, e.g. 'data[*].asin,data[*].title'.")
p_read.add_argument("--limit", type=int, default=None, help="Take at most N items (when result is a list).")
p_read.add_argument("--offset", type=int, default=None, help="Skip the first M items (when result is a list).")
p_read.add_argument("--format", choices=["json", "jsonl", "csv", "table"], default="json", help="Output format (default: json).")
p_read.set_defaults(func=cmd_read)
args = parser.parse_args()
return args.func(args)
if __name__ == "__main__":
sys.exit(main())