
Linkfox Junglescout Keyword By Keyword
- 235 installs
- 64 repo stars
- Updated August 3, 2026
- linkfox-ai/linkfox-skills
Helps with ai & agent building tasks.
About
linkfox-junglescout-keyword-by-keyword is a Claude Code skill in the AI & Agent Building category.
- linkfox-junglescout-keyword-by-keyword
- AI & Agent Building
- AI-coding skill
Linkfox Junglescout Keyword By Keyword by the numbers
- 235 all-time installs (skills.sh)
- +35 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #2,651 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/linkfox-ai/linkfox-skills --skill linkfox-junglescout-keyword-by-keywordAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 235 |
|---|---|
| repo stars | ★ 64 |
| Last updated | August 3, 2026 |
| Repository | linkfox-ai/linkfox-skills ↗ |
What it does
Helps with ai & agent building tasks.
Files
Jungle Scout — 根据关键词扩展关键词信息 (Keyword by Keyword)
This skill expands a seed keyword into a list of related keywords with search volume, trends, PPC bids, ranking difficulty, and other competitive metrics via the Jungle Scout data source, covering 10 Amazon marketplaces.
Core Concepts
Jungle Scout Keyword by Keyword 工具是亚马逊关键词研究的核心工具之一,从一个种子关键词出发,挖掘与之相关的大量关键词及其竞争指标。主要应用场景包括:
- 关键词拓展/发现:输入核心词,获取数百个相关关键词,扩充 listing 关键词库
- 长尾词挖掘:通过
minWordCount筛选 3+ 词的长尾关键词,发现低竞争高转化机会 - PPC 竞价研究:查看精确/广泛匹配 PPC 出价和品牌广告出价,规划广告预算
- 竞争度评估:通过
easeOfRankingScore和organicProductCount判断关键词排名难度 - 趋势分析:查看月度趋势和季度趋势百分比变化,识别增长型关键词
Data Fields
Output Fields (keywordInfoList)
| Field | API Name | Description | Example |
|---|---|---|---|
| 关键词 | name | 关键词名称 | yoga mat thick |
| 站点 | country | 市场代码 | us |
| 精确搜索量 | monthlySearchVolumeExact | 月均精确匹配搜索量 | 45000 |
| 广泛搜索量 | monthlySearchVolumeBroad | 月均广泛匹配搜索量 | 120000 |
| 月度趋势 | monthlyTrend | 环比月度搜索量变化百分比 | 15.3 |
| 季度趋势 | quarterlyTrend | 环比季度搜索量变化百分比 | -5.2 |
| 主类目 | dominantCategory | 搜索结果中占比最高的品类 | Sports & Outdoors |
| 相关性评分 | relevancyScore | 与种子词的相关性评分 | 856 |
| 排名难度 | easeOfRankingScore | 排名容易度评分(越高越容易) | 3 |
| 自然商品数 | organicProductCount | 搜索结果中的自然排名商品数量 | 342 |
| 广告商品数 | sponsoredProductCount | 搜索结果中的广告商品数量 | 28 |
| PPC精确出价 | ppcBidExact | 精确匹配 PPC 建议出价(美元) | 1.25 |
| PPC广泛出价 | ppcBidBroad | 广泛匹配 PPC 建议出价(美元) | 0.89 |
| 品牌广告出价 | spBrandAdBid | Sponsored Brand 广告建议出价(美元) | 2.50 |
| 推荐促销 | recommendedPromotions | 推荐促销赠品数量 | 150 |
| 消耗Token | costToken | 本次调用消耗的 token 数 | 1 |
Supported Marketplaces
10 Amazon marketplaces: us (default), uk, de, in, ca, fr, it, es, mx, jp. When the user does not specify a marketplace, use us.
| 站点 | marketplace 值 | 说明 |
|---|---|---|
| 美国 | us | Amazon.com |
| 英国 | uk | Amazon.co.uk |
| 德国 | de | Amazon.de |
| 印度 | in | Amazon.in |
| 加拿大 | ca | Amazon.ca |
| 法国 | fr | Amazon.fr |
| 意大利 | it | Amazon.it |
| 西班牙 | es | Amazon.es |
| 墨西哥 | mx | Amazon.com.mx |
| 日本 | jp | Amazon.co.jp |
API Usage
This tool calls the LinkFox tool gateway API. See references/api.md for calling conventions, request parameters, and response structure. You can also execute scripts/junglescout_keyword_by_keyword.py directly to run queries.
How to Build Queries
必填参数:marketplace、searchTerms(单个种子关键词字符串)。
Principles for Building API Calls
1. 站点映射:用户说"美国站"→ us,"日本站"→ jp,"德国站"→ de;未指定时默认 us 2. 种子关键词:原样传入用户提供的关键词(英文小写为佳),仅支持单个关键词 3. 结果数量:默认返回数量有限,若用户需要更多结果,设置 needCount 4. 排序选择:默认按精确搜索量降序 (-monthly_search_volume_exact),根据用户意图切换排序字段 5. 筛选过滤:充分利用 min/max 参数缩小结果范围,避免返回无关低质量关键词
Common Query Scenarios
1. 扩展种子关键词 — 获取相关关键词列表
{
"marketplace": "us",
"searchTerms": "yoga mat"
}2. 挖掘长尾关键词(3+ 词)
{
"marketplace": "us",
"searchTerms": "yoga mat",
"minWordCount": 3,
"needCount": 50
}3. 低竞争关键词发现
{
"marketplace": "us",
"searchTerms": "yoga mat",
"maxOrganicProductCount": 200,
"minMonthlySearchVolumeExact": 1000,
"sort": "-ease_of_ranking_score"
}4. 高搜索量关键词筛选
{
"marketplace": "us",
"searchTerms": "yoga mat",
"minMonthlySearchVolumeExact": 10000,
"sort": "-monthly_search_volume_exact",
"needCount": 30
}5. PPC 竞价研究 — 按广泛出价排序
{
"marketplace": "us",
"searchTerms": "yoga mat",
"minMonthlySearchVolumeExact": 500,
"sort": "ppc_bid_broad",
"needCount": 30
}6. 德国站广泛搜索量关键词
{
"marketplace": "de",
"searchTerms": "yogamatte",
"minMonthlySearchVolumeBroad": 5000,
"sort": "-monthly_search_volume_broad"
}Display Rules
1. 表格优先:以表格展示关键词列表,核心列包括:关键词、精确搜索量、广泛搜索量、月度趋势、PPC精确出价、排名难度 2. 按需裁剪列:根据用户意图决定展示列——PPC研究场景侧重出价列,拓词场景侧重搜索量和趋势 3. 趋势标注:月度趋势和季度趋势为正值标注上升↑,负值标注下降↓ 4. 排名难度解读:easeOfRankingScore 1-3 为困难,4-6 为中等,7-10 为容易 5. 数据洞察:在表格后提供简要总结,如高搜索量词集中在哪个类目、长尾词的竞争优势等 6. Error handling: When a query fails, explain the reason based on the error response and suggest adjusting parameters
Important Limitations
- 单次单关键词:
searchTerms仅接受一个种子关键词,多关键词需拆分多次调用 - 数据周期:搜索量为月均估算值,非实时数据
- 站点限制:仅覆盖 10 个亚马逊站点,不含澳大利亚、荷兰等
- 排序字段固定:仅支持预定义的排序字段,不支持自定义组合排序
User Expression & Scenario Quick Reference
Applicable — 关键词拓展与竞争分析:
| User Says | Scenario |
|---|---|
| "帮我拓展这个关键词" | 种子词扩展 |
| "这个词有哪些相关关键词" | 相关词挖掘 |
| "找一些长尾词" | 长尾关键词筛选(minWordCount ≥ 3) |
| "竞争度低的词有哪些" | 低竞争关键词(排名难度 + 商品数量筛选) |
| "这个词的PPC出价多少" | PPC 竞价数据查询 |
| "搜索量大的相关词" | 高搜索量关键词筛选 |
| "德国站有什么相关词" | 非美国站关键词拓展 |
| "帮我做关键词调研" | 综合关键词研究 |
Not applicable — 超出关键词拓展范围:
- 关键词历史搜索量趋势(需要 keyword-history 工具)
- ABA 搜索词排名(需要 ABA 工具)
- 商品搜索或 listing 分析
- 非亚马逊平台的关键词数据
- ASIN 反查关键词
Boundary judgment: When users say "关键词", "拓词", or "关键词研究", if they want to expand a seed keyword into a list of related keywords with metrics, this skill applies. If they want to see a single keyword's historical search volume trend over time, use the keyword-history skill instead.
Feedback:
Auto-detect and report feedback via the Feedback API when any of the following apply: 1. The functionality or purpose described in this skill does not match actual behavior 2. The skill's results do not match the user's intent 3. The user expresses dissatisfaction or praise about this skill 4. Anything you believe could be improved
Call the feedback API as specified in references/api.md. Do not interrupt the user's flow.
<!-- LF_LARGE_RESPONSE_BLOCK -->
Handling Large Responses
To avoid overflowing the agent context, persist the response to disk and extract only the fields you need:
python scripts/response_io.py run --script scripts/junglescout_keyword_by_keyword.py --out-dir <DIR> '<params>'
python scripts/response_io.py read <file> --fields "<paths>" # or --path "<JMESPath>"Pick--out-diroutside any git working tree (e.g./tmp/...on Unix,%TEMP%/...on Windows). Persisted responses may contain PII, pricing, or auth-sensitive data — do not commit them. Files are not auto-deleted; clean up when the task is done.
run writes the full response to a file and emits only a schema preview + file path. read projects specific fields, with --limit/--offset for slicing and --format json|jsonl|csv|table for output.
When to prefer this pattern — apply your judgment based on the response characteristics, e.g.:
- High field count per record, or fields you don't need
- Batch/paginated results (multiple items per call)
- Long-text fields (descriptions, reviews, HTML, time series)
- Output reused across later steps rather than consumed immediately
For small, single-use responses, calling the main script directly is fine.
⚠️ The preview is a truncated schema + sample, not the full data. Any field-level decision must read from the persisted file via read. <!-- /LF_LARGE_RESPONSE_BLOCK -->
--- For more high-quality, professional cross-border e-commerce skills, visit [LinkFox Skills](https://skill.linkfox.com/).
Jungle Scout 根据关键词扩展关键词信息 API 参考
调用规范
- 请求地址:
https://tool-gateway.linkfox.com/tool-jungle-scout/keywords/by-keyword - 请求方式:POST,Content-Type: application/json
- 认证方式:Header
Authorization: <api_key>,api_key 从环境变量LINKFOXAGENT_API_KEY读取(如未配置,提示用户前往 https://skill.linkfox.com/linkfoxskills/guide.htm 申请)
请求参数
POST Body(JSON):
必填参数
| 参数 | 类型 | 必填 | 说明 |
|---|---|---|---|
| marketplace | string | 是 | 目标市场代码。可选值:us、uk、de、in、ca、fr、it、es、mx、jp。默认 us |
| searchTerms | string | 是 | 种子关键词(单个关键词字符串) |
可选参数 — 结果控制
| 参数 | 类型 | 必填 | 说明 |
|---|---|---|---|
| needCount | int | 否 | 返回结果总数 |
| sort | string | 否 | 排序字段,默认 -monthly_search_volume_exact(精确搜索量降序) |
可选参数 — 搜索量筛选
| 参数 | 类型 | 必填 | 说明 |
|---|---|---|---|
| minMonthlySearchVolumeExact | int | 否 | 精确搜索量下限 |
| maxMonthlySearchVolumeExact | int | 否 | 精确搜索量上限 |
| minMonthlySearchVolumeBroad | int | 否 | 广泛搜索量下限 |
| maxMonthlySearchVolumeBroad | int | 否 | 广泛搜索量上限 |
可选参数 — 其他筛选
| 参数 | 类型 | 必填 | 说明 |
|---|---|---|---|
| minWordCount | int | 否 | 关键词最少词数(用于筛选长尾词) |
| maxWordCount | int | 否 | 关键词最多词数 |
| minOrganicProductCount | int | 否 | 自然排名商品数下限 |
| maxOrganicProductCount | int | 否 | 自然排名商品数上限 |
sort 可选值
| 值 | 说明 |
|---|---|
| name / -name | 关键词名称 升序/降序 |
| dominant_category / -dominant_category | 主类目 升序/降序 |
| monthly_trend / -monthly_trend | 月度趋势 升序/降序 |
| quarterly_trend / -quarterly_trend | 季度趋势 升序/降序 |
| monthly_search_volume_exact / -monthly_search_volume_exact | 精确搜索量 升序/降序(默认降序) |
| monthly_search_volume_broad / -monthly_search_volume_broad | 广泛搜索量 升序/降序 |
| recommended_promotions / -recommended_promotions | 推荐促销 升序/降序 |
| sp_brand_ad_bid / -sp_brand_ad_bid | 品牌广告出价 升序/降序 |
| ppc_bid_broad / -ppc_bid_broad | PPC广泛出价 升序/降序 |
| ppc_bid_exact / -ppc_bid_exact | PPC精确出价 升序/降序 |
| ease_of_ranking_score / -ease_of_ranking_score | 排名难度 升序/降序 |
| relevancy_score / -relevancy_score | 相关性评分 升序/降序 |
| organic_product_count / -organic_product_count | 自然商品数 升序/降序 |
站点映射
| 站点 | marketplace 值 |
|---|---|
| 美国 | us |
| 英国 | uk |
| 德国 | de |
| 印度 | in |
| 加拿大 | ca |
| 法国 | fr |
| 意大利 | it |
| 西班牙 | es |
| 墨西哥 | mx |
| 日本 | jp |
响应结构
| 字段 | 类型 | 说明 |
|---|---|---|
| costToken | integer | 消耗 token 数 |
| keywordInfoList | array | 关键词信息列表 |
keywordInfoList 数组中每个对象
| 字段 | 类型 | 说明 |
|---|---|---|
| name | string | 关键词名称 |
| country | string | 市场代码 |
| monthlySearchVolumeExact | integer | 月均精确匹配搜索量 |
| monthlySearchVolumeBroad | integer | 月均广泛匹配搜索量 |
| monthlyTrend | number | 月度搜索量变化百分比 |
| quarterlyTrend | number | 季度搜索量变化百分比 |
| dominantCategory | string | 搜索结果中占比最高的品类 |
| relevancyScore | integer | 与种子词的相关性评分 |
| easeOfRankingScore | integer | 排名容易度评分(越高越容易排名) |
| organicProductCount | integer | 自然排名商品数量 |
| sponsoredProductCount | integer | 广告商品数量 |
| ppcBidExact | number | 精确匹配 PPC 建议出价(美元) |
| ppcBidBroad | number | 广泛匹配 PPC 建议出价(美元) |
| spBrandAdBid | number | Sponsored Brand 广告建议出价(美元) |
| recommendedPromotions | integer | 推荐促销赠品数量 |
错误码
正常情况下,接口的 HTTP 状态码均为 200,业务的成功与否通过响应体中的 errorCode 字段区分(errorCode = 200 表示成功,其他值表示业务错误)。当遇到未授权等情况时,HTTP 状态码为 401,且对应的 errorCode 也是 401。
| errcode | 含义 | 处理建议 |
|---|---|---|
| 200 | 成功 | 正常解析 keywordInfoList |
| 401 | 认证失败 | 检查请求头 Authorization 是否正确携带 API Key |
| 其他非200值 | 业务异常 | 参考 errmsg 字段获取具体错误原因 |
错误响应示例:
{
"errcode": 401,
"errmsg": "authorized error"
}curl 示例
curl -X POST https://tool-gateway.linkfox.com/tool-jungle-scout/keywords/by-keyword \
-H "Authorization: $LINKFOXAGENT_API_KEY" \
-H "Content-Type: application/json" \
-d '{"marketplace": "us", "searchTerms": "yoga mat", "needCount": 20}'响应示例
{
"costToken": 1,
"keywordInfoList": [
{
"name": "yoga mat thick",
"country": "us",
"monthlySearchVolumeExact": 45000,
"monthlySearchVolumeBroad": 120000,
"monthlyTrend": 15.3,
"quarterlyTrend": -5.2,
"dominantCategory": "Sports & Outdoors",
"relevancyScore": 856,
"easeOfRankingScore": 3,
"organicProductCount": 342,
"sponsoredProductCount": 28,
"ppcBidExact": 1.25,
"ppcBidBroad": 0.89,
"spBrandAdBid": 2.50,
"recommendedPromotions": 150
},
{
"name": "yoga mat non slip",
"country": "us",
"monthlySearchVolumeExact": 38000,
"monthlySearchVolumeBroad": 95000,
"monthlyTrend": 8.1,
"quarterlyTrend": 12.4,
"dominantCategory": "Sports & Outdoors",
"relevancyScore": 920,
"easeOfRankingScore": 2,
"organicProductCount": 510,
"sponsoredProductCount": 35,
"ppcBidExact": 1.58,
"ppcBidBroad": 1.12,
"spBrandAdBid": 3.10,
"recommendedPromotions": 200
}
]
}---
Feedback API
This endpoint is separate from the tool API above. Do not mix the two base URLs.
- POST
https://skill-api.linkfox.com/api/v1/public/feedback - Content-Type:
application/json
{
"skillName": "linkfox-junglescout-keyword-by-keyword",
"sentiment": "POSITIVE",
"category": "OTHER",
"content": "Results were accurate, user was satisfied."
}Field rules:
skillName: Use this skill'snamefrom the YAML frontmattersentiment: Choose ONE —POSITIVE(praise),NEUTRAL(suggestion without emotion),NEGATIVE(complaint or error)category: Choose ONE —BUG(malfunction or wrong data),COMPLAINT(user dissatisfaction),SUGGESTION(improvement idea),OTHERcontent: Include what the user said or intended, what actually happened, and why it is a problem or praise
#!/usr/bin/env python3
"""
Jungle Scout — 根据关键词扩展关键词信息 (Keyword by Keyword) - LinkFox Skill
Calls the tool-jungle-scout/keywords/by-keyword API endpoint
Usage:
python junglescout_keyword_by_keyword.py '{"marketplace": "us", "searchTerms": "yoga mat"}'
python junglescout_keyword_by_keyword.py '{"marketplace": "us", "searchTerms": "yoga mat", "needCount": 50, "minWordCount": 3}'
"""
import json
import os
import sys
from urllib.request import urlopen, Request
from urllib.error import HTTPError, URLError
API_URL = "https://tool-gateway.linkfox.com/tool-jungle-scout/keywords/by-keyword"
REQUIRED_PARAMS = ["marketplace", "searchTerms"]
VALID_MARKETPLACES = {"us", "uk", "de", "in", "ca", "fr", "it", "es", "mx", "jp"}
VALID_SORT_VALUES = {
"name", "-name",
"dominant_category", "-dominant_category",
"monthly_trend", "-monthly_trend",
"quarterly_trend", "-quarterly_trend",
"monthly_search_volume_exact", "-monthly_search_volume_exact",
"monthly_search_volume_broad", "-monthly_search_volume_broad",
"recommended_promotions", "-recommended_promotions",
"sp_brand_ad_bid", "-sp_brand_ad_bid",
"ppc_bid_broad", "-ppc_bid_broad",
"ppc_bid_exact", "-ppc_bid_exact",
"ease_of_ranking_score", "-ease_of_ranking_score",
"relevancy_score", "-relevancy_score",
"organic_product_count", "-organic_product_count",
}
def get_api_key():
"""Retrieve the API key from environment, with a friendly prompt if missing."""
key = os.environ.get("LINKFOXAGENT_API_KEY")
if not key:
print(
"API Key not configured. Please complete authorization first:\n"
"1. Visit https://skill.linkfox.com/linkfoxskills/guide.htm to obtain your Key\n"
"2. Set the environment variable: export LINKFOXAGENT_API_KEY=your-key-here",
file=sys.stderr,
)
sys.exit(1)
return key
def validate_params(params: dict):
"""Validate required and optional parameters."""
missing = [p for p in REQUIRED_PARAMS if p not in params]
if missing:
print(f"Error: missing required parameters: {', '.join(missing)}", file=sys.stderr)
sys.exit(1)
mp = params.get("marketplace", "")
if mp not in VALID_MARKETPLACES:
print(
f"Error: invalid marketplace '{mp}'. Must be one of: {', '.join(sorted(VALID_MARKETPLACES))}",
file=sys.stderr,
)
sys.exit(1)
if "sort" in params and params["sort"] not in VALID_SORT_VALUES:
print(
f"Error: invalid sort value '{params['sort']}'. See references/api.md for valid values.",
file=sys.stderr,
)
sys.exit(1)
def call_api(params: dict) -> dict:
"""Call the tool gateway API."""
api_key = get_api_key()
data = json.dumps(params).encode("utf-8")
req = Request(
API_URL,
data=data,
headers={
"Authorization": api_key,
"Content-Type": "application/json",
"User-Agent": "LinkFox-Skill/1.0",
},
method="POST",
)
try:
with urlopen(req, timeout=60) as response:
return json.loads(response.read().decode("utf-8"))
except HTTPError as e:
body = e.read().decode("utf-8") if e.fp else ""
return {"error": f"HTTP {e.code}: {e.reason}", "details": body}
except URLError as e:
return {"error": f"Connection failed: {e.reason}"}
def main():
if len(sys.argv) < 2:
print("Usage: junglescout_keyword_by_keyword.py '<JSON parameters>'", file=sys.stderr)
print(
'Example: junglescout_keyword_by_keyword.py \'{"marketplace": "us", "searchTerms": "yoga mat"}\'',
file=sys.stderr,
)
sys.exit(1)
try:
params = json.loads(sys.argv[1])
except json.JSONDecodeError as e:
print(f"Invalid parameter format: {e}", file=sys.stderr)
sys.exit(1)
validate_params(params)
result = call_api(params)
print(json.dumps(result, indent=2, ensure_ascii=False))
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
Skill response I/O helper — wraps any main script to persist large API
responses to disk, then offers a `read` subcommand to extract specific fields
from those persisted files. Generic, business-agnostic.
This script is bundled into each skill's scripts/ directory by tools/response_io/sync.py.
The agent must pass --script <path> to identify which main script to execute.
Usage:
python scripts/response_io.py run --script <PATH> --out-dir <DIR> '<json_params>' [--label NAME] [--timeout SEC]
python scripts/response_io.py read <file> (--path "<JMESPath>" | --fields "f1,f2,...") [--limit N] [--offset M] [--format json|jsonl|csv|table]
"""
from __future__ import annotations
import sys
if sys.version_info < (3, 10):
sys.exit(
"Error: Python 3.10+ required (current: "
f"{sys.version_info.major}.{sys.version_info.minor}). "
"Please upgrade Python."
)
import argparse
import csv
import io
import json
import os
import re
import secrets
import subprocess
from datetime import datetime
from pathlib import Path
from typing import Any
# Force UTF-8 stdout/stderr so non-ASCII chars in previews and API responses
# print correctly on Windows (default cp936 / gbk).
for stream in (sys.stdout, sys.stderr):
try:
stream.reconfigure(encoding="utf-8") # type: ignore[attr-defined]
except (AttributeError, OSError):
pass
try:
import jmespath # type: ignore
HAS_JMESPATH = True
except ImportError:
HAS_JMESPATH = False
MAX_STRING_LEN = 120
MAX_DEPTH = 3
SAMPLE_KEY_CAP = 15
RAW_TEXT_PEEK = 500
DEFAULT_TIMEOUT_SEC = 300
# ---------------------------------------------------------------------------
# Shared helpers
# ---------------------------------------------------------------------------
def _err(msg: str, code: int = 1) -> None:
print(msg, file=sys.stderr)
sys.exit(code)
def _resolve_script(script_arg: str) -> Path:
p = Path(script_arg).expanduser()
if not p.is_absolute():
# Resolve relative to the current working directory the agent invoked from.
p = (Path.cwd() / p).resolve()
else:
p = p.resolve()
if not p.is_file():
_err(f"--script path not found: {p}")
return p
def _resolve_skill_name(main_script: Path) -> str:
"""Best-effort skill name extraction for filename prefixing.
main_script lives at <skill_dir>/scripts/<name>.py — return <skill_dir>'s
folder name. Fall back to the script's stem if structure differs.
"""
try:
if main_script.parent.name == "scripts":
return main_script.parents[1].name
except IndexError:
pass
return main_script.stem
def _sanitize_label(label: str) -> str:
"""Allow only safe filename chars in --label to prevent path traversal."""
cleaned = re.sub(r"[^\w\-]", "_", label)
return cleaned[:64] # cap length
def _truncate_string(s: str) -> str:
if len(s) <= MAX_STRING_LEN:
return s
return s[:MAX_STRING_LEN] + f"...(truncated, total {len(s)} chars)"
def _truncate_value(value: Any, depth: int = 0) -> Any:
"""Recursively truncate strings, deep nesting, and large arrays for preview."""
if depth >= MAX_DEPTH:
if isinstance(value, dict):
return f"<truncated nested object, keys: {list(value.keys())[:10]}>"
if isinstance(value, list):
return f"<truncated nested array, length: {len(value)}>"
if isinstance(value, str):
return _truncate_string(value)
return value
if isinstance(value, str):
return _truncate_string(value)
if isinstance(value, dict):
out = {k: _truncate_value(v, depth + 1) for k, v in value.items()}
return out
if isinstance(value, list):
if not value:
return []
truncated = [_truncate_value(value[0], depth + 1)]
if len(value) > 1:
# Note total length on the parent — keep the array type-homogeneous
# so downstream consumers can iterate without special-casing strings.
truncated.append({"_omitted_items": len(value) - 1})
return truncated
return value
def _shape_of(value: Any, top: bool = False) -> Any:
"""Lightweight schema description for the preview block."""
if isinstance(value, dict):
keys = list(value.keys())
out: dict[str, Any] = {"type": "object", "top_keys" if top else "keys": keys}
if top:
for k in keys[:8]:
out[k] = _shape_of(value[k])
return out
if isinstance(value, list):
out = {"type": "array", "length": len(value)}
if value and isinstance(value[0], dict):
out["item_keys"] = list(value[0].keys())
elif value:
out["item_type"] = type(value[0]).__name__
return out
return {"type": type(value).__name__}
def _build_sample(value: Any) -> Any:
"""First-record sample with explicit truncation marker."""
if isinstance(value, list):
if not value:
return {"_truncated_record": True, "_note": "array is empty"}
first = value[0]
if isinstance(first, dict):
sample = {"_truncated_record": True, "_note": f"first of {len(value)} items"}
sample.update(_truncate_value(first, depth=1))
return sample
return {"_truncated_record": True, "_note": f"first of {len(value)} items", "value": _truncate_value(first, depth=1)}
if isinstance(value, dict):
sample = {"_truncated_record": True, "_note": "top-level object (truncated)"}
sample.update(_truncate_value(value, depth=1))
return sample
return {"_truncated_record": True, "value": _truncate_value(value, depth=1)}
def _shrink_preview(preview: dict) -> dict:
"""Cap the sample's value fields when it has many keys.
`shape.*.item_keys` is the single source of truth for the full key list
(always complete, no truncation). The sample only ever shows up to
SAMPLE_KEY_CAP fields with their concrete values, since the agent only
needs a feel for value shapes — for the full menu of available fields,
they read `shape`.
"""
sample = preview.get("sample")
if isinstance(sample, dict):
meta_keys = {"_truncated_record", "_note"}
data_keys = [k for k in sample.keys() if k not in meta_keys]
if len(data_keys) > SAMPLE_KEY_CAP:
kept = data_keys[:SAMPLE_KEY_CAP]
new_sample = {k: v for k, v in sample.items() if k in meta_keys or k in kept}
base_note = sample.get("_note", "")
extra = (
f"showing first {SAMPLE_KEY_CAP} of {len(data_keys)} fields "
f"(see `shape` for the complete key list)"
)
new_sample["_note"] = f"{base_note}; {extra}" if base_note else extra
preview["sample"] = new_sample
return preview
# ---------------------------------------------------------------------------
# `run` subcommand
# ---------------------------------------------------------------------------
def cmd_run(args: argparse.Namespace) -> int:
main_script = _resolve_script(args.script)
skill_name = _resolve_skill_name(main_script)
out_dir = Path(args.out_dir).expanduser().resolve()
try:
out_dir.mkdir(parents=True, exist_ok=True)
except OSError as e:
_err(f"Failed to create --out-dir {out_dir}: {e}")
if not os.access(out_dir, os.W_OK):
_err(f"--out-dir is not writable: {out_dir}")
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
rand = secrets.token_hex(3)
safe_label = _sanitize_label(args.label) if args.label else ""
label_part = f"__{safe_label}" if safe_label else ""
out_file = out_dir / f"{skill_name}__{timestamp}_{rand}{label_part}.json"
# Force the child process to emit UTF-8 regardless of the host console
# encoding (Windows defaults to cp936 / gbk and would otherwise corrupt
# non-ASCII bytes when we read them back).
child_env = os.environ.copy()
child_env["PYTHONIOENCODING"] = "utf-8"
timed_out = False
try:
proc = subprocess.run(
[sys.executable, str(main_script), args.params],
capture_output=True,
text=True,
encoding="utf-8",
errors="replace",
env=child_env,
timeout=args.timeout,
)
stdout_text = proc.stdout or ""
stderr_text = proc.stderr or ""
returncode = proc.returncode
except subprocess.TimeoutExpired as e:
timed_out = True
stdout_text = (e.stdout.decode("utf-8", errors="replace") if isinstance(e.stdout, bytes) else (e.stdout or "")) or ""
stderr_text = (e.stderr.decode("utf-8", errors="replace") if isinstance(e.stderr, bytes) else (e.stderr or "")) or ""
returncode = 124 # convention for timeout
# Always write the captured stdout to disk, even if not JSON.
try:
out_file.write_text(stdout_text, encoding="utf-8")
except OSError as e:
_err(f"Failed to write output file {out_file}: {e}")
if stderr_text:
sys.stderr.write(stderr_text)
# Try to parse the captured stdout as JSON for the preview.
try:
parsed = json.loads(stdout_text) if stdout_text.strip() else None
format_kind = "json"
except json.JSONDecodeError:
parsed = None
format_kind = "raw_text"
preview: dict[str, Any] = {
"_preview": {
"is_preview": True,
"warning": (
"PREVIEW ONLY — NOT FULL DATA. The full response is saved to `file`. "
"Use `python scripts/response_io.py read <file> --fields '...'` to extract "
"specific fields, or `--path '<JMESPath>'` for complex projections."
),
},
}
# Surface failures prominently so agents don't mistake a stub preview for success.
if returncode != 0 or timed_out:
stderr_snippet = stderr_text[-500:] if stderr_text else ""
preview["_error"] = {
"exit_code": returncode,
"timed_out": timed_out,
"stderr_snippet": stderr_snippet,
"hint": "The wrapped script failed or timed out. The output file may be empty or partial.",
}
preview.update({
"file": str(out_file),
"size_bytes": out_file.stat().st_size,
"skill": skill_name,
"exit_code": returncode,
"format": format_kind,
"label": safe_label or None,
"next_steps_hint": (
"use: python scripts/response_io.py read <file> --fields '...' | --path '...'"
),
})
if format_kind == "json":
preview["shape"] = _shape_of(parsed, top=True)
preview["sample"] = _build_sample(parsed)
else:
peek = stdout_text[:RAW_TEXT_PEEK]
preview["raw_text_peek"] = peek
preview["raw_text_total_chars"] = len(stdout_text)
preview["sample"] = {
"_truncated_record": True,
"_note": f"stdout was not valid JSON; first {RAW_TEXT_PEEK} chars shown above in raw_text_peek",
}
preview = _shrink_preview(preview)
print(json.dumps(preview, ensure_ascii=False, indent=2))
return returncode
# ---------------------------------------------------------------------------
# `read` subcommand
# ---------------------------------------------------------------------------
def _load_json(path: Path) -> Any:
try:
text = path.read_text(encoding="utf-8")
except OSError as e:
_err(f"Failed to read file {path}: {e}")
try:
return json.loads(text)
except json.JSONDecodeError as e:
_err(f"File is not valid JSON: {path}\n{e}")
def _basic_dot_path(data: Any, path: str) -> Any:
"""Pure-stdlib dot-path resolver. No [*] support — callers fall back here only when jmespath is unavailable AND the path has no [*]."""
cur = data
for part in path.split("."):
if isinstance(cur, dict):
cur = cur.get(part)
else:
return None
return cur
def _resolve_field(data: Any, expr: str) -> Any:
if HAS_JMESPATH:
return jmespath.search(expr, data)
if "[" in expr or "*" in expr:
_err(
f"jmespath is required for expression '{expr}'. "
f"Install with: pip install jmespath"
)
return _basic_dot_path(data, expr)
def _project_fields(data: Any, fields: list[str]) -> Any:
"""Run each field expr; if any returns a list, zip them into list-of-dicts."""
resolved: dict[str, Any] = {f: _resolve_field(data, f) for f in fields}
list_lengths = [len(v) for v in resolved.values() if isinstance(v, list)]
if not list_lengths:
return resolved
# All list values must be same length to zip cleanly.
if len(set(list_lengths)) > 1:
# Fallback: return the dict as-is so caller can inspect mismatches.
return resolved
n = list_lengths[0]
rows = []
for i in range(n):
row = {}
for f, v in resolved.items():
row[f] = v[i] if isinstance(v, list) else v
rows.append(row)
return rows
def _apply_slice(value: Any, limit: int | None, offset: int | None) -> Any:
if not isinstance(value, list):
return value
start = offset or 0
end = (start + limit) if limit is not None else None
return value[start:end]
def _format_output(value: Any, fmt: str) -> str:
if fmt == "json":
return json.dumps(value, ensure_ascii=False, indent=2)
if fmt == "jsonl":
if isinstance(value, list):
return "\n".join(json.dumps(item, ensure_ascii=False) for item in value)
return json.dumps(value, ensure_ascii=False)
if fmt in ("csv", "table"):
if not isinstance(value, list) or not value:
_err(f"--format {fmt} requires a non-empty list result")
if not all(isinstance(item, dict) for item in value):
_err(f"--format {fmt} requires list-of-objects, got list of {type(value[0]).__name__}")
keys: list[str] = []
for item in value:
for k in item.keys():
if k not in keys:
keys.append(k)
if fmt == "csv":
buf = io.StringIO()
writer = csv.DictWriter(buf, fieldnames=keys, extrasaction="ignore")
writer.writeheader()
for item in value:
writer.writerow({k: _stringify(item.get(k)) for k in keys})
return buf.getvalue().rstrip("\n")
# table: simple aligned columns
rows = [[_stringify(item.get(k)) for k in keys] for item in value]
widths = [len(k) for k in keys]
for row in rows:
for i, cell in enumerate(row):
widths[i] = max(widths[i], len(cell))
lines = [
" ".join(k.ljust(widths[i]) for i, k in enumerate(keys)),
" ".join("-" * widths[i] for i in range(len(keys))),
]
for row in rows:
lines.append(" ".join(row[i].ljust(widths[i]) for i in range(len(keys))))
return "\n".join(lines)
_err(f"Unknown --format: {fmt}")
return "" # unreachable
def _stringify(v: Any) -> str:
if v is None:
return ""
if isinstance(v, (dict, list)):
return json.dumps(v, ensure_ascii=False)
return str(v)
def cmd_read(args: argparse.Namespace) -> int:
if not args.path and not args.fields:
_err("read: either --path or --fields is required")
if args.path and args.fields:
_err("read: --path and --fields are mutually exclusive")
file_path = Path(args.file).expanduser().resolve()
data = _load_json(file_path)
if args.path:
result = _resolve_field(data, args.path)
else:
fields = [f.strip() for f in args.fields.split(",") if f.strip()]
if not fields:
_err("--fields parsed to empty list")
result = _project_fields(data, fields)
result = _apply_slice(result, args.limit, args.offset)
print(_format_output(result, args.format))
return 0
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def main() -> int:
parser = argparse.ArgumentParser(
prog="response_io.py",
description="Persist large skill API responses to disk and read fields on demand.",
)
sub = parser.add_subparsers(dest="cmd", required=True)
p_run = sub.add_parser(
"run",
help="Execute a main script and persist its stdout to a file; "
"print only a lightweight preview to stdout.",
)
p_run.add_argument("params", help="JSON params string passed verbatim to the main script (argv[1]).")
p_run.add_argument("--script", required=True, help="Path to the main script to execute, e.g. scripts/my_api.py")
p_run.add_argument("--out-dir", required=True, help="Directory to write the response file into (created if missing).")
p_run.add_argument("--label", default=None, help="Optional filename suffix; sanitized to safe filename characters.")
p_run.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT_SEC, help=f"Subprocess timeout in seconds (default: {DEFAULT_TIMEOUT_SEC}).")
p_run.set_defaults(func=cmd_run)
p_read = sub.add_parser(
"read",
help="Extract specific fields from a previously persisted response file.",
)
p_read.add_argument("file", help="Path to the persisted JSON response file.")
g = p_read.add_mutually_exclusive_group()
g.add_argument("--path", default=None, help="JMESPath expression, e.g. 'data[*].{asin: asin, title: title}'.")
g.add_argument("--fields", default=None, help="Comma-separated field paths, e.g. 'data[*].asin,data[*].title'.")
p_read.add_argument("--limit", type=int, default=None, help="Take at most N items (when result is a list).")
p_read.add_argument("--offset", type=int, default=None, help="Skip the first M items (when result is a list).")
p_read.add_argument("--format", choices=["json", "jsonl", "csv", "table"], default="json", help="Output format (default: json).")
p_read.set_defaults(func=cmd_read)
args = parser.parse_args()
return args.func(args)
if __name__ == "__main__":
sys.exit(main())