
Linkfox Dld Product Billboard
- 252 installs
- 64 repo stars
- Updated August 3, 2026
- linkfox-ai/linkfox-skills
Monitor DLD product billboard rankings with LinkFox to track trending items, creative angles, and category movers for ongoing catalog decisions.
About
linkfox-dld-product-billboard skill accesses DLD product billboard data via LinkFox for e-commerce growth analytics. Sellers monitor trending products and leaderboard movement to refine ads, inventory, and merchandising after initial launch.
- Product billboard rankings
- Trending SKU visibility
- Category leaderboard views
- Creative trend signals
- Recurring performance checks
Linkfox Dld Product Billboard by the numbers
- 252 all-time installs (skills.sh)
- +38 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #266 of 853 Sales & Marketing skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/linkfox-ai/linkfox-skills --skill linkfox-dld-product-billboardAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 252 |
|---|---|
| repo stars | ★ 64 |
| Last updated | August 3, 2026 |
| Repository | linkfox-ai/linkfox-skills ↗ |
What it does
Monitor DLD product billboard rankings with LinkFox to track trending items, creative angles, and category movers for ongoing catalog decisions.
Files
DLD Product Billboard (1688 Bestseller Rankings)
This skill guides you on how to query 1688 platform product bestseller billboard data, helping sellers discover hot-selling wholesale products and sourcing opportunities on China's largest B2B marketplace.
Core Concepts
The DLD Product Billboard provides access to 1688 platform's product ranking data, covering both weekly and monthly bestseller lists. It enables users to discover trending wholesale products, compare suppliers, and identify sourcing opportunities in the domestic Chinese wholesale market.
Billboard types:
- Weekly Billboard (
pageType=2): Date parameter should be the Sunday of the target week (e.g.,2025-06-15). Data available for the last 90 days. - Monthly Billboard (
pageType=3): Date parameter should be the first day of the target month (e.g.,2025-06-01). Data available for the last 12 months.
Default behavior: Monthly billboard, sorted by order count descending, 20 results per page.
Parameter Guide
Core Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| keyWord | string | No | — | Product search keyword (must be in Chinese; translate first if not) |
| date | string | No | — | Query date. Weekly: Sunday date (e.g., 2025-06-15); Monthly: first of month (e.g., 2025-06-01) |
| pageType | integer | No | 3 | Billboard type: 2 = weekly, 3 = monthly |
| pageIndex | integer | No | 1 | Page number (starts from 1) |
| pageSize | integer | No | 20 | Results per page (10-100) |
Sorting
| Parameter | Type | Default | Description |
|---|---|---|---|
| sortField | string | orderCount | Sort field: orderCount (orders), saleCount (units sold), saleVolume (est. revenue), offerCreateTime (listing date), price (wholesale price), consignPrice (dropship price) |
| sortType | string | desc | Sort order: desc (descending), asc (ascending) |
Price & Volume Filters
| Parameter | Type | Description |
|---|---|---|
| beginPrice / endPrice | number | Wholesale price range |
| beginConsignPrice / endConsignPrice | number | Dropship price range |
| beginOrderCount / endOrderCount | integer | Order count range |
| beginSaleCount / endSaleCount | integer | Units sold range |
| beginSaleVolume / endSaleVolume | number | Estimated revenue range |
| beginStartQuantity / endStartQuantity | integer | Minimum order quantity range |
Product & Seller Filters
| Parameter | Type | Description |
|---|---|---|
| searchType | integer | Keyword match mode: 1 = fuzzy (default), 3 = exact |
| offerType | integer | Product tag: 0 = any, 2 = new product, 3 = 1688 Select, 4 = cross-border, 5 = customizable, 6 = store treasure |
| companyType | integer | Company type: 0 = any, 1 = store, 2 = factory |
| shiLiType | string | Seller tier (comma-separated): superFactory, Power, TrustPass |
| beginTpYear / endTpYear | integer | TrustPass membership years range |
Logistics & Service Filters
| Parameter | Type | Description |
|---|---|---|
| sendTime | string | Shipping time in hours (comma-separated): 24, 48, 72 |
| proxyRights | string | Dropship benefits (comma-separated): 4360897 (free shipping dropship), 449154 (buy-first-pay-later) |
| shopService | string | Seller services (comma-separated): 4057409 (worry-free purchase), 888777 (deep verification report) |
| buyerProtections | string | Buyer protections (comma-separated values in Chinese) |
| faceToFaceSupport | string | Shipping label support (comma-separated): 441218 (Taobao), 386434 (Douyin), 422914 (Pinduoduo), 422978 (Xiaohongshu), 386370 (Kuaishou) |
Other
| Parameter | Type | Description |
|---|---|---|
| productIds | string | Product IDs separated by Chinese comma, max 20 |
| goodsUrl | string | Direct product URL for lookup |
| beginOfferCreateTime / endOfferCreateTime | string | Listing date range (format: YYYY-MM-DD) |
Usage Examples
1. Monthly bestsellers for a keyword
"Show me the top-selling phone cases on 1688 this month"
{"keyWord": "手机壳", "pageType": 3, "date": "2026-03-01", "sortField": "orderCount", "sortType": "desc"}2. Weekly billboard sorted by revenue
"What products had the highest revenue last week in the yoga mat category?"
{"keyWord": "瑜伽垫", "pageType": 2, "date": "2026-03-22", "sortField": "saleVolume", "sortType": "desc"}3. Factory-direct products with price filter
"Find factory-direct earphone products on 1688 priced between 5 and 30 yuan"
{"keyWord": "耳机", "companyType": 2, "beginPrice": 5, "endPrice": 30, "sortField": "saleCount", "sortType": "desc"}4. Cross-border tagged products with fast shipping
"Show cross-border tagged LED light products that ship within 24 hours"
{"keyWord": "LED灯", "offerType": 4, "sendTime": "24", "sortField": "orderCount", "sortType": "desc"}5. New products from super factories
"Find newly listed products from super factories in the pet supplies category"
{"keyWord": "宠物用品", "offerType": 2, "shiLiType": "superFactory", "sortField": "offerCreateTime", "sortType": "desc"}6. Dropship-friendly products with buyer protections
"Show me dropship-friendly bag products with free shipping and return support"
{"keyWord": "包包", "proxyRights": "4360897", "buyerProtections": "商品包邮,7天包退货", "sortField": "orderCount", "sortType": "desc"}7. High-volume products in a price range
"Find products with more than 1000 orders and wholesale price under 50 yuan in the toy category"
{"keyWord": "玩具", "beginOrderCount": 1000, "endPrice": 50, "sortField": "orderCount", "sortType": "desc"}8. Browse by product IDs
"Look up these specific 1688 product IDs: 123456、789012"
{"productIds": "123456、789012"}Display Rules
1. Present data clearly: Show product results in well-organized tables including product title, wholesale price, dropship price, order count, units sold, estimated revenue, supplier name, and listing date 2. Image display: When imageUrl is available, display product images to help users visually identify products 3. Link provision: Include product links (asinUrl) and shop links (shopUrl) so users can navigate directly to the 1688 listing 4. Price formatting: Always show prices with the currency (CNY/RMB) and clarify whether the price is wholesale or dropship 5. Volume context: When presenting sales data, clearly label whether it is weekly or monthly data based on the dataType field 6. Pagination guidance: When total results exceed the current page, inform the user of the total count and offer to fetch more pages 7. Keyword translation: If the user provides a keyword in a non-Chinese language, translate it to Chinese before querying, and inform the user of the translated keyword 8. Error handling: When a query fails, explain the issue and suggest adjusting parameters (e.g., broadening filters, checking date format)
Important Limitations
- Keyword language: The
keyWordparameter must be in Chinese. Always translate non-Chinese keywords before querying. - Date format matters: Weekly billboard dates must be a Sunday; monthly billboard dates must be the 1st of the month.
- Weekly data range: Last 90 days only.
- Monthly data range: Last 12 months only.
- Page size cap: Maximum 100 results per request.
- Product ID limit: Maximum 20 product IDs per lookup.
User Expression & Scenario Quick Reference
Applicable -- 1688 wholesale product discovery and sourcing:
| User Says | Scenario |
|---|---|
| "What's hot on 1688", "1688 trending products" | Bestseller discovery |
| "Find cheap suppliers for XX", "wholesale source for XX" | Sourcing by keyword |
| "Factory-direct products", "OEM suppliers" | Factory filtering |
| "Cross-border sourcing", "products for export" | Cross-border tag filtering |
| "Dropshipping products on 1688" | Dropship-enabled product search |
| "New products on 1688", "recently listed items" | New product discovery |
| "Compare suppliers for XX" | Multi-result comparison |
| "1688 product ranking", "bestseller list" | Billboard browsing |
Not applicable -- Needs beyond 1688 product billboard data:
- Amazon product research or keyword analysis (use ABA tools instead)
- 1688 store/shop-level analytics (shop rankings, shop scores)
- Alibaba.com (international) data
- Price negotiation or order placement
- Product quality reviews or certifications
- Logistics cost calculation or freight forwarding
Feedback:
Auto-detect and report feedback via the Feedback API when any of the following apply: 1. The functionality or purpose described in this skill does not match actual behavior 2. The skill's results do not match the user's intent 3. The user expresses dissatisfaction or praise about this skill 4. Anything you believe could be improved
Call the feedback API as specified in references/api.md. Do not interrupt the user's flow.
<!-- LF_LARGE_RESPONSE_BLOCK -->
Handling Large Responses
To avoid overflowing the agent context, persist the response to disk and extract only the fields you need:
python scripts/response_io.py run --script scripts/dld_product_billboard.py --out-dir <DIR> '<params>'
python scripts/response_io.py read <file> --fields "<paths>" # or --path "<JMESPath>"Pick--out-diroutside any git working tree (e.g./tmp/...on Unix,%TEMP%/...on Windows). Persisted responses may contain PII, pricing, or auth-sensitive data — do not commit them. Files are not auto-deleted; clean up when the task is done.
run writes the full response to a file and emits only a schema preview + file path. read projects specific fields, with --limit/--offset for slicing and --format json|jsonl|csv|table for output.
When to prefer this pattern — apply your judgment based on the response characteristics, e.g.:
- High field count per record, or fields you don't need
- Batch/paginated results (multiple items per call)
- Long-text fields (descriptions, reviews, HTML, time series)
- Output reused across later steps rather than consumed immediately
For small, single-use responses, calling the main script directly is fine.
⚠️ The preview is a truncated schema + sample, not the full data. Any field-level decision must read from the persisted file via read. <!-- /LF_LARGE_RESPONSE_BLOCK -->
--- For more high-quality, professional cross-border e-commerce skills, set [LinkFox Skills](https://skill.linkfox.com/).
店雷达-1688商品榜单 API 参考
调用规范
- 请求地址:
https://tool-gateway.linkfox.com/dld/productBillboard - 请求方式:POST,Content-Type: application/json
- 认证方式:Header
Authorization: <api_key>,api_key 从环境变量LINKFOXAGENT_API_KEY读取(如未配置,提示用户前往 https://skill.linkfox.com/linkfoxskills/guide.htm 申请)
请求参数
POST Body(JSON):
| 参数 | 类型 | 必填 | 说明 |
|---|---|---|---|
| keyWord | string | 否 | 商品搜索关键字(搜索关键词必须是中文,如果不是请先翻译),最大长度50 |
| date | string | 否 | 查询时间。周榜:传入该周的周天日期,如 2025-06-15(最长近90天);月榜:传入该月第一天,如 2025-06-01(最长近一年) |
| pageType | integer | 否 | 榜单类型:2 = 周榜,3 = 月榜。默认 3 |
| pageIndex | integer | 否 | 页码(从1开始),默认 1 |
| pageSize | integer | 否 | 每页返回数量(10-100),默认 20 |
| sortField | string | 否 | 排序字段,默认 orderCount。可选值:orderCount(销售笔数)、saleCount(销售件数)、saleVolume(预估销售额)、offerCreateTime(上架时间)、price(批发价)、consignPrice(代发价) |
| sortType | string | 否 | 排序类型:desc(降序)、asc(升序),默认 desc |
| searchType | integer | 否 | 商品关键词搜索类型:1 = 模糊匹配,3 = 精准匹配。默认 1 |
| offerType | integer | 否 | 商品标识:0 = 不限制,2 = 新品,3 = 1688严选,4 = 跨境,5 = 支持定制,6 = 镇店之宝。默认 0 |
| companyType | integer | 否 | 公司类型:0 = 不限,1 = 店铺,2 = 工厂 |
| shiLiType | string | 否 | 卖家会员类型(多选),多个使用","号隔开。可选值:superFactory(超级工厂)、Power(实力商家)、TrustPass(仅诚信通会员) |
| beginTpYear | integer | 否 | 开始诚信通年限 |
| endTpYear | integer | 否 | 结束诚信通年限 |
| beginPrice | number | 否 | 批发价(起始) |
| endPrice | number | 否 | 批发价(结束) |
| beginConsignPrice | number | 否 | 代发价(起始) |
| endConsignPrice | number | 否 | 代发价(结束) |
| beginOrderCount | integer | 否 | 销售笔数(起始) |
| endOrderCount | integer | 否 | 销售笔数(结束) |
| beginSaleCount | integer | 否 | 销售件数(起始) |
| endSaleCount | integer | 否 | 销售件数(结束) |
| beginSaleVolume | number | 否 | 销售额(起始) |
| endSaleVolume | number | 否 | 销售额(结束) |
| beginStartQuantity | integer | 否 | 起始起批量 |
| endStartQuantity | integer | 否 | 结束起批量 |
| beginOfferCreateTime | string | 否 | 上架时间(起始),格式:YYYY-MM-DD |
| endOfferCreateTime | string | 否 | 上架时间(结束),格式:YYYY-MM-DD |
| sendTime | string | 否 | 发货时间(多选),多个使用","号隔开。可选值:24(24小时)、48(48小时)、72(72小时) |
| proxyRights | string | 否 | 代发权益(多选),多个使用","号隔开。可选值:4360897(一件代发包邮)、449154(先采后付) |
| shopService | string | 否 | 卖家服务(多选),多个使用","号隔开。可选值:4057409(安心购)、888777(深度认证报告) |
| buyerProtections | string | 否 | 权益保障(多选),多个用","隔开。可选值:商品包邮、7天包退货、支持运费险 |
| faceToFaceSupport | string | 否 | 面单支持(多选),多个使用","号隔开。可选值:441218(淘宝)、386434(抖音)、422914(拼多多)、422978(小红书)、386370(快手) |
| productIds | string | 否 | 商品ID,顿号隔开搜索多个,最多20个 |
| goodsUrl | string | 否 | 商品链接地址 |
响应结构
| 字段 | 类型 | 说明 |
|---|---|---|
| total | integer | 记录数 |
| type | string | 渲染的样式 |
| columns | array | 渲染的列 |
| products | array | 商品列表(见下方商品对象) |
商品对象
| 字段 | 类型 | 说明 |
|---|---|---|
| offerId | string | 商品id |
| asin | string | 商品编号 |
| title | string | 商品标题 |
| price | number | 批发价 |
| consignPrice | number | 代发价 |
| currency | string | 币种 |
| unit | string | 单位 |
| quantityBegin | integer | 起批量 |
| quantityPrices | string | 价格区间 |
| salesOrderCount | integer | 销售笔数(按统计周期返回对应的值) |
| salesQuantity | integer | 销售件数(按统计周期返回对应的值) |
| estimatedSalesAmount | integer | 预估销售额(按统计周期返回对应的值) |
| dataType | string | 数据类型:weeklyData = 周数据,monthlyData = 月数据 |
| availableDate | string | 商品上架时间,格式为 yyyy-MM-dd HH:mm:ss |
| deliveryTime | string | 发货时间 |
| levelName | string | 类目层级名称 |
| company | string | 店铺名称 |
| shopId | string | 店铺id |
| shopUrl | string | 店铺链接地址 |
| asinUrl | string | 商品链接地址 |
| imageUrl | string | 图片地址 |
| sourceType | string | 来源平台(1688) |
| sourceTool | string | 来源工具 |
错误码
正常情况下,接口的 HTTP 状态码均为 200,业务的成功与否通过响应体中的 errorCode 字段区分(errorCode = 200 表示成功,其他值表示业务错误)。当遇到未授权等情况时,HTTP 状态码为 401,且对应的 errorCode 也是 401。
| errcode | 含义 | 处理建议 |
|---|---|---|
| 200 | 成功 | 正常解析业务字段 |
| 401 | 认证失败 | 检查请求头 Authorization 是否正确携带 API Key;API Key 申请方式请参考上述调用规范下的认证方式。 |
| 其他非200值 | 业务异常 | 参考 errmsg 字段获取具体错误原因 |
错误响应示例:
{
"errcode": 401,
"errmsg": "authorized error"
}curl 示例
curl -X POST https://tool-gateway.linkfox.com/dld/productBillboard \
-H "Authorization: $LINKFOXAGENT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"keyWord": "手机壳",
"pageType": 3,
"date": "2026-03-01",
"sortField": "orderCount",
"sortType": "desc",
"pageSize": 20,
"pageIndex": 1
}'---
Feedback API
This endpoint is separate from the tool API above. Do not mix the two base URLs.
- POST
https://skill-api.linkfox.com/api/v1/public/feedback - Content-Type:
application/json
{
"skillName": "linkfox-xxx-xxx",
"sentiment": "POSITIVE",
"category": "OTHER",
"content": "Results were accurate, user was satisfied."
}Field rules:
skillName: Use this skill'snamefrom the YAML frontmattersentiment: Choose ONE —POSITIVE(praise),NEUTRAL(suggestion without emotion),NEGATIVE(complaint or error)category: Choose ONE —BUG(malfunction or wrong data),COMPLAINT(user dissatisfaction),SUGGESTION(improvement idea),OTHERcontent: Include what the user said or intended, what actually happened, and why it is a problem or praise
#!/usr/bin/env python3
"""
DLD Product Billboard - LinkFox Skill
Calls the dld/productBillboard API endpoint to query 1688 bestseller rankings.
Usage:
python dld_product_billboard.py '<JSON parameters>'
Examples:
# Monthly bestsellers for phone cases
python dld_product_billboard.py '{"keyWord": "手机壳", "pageType": 3, "date": "2026-03-01"}'
# Weekly billboard sorted by revenue
python dld_product_billboard.py '{"keyWord": "瑜伽垫", "pageType": 2, "date": "2026-03-22", "sortField": "saleVolume"}'
# Factory-direct products with price filter
python dld_product_billboard.py '{"keyWord": "耳机", "companyType": 2, "beginPrice": 5, "endPrice": 30}'
"""
import json
import os
import sys
from urllib.request import urlopen, Request
from urllib.error import HTTPError, URLError
API_URL = "https://tool-gateway.linkfox.com/dld/productBillboard"
def get_api_key():
"""Retrieve the API key from environment, with a friendly prompt if missing."""
key = os.environ.get("LINKFOXAGENT_API_KEY")
if not key:
print(
"API Key not configured. Please complete authorization first:\n"
"1. Visit https://skill.linkfox.com/linkfoxskills/guide.htm to obtain your Key\n"
"2. Set the environment variable: export LINKFOXAGENT_API_KEY=your-key-here",
file=sys.stderr,
)
sys.exit(1)
return key
def call_api(params: dict) -> dict:
"""Send a POST request to the DLD product billboard API."""
api_key = get_api_key()
data = json.dumps(params).encode("utf-8")
req = Request(
API_URL,
data=data,
headers={
"Authorization": api_key,
"Content-Type": "application/json",
"User-Agent": "LinkFox-Skill/1.0",
},
method="POST",
)
try:
with urlopen(req, timeout=60) as response:
return json.loads(response.read().decode("utf-8"))
except HTTPError as e:
body = e.read().decode("utf-8") if e.fp else ""
return {"error": f"HTTP {e.code}: {e.reason}", "details": body}
except URLError as e:
return {"error": f"Connection failed: {e.reason}"}
def format_product_summary(product: dict) -> str:
"""Format a single product record into a readable summary line."""
title = product.get("title", "N/A")
price = product.get("price", "N/A")
consign_price = product.get("consignPrice", "N/A")
orders = product.get("salesOrderCount", "N/A")
sales_qty = product.get("salesQuantity", "N/A")
revenue = product.get("estimatedSalesAmount", "N/A")
company = product.get("company", "N/A")
url = product.get("asinUrl", "")
return (
f" Title: {title}\n"
f" Wholesale Price: {price} | Dropship Price: {consign_price}\n"
f" Orders: {orders} | Units Sold: {sales_qty} | Est. Revenue: {revenue}\n"
f" Supplier: {company}\n"
f" URL: {url}"
)
def print_results(result: dict):
"""Print API results in a human-readable format."""
# Check for errors
if "error" in result:
print(f"Error: {result['error']}", file=sys.stderr)
if "details" in result:
print(f"Details: {result['details']}", file=sys.stderr)
return
total = result.get("total", 0)
products = result.get("products", [])
print(f"Total records: {total}")
print(f"Returned: {len(products)} products")
print("-" * 60)
for i, product in enumerate(products, 1):
print(f"\n[{i}]")
print(format_product_summary(product))
if not products:
print("No products found. Try broadening your filters or changing the keyword.")
def main():
if len(sys.argv) < 2:
print("Usage: dld_product_billboard.py '<JSON parameters>'", file=sys.stderr)
print(
"\nExamples:",
file=sys.stderr,
)
print(
' dld_product_billboard.py \'{"keyWord": "手机壳", "pageType": 3, "date": "2026-03-01"}\'',
file=sys.stderr,
)
print(
' dld_product_billboard.py \'{"keyWord": "耳机", "companyType": 2, "beginPrice": 5, "endPrice": 30}\'',
file=sys.stderr,
)
sys.exit(1)
# Parse input JSON
try:
params = json.loads(sys.argv[1])
except json.JSONDecodeError as e:
print(f"Invalid parameter format: {e}", file=sys.stderr)
sys.exit(1)
# Call the API
result = call_api(params)
# Print both raw JSON and a formatted summary
print("=== Raw JSON Response ===")
print(json.dumps(result, indent=2, ensure_ascii=False))
print("\n=== Formatted Summary ===")
print_results(result)
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
Skill response I/O helper — wraps any main script to persist large API
responses to disk, then offers a `read` subcommand to extract specific fields
from those persisted files. Generic, business-agnostic.
This script is bundled into each skill's scripts/ directory by tools/response_io/sync.py.
The agent must pass --script <path> to identify which main script to execute.
Usage:
python scripts/response_io.py run --script <PATH> --out-dir <DIR> '<json_params>' [--label NAME] [--timeout SEC]
python scripts/response_io.py read <file> (--path "<JMESPath>" | --fields "f1,f2,...") [--limit N] [--offset M] [--format json|jsonl|csv|table]
"""
from __future__ import annotations
import sys
if sys.version_info < (3, 10):
sys.exit(
"Error: Python 3.10+ required (current: "
f"{sys.version_info.major}.{sys.version_info.minor}). "
"Please upgrade Python."
)
import argparse
import csv
import io
import json
import os
import re
import secrets
import subprocess
from datetime import datetime
from pathlib import Path
from typing import Any
# Force UTF-8 stdout/stderr so non-ASCII chars in previews and API responses
# print correctly on Windows (default cp936 / gbk).
for stream in (sys.stdout, sys.stderr):
try:
stream.reconfigure(encoding="utf-8") # type: ignore[attr-defined]
except (AttributeError, OSError):
pass
try:
import jmespath # type: ignore
HAS_JMESPATH = True
except ImportError:
HAS_JMESPATH = False
MAX_STRING_LEN = 120
MAX_DEPTH = 3
SAMPLE_KEY_CAP = 15
RAW_TEXT_PEEK = 500
DEFAULT_TIMEOUT_SEC = 300
# ---------------------------------------------------------------------------
# Shared helpers
# ---------------------------------------------------------------------------
def _err(msg: str, code: int = 1) -> None:
print(msg, file=sys.stderr)
sys.exit(code)
def _resolve_script(script_arg: str) -> Path:
p = Path(script_arg).expanduser()
if not p.is_absolute():
# Resolve relative to the current working directory the agent invoked from.
p = (Path.cwd() / p).resolve()
else:
p = p.resolve()
if not p.is_file():
_err(f"--script path not found: {p}")
return p
def _resolve_skill_name(main_script: Path) -> str:
"""Best-effort skill name extraction for filename prefixing.
main_script lives at <skill_dir>/scripts/<name>.py — return <skill_dir>'s
folder name. Fall back to the script's stem if structure differs.
"""
try:
if main_script.parent.name == "scripts":
return main_script.parents[1].name
except IndexError:
pass
return main_script.stem
def _sanitize_label(label: str) -> str:
"""Allow only safe filename chars in --label to prevent path traversal."""
cleaned = re.sub(r"[^\w\-]", "_", label)
return cleaned[:64] # cap length
def _truncate_string(s: str) -> str:
if len(s) <= MAX_STRING_LEN:
return s
return s[:MAX_STRING_LEN] + f"...(truncated, total {len(s)} chars)"
def _truncate_value(value: Any, depth: int = 0) -> Any:
"""Recursively truncate strings, deep nesting, and large arrays for preview."""
if depth >= MAX_DEPTH:
if isinstance(value, dict):
return f"<truncated nested object, keys: {list(value.keys())[:10]}>"
if isinstance(value, list):
return f"<truncated nested array, length: {len(value)}>"
if isinstance(value, str):
return _truncate_string(value)
return value
if isinstance(value, str):
return _truncate_string(value)
if isinstance(value, dict):
out = {k: _truncate_value(v, depth + 1) for k, v in value.items()}
return out
if isinstance(value, list):
if not value:
return []
truncated = [_truncate_value(value[0], depth + 1)]
if len(value) > 1:
# Note total length on the parent — keep the array type-homogeneous
# so downstream consumers can iterate without special-casing strings.
truncated.append({"_omitted_items": len(value) - 1})
return truncated
return value
def _shape_of(value: Any, top: bool = False) -> Any:
"""Lightweight schema description for the preview block."""
if isinstance(value, dict):
keys = list(value.keys())
out: dict[str, Any] = {"type": "object", "top_keys" if top else "keys": keys}
if top:
for k in keys[:8]:
out[k] = _shape_of(value[k])
return out
if isinstance(value, list):
out = {"type": "array", "length": len(value)}
if value and isinstance(value[0], dict):
out["item_keys"] = list(value[0].keys())
elif value:
out["item_type"] = type(value[0]).__name__
return out
return {"type": type(value).__name__}
def _build_sample(value: Any) -> Any:
"""First-record sample with explicit truncation marker."""
if isinstance(value, list):
if not value:
return {"_truncated_record": True, "_note": "array is empty"}
first = value[0]
if isinstance(first, dict):
sample = {"_truncated_record": True, "_note": f"first of {len(value)} items"}
sample.update(_truncate_value(first, depth=1))
return sample
return {"_truncated_record": True, "_note": f"first of {len(value)} items", "value": _truncate_value(first, depth=1)}
if isinstance(value, dict):
sample = {"_truncated_record": True, "_note": "top-level object (truncated)"}
sample.update(_truncate_value(value, depth=1))
return sample
return {"_truncated_record": True, "value": _truncate_value(value, depth=1)}
def _shrink_preview(preview: dict) -> dict:
"""Cap the sample's value fields when it has many keys.
`shape.*.item_keys` is the single source of truth for the full key list
(always complete, no truncation). The sample only ever shows up to
SAMPLE_KEY_CAP fields with their concrete values, since the agent only
needs a feel for value shapes — for the full menu of available fields,
they read `shape`.
"""
sample = preview.get("sample")
if isinstance(sample, dict):
meta_keys = {"_truncated_record", "_note"}
data_keys = [k for k in sample.keys() if k not in meta_keys]
if len(data_keys) > SAMPLE_KEY_CAP:
kept = data_keys[:SAMPLE_KEY_CAP]
new_sample = {k: v for k, v in sample.items() if k in meta_keys or k in kept}
base_note = sample.get("_note", "")
extra = (
f"showing first {SAMPLE_KEY_CAP} of {len(data_keys)} fields "
f"(see `shape` for the complete key list)"
)
new_sample["_note"] = f"{base_note}; {extra}" if base_note else extra
preview["sample"] = new_sample
return preview
# ---------------------------------------------------------------------------
# `run` subcommand
# ---------------------------------------------------------------------------
def cmd_run(args: argparse.Namespace) -> int:
main_script = _resolve_script(args.script)
skill_name = _resolve_skill_name(main_script)
out_dir = Path(args.out_dir).expanduser().resolve()
try:
out_dir.mkdir(parents=True, exist_ok=True)
except OSError as e:
_err(f"Failed to create --out-dir {out_dir}: {e}")
if not os.access(out_dir, os.W_OK):
_err(f"--out-dir is not writable: {out_dir}")
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
rand = secrets.token_hex(3)
safe_label = _sanitize_label(args.label) if args.label else ""
label_part = f"__{safe_label}" if safe_label else ""
out_file = out_dir / f"{skill_name}__{timestamp}_{rand}{label_part}.json"
# Force the child process to emit UTF-8 regardless of the host console
# encoding (Windows defaults to cp936 / gbk and would otherwise corrupt
# non-ASCII bytes when we read them back).
child_env = os.environ.copy()
child_env["PYTHONIOENCODING"] = "utf-8"
timed_out = False
try:
proc = subprocess.run(
[sys.executable, str(main_script), args.params],
capture_output=True,
text=True,
encoding="utf-8",
errors="replace",
env=child_env,
timeout=args.timeout,
)
stdout_text = proc.stdout or ""
stderr_text = proc.stderr or ""
returncode = proc.returncode
except subprocess.TimeoutExpired as e:
timed_out = True
stdout_text = (e.stdout.decode("utf-8", errors="replace") if isinstance(e.stdout, bytes) else (e.stdout or "")) or ""
stderr_text = (e.stderr.decode("utf-8", errors="replace") if isinstance(e.stderr, bytes) else (e.stderr or "")) or ""
returncode = 124 # convention for timeout
# Always write the captured stdout to disk, even if not JSON.
try:
out_file.write_text(stdout_text, encoding="utf-8")
except OSError as e:
_err(f"Failed to write output file {out_file}: {e}")
if stderr_text:
sys.stderr.write(stderr_text)
# Try to parse the captured stdout as JSON for the preview.
try:
parsed = json.loads(stdout_text) if stdout_text.strip() else None
format_kind = "json"
except json.JSONDecodeError:
parsed = None
format_kind = "raw_text"
preview: dict[str, Any] = {
"_preview": {
"is_preview": True,
"warning": (
"PREVIEW ONLY — NOT FULL DATA. The full response is saved to `file`. "
"Use `python scripts/response_io.py read <file> --fields '...'` to extract "
"specific fields, or `--path '<JMESPath>'` for complex projections."
),
},
}
# Surface failures prominently so agents don't mistake a stub preview for success.
if returncode != 0 or timed_out:
stderr_snippet = stderr_text[-500:] if stderr_text else ""
preview["_error"] = {
"exit_code": returncode,
"timed_out": timed_out,
"stderr_snippet": stderr_snippet,
"hint": "The wrapped script failed or timed out. The output file may be empty or partial.",
}
preview.update({
"file": str(out_file),
"size_bytes": out_file.stat().st_size,
"skill": skill_name,
"exit_code": returncode,
"format": format_kind,
"label": safe_label or None,
"next_steps_hint": (
"use: python scripts/response_io.py read <file> --fields '...' | --path '...'"
),
})
if format_kind == "json":
preview["shape"] = _shape_of(parsed, top=True)
preview["sample"] = _build_sample(parsed)
else:
peek = stdout_text[:RAW_TEXT_PEEK]
preview["raw_text_peek"] = peek
preview["raw_text_total_chars"] = len(stdout_text)
preview["sample"] = {
"_truncated_record": True,
"_note": f"stdout was not valid JSON; first {RAW_TEXT_PEEK} chars shown above in raw_text_peek",
}
preview = _shrink_preview(preview)
print(json.dumps(preview, ensure_ascii=False, indent=2))
return returncode
# ---------------------------------------------------------------------------
# `read` subcommand
# ---------------------------------------------------------------------------
def _load_json(path: Path) -> Any:
try:
text = path.read_text(encoding="utf-8")
except OSError as e:
_err(f"Failed to read file {path}: {e}")
try:
return json.loads(text)
except json.JSONDecodeError as e:
_err(f"File is not valid JSON: {path}\n{e}")
def _basic_dot_path(data: Any, path: str) -> Any:
"""Pure-stdlib dot-path resolver. No [*] support — callers fall back here only when jmespath is unavailable AND the path has no [*]."""
cur = data
for part in path.split("."):
if isinstance(cur, dict):
cur = cur.get(part)
else:
return None
return cur
def _resolve_field(data: Any, expr: str) -> Any:
if HAS_JMESPATH:
return jmespath.search(expr, data)
if "[" in expr or "*" in expr:
_err(
f"jmespath is required for expression '{expr}'. "
f"Install with: pip install jmespath"
)
return _basic_dot_path(data, expr)
def _project_fields(data: Any, fields: list[str]) -> Any:
"""Run each field expr; if any returns a list, zip them into list-of-dicts."""
resolved: dict[str, Any] = {f: _resolve_field(data, f) for f in fields}
list_lengths = [len(v) for v in resolved.values() if isinstance(v, list)]
if not list_lengths:
return resolved
# All list values must be same length to zip cleanly.
if len(set(list_lengths)) > 1:
# Fallback: return the dict as-is so caller can inspect mismatches.
return resolved
n = list_lengths[0]
rows = []
for i in range(n):
row = {}
for f, v in resolved.items():
row[f] = v[i] if isinstance(v, list) else v
rows.append(row)
return rows
def _apply_slice(value: Any, limit: int | None, offset: int | None) -> Any:
if not isinstance(value, list):
return value
start = offset or 0
end = (start + limit) if limit is not None else None
return value[start:end]
def _format_output(value: Any, fmt: str) -> str:
if fmt == "json":
return json.dumps(value, ensure_ascii=False, indent=2)
if fmt == "jsonl":
if isinstance(value, list):
return "\n".join(json.dumps(item, ensure_ascii=False) for item in value)
return json.dumps(value, ensure_ascii=False)
if fmt in ("csv", "table"):
if not isinstance(value, list) or not value:
_err(f"--format {fmt} requires a non-empty list result")
if not all(isinstance(item, dict) for item in value):
_err(f"--format {fmt} requires list-of-objects, got list of {type(value[0]).__name__}")
keys: list[str] = []
for item in value:
for k in item.keys():
if k not in keys:
keys.append(k)
if fmt == "csv":
buf = io.StringIO()
writer = csv.DictWriter(buf, fieldnames=keys, extrasaction="ignore")
writer.writeheader()
for item in value:
writer.writerow({k: _stringify(item.get(k)) for k in keys})
return buf.getvalue().rstrip("\n")
# table: simple aligned columns
rows = [[_stringify(item.get(k)) for k in keys] for item in value]
widths = [len(k) for k in keys]
for row in rows:
for i, cell in enumerate(row):
widths[i] = max(widths[i], len(cell))
lines = [
" ".join(k.ljust(widths[i]) for i, k in enumerate(keys)),
" ".join("-" * widths[i] for i in range(len(keys))),
]
for row in rows:
lines.append(" ".join(row[i].ljust(widths[i]) for i in range(len(keys))))
return "\n".join(lines)
_err(f"Unknown --format: {fmt}")
return "" # unreachable
def _stringify(v: Any) -> str:
if v is None:
return ""
if isinstance(v, (dict, list)):
return json.dumps(v, ensure_ascii=False)
return str(v)
def cmd_read(args: argparse.Namespace) -> int:
if not args.path and not args.fields:
_err("read: either --path or --fields is required")
if args.path and args.fields:
_err("read: --path and --fields are mutually exclusive")
file_path = Path(args.file).expanduser().resolve()
data = _load_json(file_path)
if args.path:
result = _resolve_field(data, args.path)
else:
fields = [f.strip() for f in args.fields.split(",") if f.strip()]
if not fields:
_err("--fields parsed to empty list")
result = _project_fields(data, fields)
result = _apply_slice(result, args.limit, args.offset)
print(_format_output(result, args.format))
return 0
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def main() -> int:
parser = argparse.ArgumentParser(
prog="response_io.py",
description="Persist large skill API responses to disk and read fields on demand.",
)
sub = parser.add_subparsers(dest="cmd", required=True)
p_run = sub.add_parser(
"run",
help="Execute a main script and persist its stdout to a file; "
"print only a lightweight preview to stdout.",
)
p_run.add_argument("params", help="JSON params string passed verbatim to the main script (argv[1]).")
p_run.add_argument("--script", required=True, help="Path to the main script to execute, e.g. scripts/my_api.py")
p_run.add_argument("--out-dir", required=True, help="Directory to write the response file into (created if missing).")
p_run.add_argument("--label", default=None, help="Optional filename suffix; sanitized to safe filename characters.")
p_run.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT_SEC, help=f"Subprocess timeout in seconds (default: {DEFAULT_TIMEOUT_SEC}).")
p_run.set_defaults(func=cmd_run)
p_read = sub.add_parser(
"read",
help="Extract specific fields from a previously persisted response file.",
)
p_read.add_argument("file", help="Path to the persisted JSON response file.")
g = p_read.add_mutually_exclusive_group()
g.add_argument("--path", default=None, help="JMESPath expression, e.g. 'data[*].{asin: asin, title: title}'.")
g.add_argument("--fields", default=None, help="Comma-separated field paths, e.g. 'data[*].asin,data[*].title'.")
p_read.add_argument("--limit", type=int, default=None, help="Take at most N items (when result is a list).")
p_read.add_argument("--offset", type=int, default=None, help="Skip the first M items (when result is a list).")
p_read.add_argument("--format", choices=["json", "jsonl", "csv", "table"], default="json", help="Output format (default: json).")
p_read.set_defaults(func=cmd_read)
args = parser.parse_args()
return args.func(args)
if __name__ == "__main__":
sys.exit(main())