
Linkfox Dld Product Search
- 275 installs
- 64 repo stars
- Updated August 3, 2026
- linkfox-ai/linkfox-skills
Search the DLD marketplace catalog to find products, compare listings, and collect early sourcing signals before choosing SKUs or suppliers for an e-commerce or marketplace launch.
About
Linkfox skill for searching products on the DLD marketplace, returning listing results agents can use to discover items, compare offers, and inform early e-commerce sourcing and niche-selection decisions.
- DLD marketplace product search
- Catalog and listing discovery
- E-commerce sourcing research
- Agent-callable marketplace lookup
- Early SKU opportunity scanning
Linkfox Dld Product Search by the numbers
- 275 all-time installs (skills.sh)
- +41 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #508 of 2,715 Automation & Workflows skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/linkfox-ai/linkfox-skills --skill linkfox-dld-product-searchAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 275 |
|---|---|
| repo stars | ★ 64 |
| Last updated | August 3, 2026 |
| Repository | linkfox-ai/linkfox-skills ↗ |
What it does
Search the DLD marketplace catalog to find products, compare listings, and collect early sourcing signals before choosing SKUs or suppliers for an e-commerce or marketplace launch.
Files
1688 Product Search (DianLeiDa)
This skill guides you on how to search and analyze products on the 1688 wholesale platform, helping e-commerce sellers and sourcing professionals find quality suppliers and profitable products.
Core Concepts
This tool provides keyword-based product search on 1688 (China's largest B2B wholesale marketplace). It aggregates product listings with sales data, pricing tiers, supplier credentials, and fulfillment options. Data is sourced from DianLeiDa (store radar analytics) and covers real-time product listings with 7-day and 30-day sales metrics.
Key terminology:
- Wholesale price (
price): The price per unit when ordering at the minimum batch quantity - Dropship price (
consignPrice): The price per unit for single-item dropshipping (typically higher than wholesale) - TrustPass years (
tpYear): The number of years the supplier has held Alibaba's TrustPass membership, indicating business longevity - Sales count (
salesQuantity): Total units sold in the selected time period - Order count (
salesOrderCount): Total number of separate orders in the selected time period - Estimated sales volume (
estimatedSalesAmount): Estimated revenue in the selected time period
Data Fields
| Field | API Name | Description | Example |
|---|---|---|---|
| Product Title | title | Full product listing title | ... |
| Product ID | offerId | Unique 1688 product identifier | 805578065498 |
| Product URL | asinUrl | Direct link to the product listing | https://detail.1688.com/... |
| Image URL | imageUrl | Product main image | https://cbu01.alicdn.com/... |
| Wholesale Price | price | Unit price at batch quantity (CNY) | 12.50 |
| Dropship Price | consignPrice | Unit price for single-piece dropship (CNY) | 18.90 |
| Price Tiers | quantityPrices | Volume-based pricing breakdown | ... |
| Min Order Qty | quantityBegin | Minimum order quantity | 2 |
| Sales Order Count | salesOrderCount | Number of orders in the period | 350 |
| Sales Quantity | salesQuantity | Units sold in the period | 1200 |
| Est. Sales Amount | estimatedSalesAmount | Estimated revenue in the period | 45000 |
| Delivery Time | deliveryTime | Promised shipping time | 24h |
| Listing Date | availableDate | When the product was first listed | 2025-03-15 |
| Category | levelName | Product category path | ... |
| Shop Name | company | Supplier/store name | ... |
| Shop ID | shopId | Unique store identifier | ... |
| Shop URL | shopUrl | Link to the supplier's store | https://shop... |
| Currency | currency | Price currency (always CNY) | CNY |
| Data Type | dataType | Period indicator: weeklyData or monthlyData | monthlyData |
Parameter Guide
Search Keyword
The most important parameter. Keywords must be in Chinese. If the user provides an English term, translate it to Chinese before querying.
keyWord(string, max 50 chars): The Chinese search termsearchType: 1 = fuzzy match (default), 3 = exact matchgoodsUrl: Search by product URL instead of keywordproductIds: Search by specific product IDs (comma-separated, max 20)
Time Period
cycle:"7"for last 7 days,"30"for last 30 days
Sorting
sortField: Field to sort by. Default:orderCount30dorderCount7d/orderCount30d-- order countsaleCount7d/saleCount30d-- units soldsaleVolume7d/saleVolume30d-- estimated revenueofferCreateTime-- listing dateprice-- wholesale priceconsignPrice-- dropship pricesortType:"desc"(default) or"asc"
Price Filters
beginPrice/endPrice: Wholesale price range (CNY)beginConsignPrice/endConsignPrice: Dropship price range (CNY)
Sales Filters
beginOrderCount/endOrderCount: Order count rangebeginSaleCount/endSaleCount: Units sold rangebeginSaleVolume/endSaleVolume: Revenue range (CNY)
Supplier Filters
companyType: 0 = any (default), 1 = store, 2 = factoryshiLiType: Seller tier. Comma-separated multi-select:superFactory-- Super FactoryPower-- Power MerchantTrustPass-- TrustPass members onlybeginTpYear/endTpYear: TrustPass membership year range
Product Tags
offerType: 0 = any, 2 = new product, 3 = 1688 Select, 4 = cross-border, 5 = customizable, 6 = top store pick
Fulfillment & Services
sendTime: Shipping speed. Comma-separated:"24","48","72"faceToFaceSupport: Platform face-sheet support. Comma-separated:441218(Taobao),386434(Douyin),422914(Pinduoduo),422978(Xiaohongshu),386370(Kuaishou)proxyRights: Dropship benefits. Comma-separated:4360897(free shipping dropship),449154(buy now pay later)shopService: Seller services. Comma-separated:4057409(worry-free purchase),888777(deep verification report)buyerProtections: Buyer guarantees. Comma-separated Chinese strings:商品包邮(free shipping),7天包退货(7-day returns),支持运费险(shipping insurance)
Listing Date Filter
beginOfferCreateTime/endOfferCreateTime: Date range inYYYY-MM-DDformat
Pagination
pageIndex: Page number, starting from 1 (default: 1)pageSize: Results per page, 10-100 (default: 20)
Example Queries
1. Basic keyword search -- top sellers for "yoga mat"
{"keyWord": "瑜伽垫", "cycle": "30", "sortField": "saleCount30d", "sortType": "desc", "pageSize": 20}2. Factory-only search with price range
{"keyWord": "蓝牙耳机", "companyType": 2, "beginPrice": 10, "endPrice": 50, "cycle": "30", "sortField": "orderCount30d"}3. Cross-border products with dropship support
{"keyWord": "手机壳", "offerType": 4, "proxyRights": "4360897", "cycle": "7", "sortField": "saleVolume7d"}4. New products listed recently, sorted by listing date
{"keyWord": "夏季连衣裙", "offerType": 2, "sortField": "offerCreateTime", "sortType": "desc", "pageSize": 50}5. High-volume products from Super Factories
{"keyWord": "数据线", "shiLiType": "superFactory", "beginSaleCount": 1000, "cycle": "30", "sortField": "saleCount30d"}6. Search by product URL
{"goodsUrl": "https://detail.1688.com/offer/805578065498.html", "cycle": "30"}Display Rules
1. Present data clearly: Show results in structured tables with product title, price, sales metrics, and supplier info. Include product URLs so users can visit listings directly. 2. Price context: Always show both wholesale price and dropship price when available, so users can compare margins. 3. Sales metrics: Clearly label whether metrics are 7-day or 30-day figures based on the cycle parameter used. 4. Image display: When imageUrl is available, display product images to help users visually identify products. 5. Pagination notice: When total exceeds the returned page size, inform users of the total result count and that they can request additional pages. 6. Error handling: If a query returns an error, explain the issue and suggest adjusting parameters (e.g., broadening filters, checking keyword spelling). 7. Keyword translation: If the user provides English product terms, translate to Chinese before calling the API, and note this translation in your response.
Applicable Scenarios
| User Says | Scenario |
|---|---|
| "Find suppliers for XX on 1688" | Keyword product search |
| "What's selling well on 1688" | Top-selling product discovery |
| "Find cheap XX from factories" | Factory sourcing with price filters |
| "1688 dropshipping suppliers for XX" | Dropship-enabled product search |
| "Compare prices for XX on 1688" | Price comparison across suppliers |
| "New products on 1688 for XX" | New product discovery |
| "Find 1688 products for cross-border" | Cross-border product sourcing |
| "Which 1688 suppliers ship within 24h" | Fulfillment speed filtering |
| "Top factories for XX" | Super Factory / Power Merchant search |
| "Find this 1688 product" (with URL or ID) | Direct product lookup |
Not Applicable
- Amazon or other non-1688 platform product research
- 1688 store/shop-level analytics (store traffic, store rankings)
- 1688 advertising or promotion strategies
- Product review or rating analysis
- Logistics cost calculation or freight estimation
- Order placement or transaction processing
- User already has local product data files to analyze
Boundary judgment: When users say "sourcing", "find suppliers", or "wholesale products", if it involves searching for specific products on the 1688 platform with filters like price, sales, or supplier type, this skill applies. If they need store-level analytics, advertising optimization, or operations beyond product search, it does not apply.
Feedback:
Auto-detect and report feedback via the Feedback API when any of the following apply: 1. The functionality or purpose described in this skill does not match actual behavior 2. The skill's results do not match the user's intent 3. The user expresses dissatisfaction or praise about this skill 4. Anything you believe could be improved
Call the feedback API as specified in references/api.md. Do not interrupt the user's flow.
<!-- LF_LARGE_RESPONSE_BLOCK -->
Handling Large Responses
To avoid overflowing the agent context, persist the response to disk and extract only the fields you need:
python scripts/response_io.py run --script scripts/dld_product_search.py --out-dir <DIR> '<params>'
python scripts/response_io.py read <file> --fields "<paths>" # or --path "<JMESPath>"Pick--out-diroutside any git working tree (e.g./tmp/...on Unix,%TEMP%/...on Windows). Persisted responses may contain PII, pricing, or auth-sensitive data — do not commit them. Files are not auto-deleted; clean up when the task is done.
run writes the full response to a file and emits only a schema preview + file path. read projects specific fields, with --limit/--offset for slicing and --format json|jsonl|csv|table for output.
When to prefer this pattern — apply your judgment based on the response characteristics, e.g.:
- High field count per record, or fields you don't need
- Batch/paginated results (multiple items per call)
- Long-text fields (descriptions, reviews, HTML, time series)
- Output reused across later steps rather than consumed immediately
For small, single-use responses, calling the main script directly is fine.
⚠️ The preview is a truncated schema + sample, not the full data. Any field-level decision must read from the persisted file via read. <!-- /LF_LARGE_RESPONSE_BLOCK -->
--- For more high-quality, professional cross-border e-commerce skills, set [LinkFox Skills](https://skill.linkfox.com/).
店雷达 1688 选品库 API 参考
调用规范
- 请求地址:
https://tool-gateway.linkfox.com/dld/productSearch - 请求方式:POST,Content-Type: application/json
- 认证方式:Header
Authorization: <api_key>,api_key 从环境变量LINKFOXAGENT_API_KEY读取(如未配置,提示用户前往 https://skill.linkfox.com/linkfoxskills/guide.htm 申请)
请求参数
POST Body(JSON):
| 参数 | 类型 | 必填 | 默认值 | 说明 |
|---|---|---|---|---|
| keyWord | string | 否 | - | 搜索关键词(必须为中文,最多50字符) |
| goodsUrl | string | 否 | - | 商品链接地址(与keyWord二选一) |
| productIds | string | 否 | - | 商品ID,多个逗号隔开,最多20个 |
| cycle | string | 否 | - | 统计周期:7(近7天)或 30(近30天) |
| searchType | integer | 否 | 1 | 搜索类型:1-模糊匹配,3-精准匹配 |
| sortField | string | 否 | orderCount30d | 排序字段:orderCount7d, saleCount7d, saleVolume7d, orderCount30d, saleCount30d, saleVolume30d, offerCreateTime, price, consignPrice |
| sortType | string | 否 | desc | 排序类型:desc(降序)、asc(升序) |
| pageIndex | integer | 否 | 1 | 页码(从1开始) |
| pageSize | integer | 否 | 20 | 每页数量(10-100) |
| beginPrice | number | 否 | - | 批发价(起始) |
| endPrice | number | 否 | - | 批发价(结束) |
| beginConsignPrice | number | 否 | - | 代发价(起始) |
| endConsignPrice | number | 否 | - | 代发价(结束) |
| beginOrderCount | integer | 否 | - | 销售笔数(起始) |
| endOrderCount | integer | 否 | - | 销售笔数(结束) |
| beginSaleCount | integer | 否 | - | 销售件数(起始) |
| endSaleCount | integer | 否 | - | 销售件数(结束) |
| beginSaleVolume | number | 否 | - | 销售额(起始) |
| endSaleVolume | number | 否 | - | 销售额(结束) |
| beginStartQuantity | integer | 否 | - | 起购数量(起始) |
| endStartQuantity | integer | 否 | - | 起购数量(结束) |
| beginTpYear | integer | 否 | - | 诚信通年限(起始) |
| endTpYear | integer | 否 | - | 诚信通年限(结束) |
| beginOfferCreateTime | string | 否 | - | 上架时间起始(格式:YYYY-MM-DD) |
| endOfferCreateTime | string | 否 | - | 上架时间结束(格式:YYYY-MM-DD) |
| companyType | integer | 否 | 0 | 公司类型:0-不限,1-店铺,2-工厂 |
| offerType | integer | 否 | 0 | 商品标识:0-不限,2-新品,3-1688严选,4-跨境,5-支持定制,6-镇店之宝 |
| shiLiType | string | 否 | - | 卖家类型(多选逗号隔开):superFactory(超级工厂)、Power(实力商家)、TrustPass(诚信通) |
| sendTime | string | 否 | - | 发货时间(多选逗号隔开):24、48、72 |
| faceToFaceSupport | string | 否 | - | 面单支持(多选逗号隔开):441218(淘宝)、386434(抖音)、422914(拼多多)、422978(小红书)、386370(快手) |
| proxyRights | string | 否 | - | 代发权益(多选逗号隔开):4360897(一件代发包邮)、449154(先采后付) |
| shopService | string | 否 | - | 卖家服务(多选逗号隔开):4057409(安心购)、888777(深度认证报告) |
| buyerProtections | string | 否 | - | 权益保障(多选逗号隔开):商品包邮、7天包退货、支持运费险 |
响应结构
| 字段 | 类型 | 说明 |
|---|---|---|
| total | integer | 总记录数 |
| type | string | 渲染样式 |
| columns | array | 渲染列定义 |
| products | array | 商品列表(详见下方字段) |
products 数组元素字段
| 字段 | 类型 | 说明 |
|---|---|---|
| offerId | string | 商品ID |
| asin | string | 商品编号 |
| title | string | 商品标题 |
| asinUrl | string | 商品链接地址 |
| imageUrl | string | 商品图片地址 |
| price | number | 批发价 |
| consignPrice | number | 代发价 |
| quantityPrices | string | 价格区间 |
| quantityBegin | integer | 起批量 |
| unit | string | 单位 |
| currency | string | 币种 |
| salesOrderCount | integer | 销售笔数(按统计周期) |
| salesQuantity | integer | 销售件数(按统计周期) |
| estimatedSalesAmount | integer | 预估销售额(按统计周期) |
| deliveryTime | string | 发货时间 |
| availableDate | string | 商品上架时间 |
| levelName | string | 类目层级名称 |
| company | string | 店铺名称 |
| shopId | string | 店铺ID |
| shopUrl | string | 店铺链接地址 |
| dataType | string | 数据类型:weeklyData(周数据)、monthlyData(月数据) |
| sourceType | string | 来源平台(1688) |
| sourceTool | string | 来源工具 |
错误码
正常情况下,接口的 HTTP 状态码均为 200,业务的成功与否通过响应体中的 errorCode 字段区分(errorCode = 200 表示成功,其他值表示业务错误)。当遇到未授权等情况时,HTTP 状态码为 401,且对应的 errorCode 也是 401。
| errcode | 含义 | 处理建议 |
|---|---|---|
| 200 | 成功 | 正常解析业务字段 |
| 401 | 认证失败 | 检查请求头 Authorization 是否正确携带 API Key;API Key 申请方式请参考上述调用规范下的认证方式。 |
| 其他非200值 | 业务异常 | 参考 errmsg 字段获取具体错误原因 |
错误响应示例:
{
"errcode": 401,
"errmsg": "authorized error"
}curl 示例
curl -X POST https://tool-gateway.linkfox.com/dld/productSearch \
-H "Authorization: $LINKFOXAGENT_API_KEY" \
-H "Content-Type: application/json" \
-d '{"keyWord": "瑜伽垫", "cycle": "30", "sortField": "saleCount30d", "sortType": "desc", "pageSize": 20}'---
Feedback API
This endpoint is separate from the tool API above. Do not mix the two base URLs.
- POST
https://skill-api.linkfox.com/api/v1/public/feedback - Content-Type:
application/json
{
"skillName": "linkfox-xxx-xxx",
"sentiment": "POSITIVE",
"category": "OTHER",
"content": "Results were accurate, user was satisfied."
}Field rules:
skillName: Use this skill'snamefrom the YAML frontmattersentiment: Choose ONE —POSITIVE(praise),NEUTRAL(suggestion without emotion),NEGATIVE(complaint or error)category: Choose ONE —BUG(malfunction or wrong data),COMPLAINT(user dissatisfaction),SUGGESTION(improvement idea),OTHERcontent: Include what the user said or intended, what actually happened, and why it is a problem or praise
#!/usr/bin/env python3
"""
DianLeiDa 1688 Product Search - LinkFox Skill
Calls the dld/productSearch API endpoint
Usage:
python dld_product_search.py '{"keyWord": "瑜伽垫", "cycle": "30", "sortField": "saleCount30d", "pageSize": 20}'
"""
import json
import os
import sys
from urllib.request import urlopen, Request
from urllib.error import HTTPError, URLError
API_URL = "https://tool-gateway.linkfox.com/dld/productSearch"
def get_api_key():
"""Retrieve the API key from environment, with a friendly prompt if missing."""
key = os.environ.get("LINKFOXAGENT_API_KEY")
if not key:
print(
"API Key not configured. Please complete authorization first:\n"
"1. Visit https://skill.linkfox.com/linkfoxskills/guide.htm to obtain your Key\n"
"2. Set the environment variable: export LINKFOXAGENT_API_KEY=your-key-here",
file=sys.stderr,
)
sys.exit(1)
return key
def call_api(params: dict) -> dict:
"""Call the tool gateway API."""
api_key = get_api_key()
data = json.dumps(params).encode("utf-8")
req = Request(
API_URL,
data=data,
headers={
"Authorization": api_key,
"Content-Type": "application/json",
"User-Agent": "LinkFox-Skill/1.0",
},
method="POST",
)
try:
with urlopen(req, timeout=60) as response:
return json.loads(response.read().decode("utf-8"))
except HTTPError as e:
body = e.read().decode("utf-8") if e.fp else ""
return {"error": f"HTTP {e.code}: {e.reason}", "details": body}
except URLError as e:
return {"error": f"Connection failed: {e.reason}"}
def main():
if len(sys.argv) < 2:
print("Usage: dld_product_search.py '<JSON parameters>'", file=sys.stderr)
print(
'Example: dld_product_search.py \'{"keyWord": "瑜伽垫", "cycle": "30", "sortField": "saleCount30d", "pageSize": 20}\'',
file=sys.stderr,
)
sys.exit(1)
try:
params = json.loads(sys.argv[1])
except json.JSONDecodeError as e:
print(f"Invalid parameter format: {e}", file=sys.stderr)
sys.exit(1)
result = call_api(params)
print(json.dumps(result, indent=2, ensure_ascii=False))
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
Skill response I/O helper — wraps any main script to persist large API
responses to disk, then offers a `read` subcommand to extract specific fields
from those persisted files. Generic, business-agnostic.
This script is bundled into each skill's scripts/ directory by tools/response_io/sync.py.
The agent must pass --script <path> to identify which main script to execute.
Usage:
python scripts/response_io.py run --script <PATH> --out-dir <DIR> '<json_params>' [--label NAME] [--timeout SEC]
python scripts/response_io.py read <file> (--path "<JMESPath>" | --fields "f1,f2,...") [--limit N] [--offset M] [--format json|jsonl|csv|table]
"""
from __future__ import annotations
import sys
if sys.version_info < (3, 10):
sys.exit(
"Error: Python 3.10+ required (current: "
f"{sys.version_info.major}.{sys.version_info.minor}). "
"Please upgrade Python."
)
import argparse
import csv
import io
import json
import os
import re
import secrets
import subprocess
from datetime import datetime
from pathlib import Path
from typing import Any
# Force UTF-8 stdout/stderr so non-ASCII chars in previews and API responses
# print correctly on Windows (default cp936 / gbk).
for stream in (sys.stdout, sys.stderr):
try:
stream.reconfigure(encoding="utf-8") # type: ignore[attr-defined]
except (AttributeError, OSError):
pass
try:
import jmespath # type: ignore
HAS_JMESPATH = True
except ImportError:
HAS_JMESPATH = False
MAX_STRING_LEN = 120
MAX_DEPTH = 3
SAMPLE_KEY_CAP = 15
RAW_TEXT_PEEK = 500
DEFAULT_TIMEOUT_SEC = 300
# ---------------------------------------------------------------------------
# Shared helpers
# ---------------------------------------------------------------------------
def _err(msg: str, code: int = 1) -> None:
print(msg, file=sys.stderr)
sys.exit(code)
def _resolve_script(script_arg: str) -> Path:
p = Path(script_arg).expanduser()
if not p.is_absolute():
# Resolve relative to the current working directory the agent invoked from.
p = (Path.cwd() / p).resolve()
else:
p = p.resolve()
if not p.is_file():
_err(f"--script path not found: {p}")
return p
def _resolve_skill_name(main_script: Path) -> str:
"""Best-effort skill name extraction for filename prefixing.
main_script lives at <skill_dir>/scripts/<name>.py — return <skill_dir>'s
folder name. Fall back to the script's stem if structure differs.
"""
try:
if main_script.parent.name == "scripts":
return main_script.parents[1].name
except IndexError:
pass
return main_script.stem
def _sanitize_label(label: str) -> str:
"""Allow only safe filename chars in --label to prevent path traversal."""
cleaned = re.sub(r"[^\w\-]", "_", label)
return cleaned[:64] # cap length
def _truncate_string(s: str) -> str:
if len(s) <= MAX_STRING_LEN:
return s
return s[:MAX_STRING_LEN] + f"...(truncated, total {len(s)} chars)"
def _truncate_value(value: Any, depth: int = 0) -> Any:
"""Recursively truncate strings, deep nesting, and large arrays for preview."""
if depth >= MAX_DEPTH:
if isinstance(value, dict):
return f"<truncated nested object, keys: {list(value.keys())[:10]}>"
if isinstance(value, list):
return f"<truncated nested array, length: {len(value)}>"
if isinstance(value, str):
return _truncate_string(value)
return value
if isinstance(value, str):
return _truncate_string(value)
if isinstance(value, dict):
out = {k: _truncate_value(v, depth + 1) for k, v in value.items()}
return out
if isinstance(value, list):
if not value:
return []
truncated = [_truncate_value(value[0], depth + 1)]
if len(value) > 1:
# Note total length on the parent — keep the array type-homogeneous
# so downstream consumers can iterate without special-casing strings.
truncated.append({"_omitted_items": len(value) - 1})
return truncated
return value
def _shape_of(value: Any, top: bool = False) -> Any:
"""Lightweight schema description for the preview block."""
if isinstance(value, dict):
keys = list(value.keys())
out: dict[str, Any] = {"type": "object", "top_keys" if top else "keys": keys}
if top:
for k in keys[:8]:
out[k] = _shape_of(value[k])
return out
if isinstance(value, list):
out = {"type": "array", "length": len(value)}
if value and isinstance(value[0], dict):
out["item_keys"] = list(value[0].keys())
elif value:
out["item_type"] = type(value[0]).__name__
return out
return {"type": type(value).__name__}
def _build_sample(value: Any) -> Any:
"""First-record sample with explicit truncation marker."""
if isinstance(value, list):
if not value:
return {"_truncated_record": True, "_note": "array is empty"}
first = value[0]
if isinstance(first, dict):
sample = {"_truncated_record": True, "_note": f"first of {len(value)} items"}
sample.update(_truncate_value(first, depth=1))
return sample
return {"_truncated_record": True, "_note": f"first of {len(value)} items", "value": _truncate_value(first, depth=1)}
if isinstance(value, dict):
sample = {"_truncated_record": True, "_note": "top-level object (truncated)"}
sample.update(_truncate_value(value, depth=1))
return sample
return {"_truncated_record": True, "value": _truncate_value(value, depth=1)}
def _shrink_preview(preview: dict) -> dict:
"""Cap the sample's value fields when it has many keys.
`shape.*.item_keys` is the single source of truth for the full key list
(always complete, no truncation). The sample only ever shows up to
SAMPLE_KEY_CAP fields with their concrete values, since the agent only
needs a feel for value shapes — for the full menu of available fields,
they read `shape`.
"""
sample = preview.get("sample")
if isinstance(sample, dict):
meta_keys = {"_truncated_record", "_note"}
data_keys = [k for k in sample.keys() if k not in meta_keys]
if len(data_keys) > SAMPLE_KEY_CAP:
kept = data_keys[:SAMPLE_KEY_CAP]
new_sample = {k: v for k, v in sample.items() if k in meta_keys or k in kept}
base_note = sample.get("_note", "")
extra = (
f"showing first {SAMPLE_KEY_CAP} of {len(data_keys)} fields "
f"(see `shape` for the complete key list)"
)
new_sample["_note"] = f"{base_note}; {extra}" if base_note else extra
preview["sample"] = new_sample
return preview
# ---------------------------------------------------------------------------
# `run` subcommand
# ---------------------------------------------------------------------------
def cmd_run(args: argparse.Namespace) -> int:
main_script = _resolve_script(args.script)
skill_name = _resolve_skill_name(main_script)
out_dir = Path(args.out_dir).expanduser().resolve()
try:
out_dir.mkdir(parents=True, exist_ok=True)
except OSError as e:
_err(f"Failed to create --out-dir {out_dir}: {e}")
if not os.access(out_dir, os.W_OK):
_err(f"--out-dir is not writable: {out_dir}")
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
rand = secrets.token_hex(3)
safe_label = _sanitize_label(args.label) if args.label else ""
label_part = f"__{safe_label}" if safe_label else ""
out_file = out_dir / f"{skill_name}__{timestamp}_{rand}{label_part}.json"
# Force the child process to emit UTF-8 regardless of the host console
# encoding (Windows defaults to cp936 / gbk and would otherwise corrupt
# non-ASCII bytes when we read them back).
child_env = os.environ.copy()
child_env["PYTHONIOENCODING"] = "utf-8"
timed_out = False
try:
proc = subprocess.run(
[sys.executable, str(main_script), args.params],
capture_output=True,
text=True,
encoding="utf-8",
errors="replace",
env=child_env,
timeout=args.timeout,
)
stdout_text = proc.stdout or ""
stderr_text = proc.stderr or ""
returncode = proc.returncode
except subprocess.TimeoutExpired as e:
timed_out = True
stdout_text = (e.stdout.decode("utf-8", errors="replace") if isinstance(e.stdout, bytes) else (e.stdout or "")) or ""
stderr_text = (e.stderr.decode("utf-8", errors="replace") if isinstance(e.stderr, bytes) else (e.stderr or "")) or ""
returncode = 124 # convention for timeout
# Always write the captured stdout to disk, even if not JSON.
try:
out_file.write_text(stdout_text, encoding="utf-8")
except OSError as e:
_err(f"Failed to write output file {out_file}: {e}")
if stderr_text:
sys.stderr.write(stderr_text)
# Try to parse the captured stdout as JSON for the preview.
try:
parsed = json.loads(stdout_text) if stdout_text.strip() else None
format_kind = "json"
except json.JSONDecodeError:
parsed = None
format_kind = "raw_text"
preview: dict[str, Any] = {
"_preview": {
"is_preview": True,
"warning": (
"PREVIEW ONLY — NOT FULL DATA. The full response is saved to `file`. "
"Use `python scripts/response_io.py read <file> --fields '...'` to extract "
"specific fields, or `--path '<JMESPath>'` for complex projections."
),
},
}
# Surface failures prominently so agents don't mistake a stub preview for success.
if returncode != 0 or timed_out:
stderr_snippet = stderr_text[-500:] if stderr_text else ""
preview["_error"] = {
"exit_code": returncode,
"timed_out": timed_out,
"stderr_snippet": stderr_snippet,
"hint": "The wrapped script failed or timed out. The output file may be empty or partial.",
}
preview.update({
"file": str(out_file),
"size_bytes": out_file.stat().st_size,
"skill": skill_name,
"exit_code": returncode,
"format": format_kind,
"label": safe_label or None,
"next_steps_hint": (
"use: python scripts/response_io.py read <file> --fields '...' | --path '...'"
),
})
if format_kind == "json":
preview["shape"] = _shape_of(parsed, top=True)
preview["sample"] = _build_sample(parsed)
else:
peek = stdout_text[:RAW_TEXT_PEEK]
preview["raw_text_peek"] = peek
preview["raw_text_total_chars"] = len(stdout_text)
preview["sample"] = {
"_truncated_record": True,
"_note": f"stdout was not valid JSON; first {RAW_TEXT_PEEK} chars shown above in raw_text_peek",
}
preview = _shrink_preview(preview)
print(json.dumps(preview, ensure_ascii=False, indent=2))
return returncode
# ---------------------------------------------------------------------------
# `read` subcommand
# ---------------------------------------------------------------------------
def _load_json(path: Path) -> Any:
try:
text = path.read_text(encoding="utf-8")
except OSError as e:
_err(f"Failed to read file {path}: {e}")
try:
return json.loads(text)
except json.JSONDecodeError as e:
_err(f"File is not valid JSON: {path}\n{e}")
def _basic_dot_path(data: Any, path: str) -> Any:
"""Pure-stdlib dot-path resolver. No [*] support — callers fall back here only when jmespath is unavailable AND the path has no [*]."""
cur = data
for part in path.split("."):
if isinstance(cur, dict):
cur = cur.get(part)
else:
return None
return cur
def _resolve_field(data: Any, expr: str) -> Any:
if HAS_JMESPATH:
return jmespath.search(expr, data)
if "[" in expr or "*" in expr:
_err(
f"jmespath is required for expression '{expr}'. "
f"Install with: pip install jmespath"
)
return _basic_dot_path(data, expr)
def _project_fields(data: Any, fields: list[str]) -> Any:
"""Run each field expr; if any returns a list, zip them into list-of-dicts."""
resolved: dict[str, Any] = {f: _resolve_field(data, f) for f in fields}
list_lengths = [len(v) for v in resolved.values() if isinstance(v, list)]
if not list_lengths:
return resolved
# All list values must be same length to zip cleanly.
if len(set(list_lengths)) > 1:
# Fallback: return the dict as-is so caller can inspect mismatches.
return resolved
n = list_lengths[0]
rows = []
for i in range(n):
row = {}
for f, v in resolved.items():
row[f] = v[i] if isinstance(v, list) else v
rows.append(row)
return rows
def _apply_slice(value: Any, limit: int | None, offset: int | None) -> Any:
if not isinstance(value, list):
return value
start = offset or 0
end = (start + limit) if limit is not None else None
return value[start:end]
def _format_output(value: Any, fmt: str) -> str:
if fmt == "json":
return json.dumps(value, ensure_ascii=False, indent=2)
if fmt == "jsonl":
if isinstance(value, list):
return "\n".join(json.dumps(item, ensure_ascii=False) for item in value)
return json.dumps(value, ensure_ascii=False)
if fmt in ("csv", "table"):
if not isinstance(value, list) or not value:
_err(f"--format {fmt} requires a non-empty list result")
if not all(isinstance(item, dict) for item in value):
_err(f"--format {fmt} requires list-of-objects, got list of {type(value[0]).__name__}")
keys: list[str] = []
for item in value:
for k in item.keys():
if k not in keys:
keys.append(k)
if fmt == "csv":
buf = io.StringIO()
writer = csv.DictWriter(buf, fieldnames=keys, extrasaction="ignore")
writer.writeheader()
for item in value:
writer.writerow({k: _stringify(item.get(k)) for k in keys})
return buf.getvalue().rstrip("\n")
# table: simple aligned columns
rows = [[_stringify(item.get(k)) for k in keys] for item in value]
widths = [len(k) for k in keys]
for row in rows:
for i, cell in enumerate(row):
widths[i] = max(widths[i], len(cell))
lines = [
" ".join(k.ljust(widths[i]) for i, k in enumerate(keys)),
" ".join("-" * widths[i] for i in range(len(keys))),
]
for row in rows:
lines.append(" ".join(row[i].ljust(widths[i]) for i in range(len(keys))))
return "\n".join(lines)
_err(f"Unknown --format: {fmt}")
return "" # unreachable
def _stringify(v: Any) -> str:
if v is None:
return ""
if isinstance(v, (dict, list)):
return json.dumps(v, ensure_ascii=False)
return str(v)
def cmd_read(args: argparse.Namespace) -> int:
if not args.path and not args.fields:
_err("read: either --path or --fields is required")
if args.path and args.fields:
_err("read: --path and --fields are mutually exclusive")
file_path = Path(args.file).expanduser().resolve()
data = _load_json(file_path)
if args.path:
result = _resolve_field(data, args.path)
else:
fields = [f.strip() for f in args.fields.split(",") if f.strip()]
if not fields:
_err("--fields parsed to empty list")
result = _project_fields(data, fields)
result = _apply_slice(result, args.limit, args.offset)
print(_format_output(result, args.format))
return 0
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def main() -> int:
parser = argparse.ArgumentParser(
prog="response_io.py",
description="Persist large skill API responses to disk and read fields on demand.",
)
sub = parser.add_subparsers(dest="cmd", required=True)
p_run = sub.add_parser(
"run",
help="Execute a main script and persist its stdout to a file; "
"print only a lightweight preview to stdout.",
)
p_run.add_argument("params", help="JSON params string passed verbatim to the main script (argv[1]).")
p_run.add_argument("--script", required=True, help="Path to the main script to execute, e.g. scripts/my_api.py")
p_run.add_argument("--out-dir", required=True, help="Directory to write the response file into (created if missing).")
p_run.add_argument("--label", default=None, help="Optional filename suffix; sanitized to safe filename characters.")
p_run.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT_SEC, help=f"Subprocess timeout in seconds (default: {DEFAULT_TIMEOUT_SEC}).")
p_run.set_defaults(func=cmd_run)
p_read = sub.add_parser(
"read",
help="Extract specific fields from a previously persisted response file.",
)
p_read.add_argument("file", help="Path to the persisted JSON response file.")
g = p_read.add_mutually_exclusive_group()
g.add_argument("--path", default=None, help="JMESPath expression, e.g. 'data[*].{asin: asin, title: title}'.")
g.add_argument("--fields", default=None, help="Comma-separated field paths, e.g. 'data[*].asin,data[*].title'.")
p_read.add_argument("--limit", type=int, default=None, help="Take at most N items (when result is a list).")
p_read.add_argument("--offset", type=int, default=None, help="Skip the first M items (when result is a list).")
p_read.add_argument("--format", choices=["json", "jsonl", "csv", "table"], default="json", help="Output format (default: json).")
p_read.set_defaults(func=cmd_read)
args = parser.parse_args()
return args.func(args)
if __name__ == "__main__":
sys.exit(main())