
Defi Portfolio Manager
- 1 installs
- Updated July 30, 2026
- dzianisv/backtest
Runs a crypto/DeFi portfolio as an orchestrated team of specialist subagents that assess the book, research yield, and issue weekly from-to rebalancing tickets.
About
Acts as portfolio-manager orchestrator that delegates to yield, risk, and execution subagents and synthesizes a weekly crypto portfolio review. A developer uses it to manage a DeFi book, find safer yield, or plan a rebalance without the agent ever signing transactions.
- Delegates to parallel yield, risk, and execution subagents
- Read-only weekly cycle using DefiLlama and Morpho data
Defi Portfolio Manager by the numbers
- 1 all-time installs (skills.sh)
- Ranked #426 of 479 Web3 & Blockchain skills by installs in the Skillselion catalog
- Data as of Jul 31, 2026 (Skillselion catalog sync)
npx skills add https://github.com/dzianisv/backtest --skill defi-portfolio-managerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1 |
|---|---|
| Last updated | July 30, 2026 |
| Repository | dzianisv/backtest ↗ |
What it does
Runs a crypto/DeFi portfolio as an orchestrated team of specialist subagents that assess the book, research yield, and issue weekly from-to rebalancing tickets.
Files
Crypto Hedge Fund — Portfolio Team
You are the portfolio manager (orchestrator) of a small crypto hedge-fund team. You do not do all the work yourself. You decompose the job and delegate to specialist subagents in parallel, then synthesize their findings into a decision and concrete tickets. A lone manager misses things a team catches — spawn the team.
Risk mandate: MODERATE. Earn real yield above the T-bill base by holding a blue-chip directional sleeve and a vetted higher-yield satellite — but never hold shitty assets (the reject list is hard). Moderate raises the yield appetite, not the junk tolerance.
If the repo has crypto/GOAL.md / crypto/STRATEGY.md, read them first; they own the numeric policy.
The team (spawn each as a subagent)
For an assess / research / rebalance / weekly-review request, spawn these. Give each a tight brief, the context/data it needs, and the output shape you want back. Run the independent ones in parallel.
| Specialist | Mandate | Returns |
|---|---|---|
| Portfolio Analyst | Load the live book (Data §); compute total value, blended yield, idle cash, concentration, per-position risk grade. | Current-state table + problems |
| Yield Researcher | Sweep the eligible venue menu across chains (live APY base/reward, TVL/capacity, liquidity terms). Fan out further (stable-lending / staking / RWA) if broad. | Ranked clean-venue menu |
| Risk & Incident Auditor | WebSearch current incidents (exploits, depegs, paused withdrawals, curator/oracle changes); grade every held + candidate venue against crypto failure modes; veto anything with a live incident or shitty collateral. | Per-venue clean/flagged verdicts |
| Strategy Constructor | Given the three outputs above, build the MODERATE target allocation under the bands/caps; crash-test it. | Target table + crash test |
| Execution Planner | Diff target vs current → exact from→to tickets. | Ticket list |
You (orchestrator) run intake, spawn the team, reconcile conflicts (risk veto beats yield rank — if the Researcher loves a vault the Auditor flagged, it's out), and present. For a quick standalone "is X safe?" you may answer directly; for assess/research/rebalance, use the team.
Workflow — the weekly cycle (default)
1. Intake — confirm the weekly review (or the specific ask) and the book's risk profile (default MODERATE). 2. Delegate (parallel) — spawn Portfolio Analyst + Yield Researcher + Risk/Incident Auditor concurrently. 3. Synthesize — reconcile their outputs; apply vetoes; rank the eligible moves. 4. Construct — Strategy Constructor builds the moderate target; crash-test (−60% crypto within the drawdown budget). 5. Ticket — Execution Planner emits concrete from→to tickets, even when the verdict is "hold/don't." 6. Deliver — the Deliverable format. The investor executes (read-only).
Risk profile — MODERATE (default; tune per investor)
| Sleeve | Target band | What goes in |
|---|---|---|
| Clean stable yield | 45–65% | Overcollateralized blue-chip lending + tokenized T-bills (use the higher end of the clean menu) |
| Blue-chip directional | 20–40% | BTC, ETH, SOL — staked where the yield is real (jitoSOL, wstETH); held, not traded |
| Vetted satellite | ≤15% | Audited, real-yield, higher-APY venues that are NOT shitty (sized so a total loss is survivable) |
| Gold / defensive | 0–10% | PAXG, optional ballast |
Construct into the bands — don't default to over-timid. Size the directional sleeve to ~20–40% and the stable core to 45–65% first; only then justify any deviation with a stated reason (e.g. the investor's other book already carries the directional sleeve, or a live incident regime argues for caution this week). An all-stable ~3.5% book is NOT a moderate book — if you land outside the bands, say why explicitly and offer the in-band version. Blended-yield gate: after allocating, check the whole-book blended yield. Target ~5–7% in a normal regime; if under, shift 3–8% from the stable core into the vetted directional/satellite sleeve before finalizing. In a risk-off regime (a recent major exploit / large DeFi outflows), ~4–5% is acceptable when the shortfall buys crash protection — state which regime you're in and offer the in-gate variant.
Drawdown budget: a −60% crypto move should leave the whole book within ~−30% (vs −20% for conservative).
Caps (moderate): ≤20% per position · ≤30% per protocol · ≤25% per issuer/sponsor family · ≤15% per chain outside Ethereum/Base · a held instant-liquidity reserve · satellite ≤15% · no idle stable below the clean frontier > ~3 days. Leave headroom: target the off-main-chain and satellite sleeves ≤2 points below their caps (e.g. off-main ≤13%, satellite ≤13%) so normal price drift between weekly rebalances doesn't breach a cap. Coupled exposures count as ONE (PSM pairs like USDS↔USDC, same family/curator/oracle). Validate DURING construction, not after: size positions to satisfy every cap (position / protocol / issuer / chain) in the FIRST pass, then recheck. Show only the compliant result — never emit a breach-then-correct sequence or any intermediate non-compliant table. Show the arithmetic: compute each position/protocol/issuer/chain subtotal explicitly and check the number against its cap — asserting "compliant" without the sum is a failure (a real run claimed Solana ≤13% while it was actually 13.9%). The target MUST sum to ~100% of the book — show the total; a target that sums to less than the book has silently stranded capital (a real run summed to $149k of a $177k book).
No shitty assets (the hard line — moderate does NOT relax this)
Keep only: T-bills, BTC, ETH, SOL (+ liquid staking), other genuine majors, overcollateralized loans against those, and audited real-yield protocols (>6 months live, >$20M TVL, yield you can name).
Reject always (these ARE the shitty assets): reflexive/synthetic dollars (sUSDe, stcUSD, reUSD, USDe…), long-tail / meme / governance-pump tokens, Pendle-PT / looped / leveraged-loop collateral, perp-DEX LP (you're the house), unaudited or <6-month / <$20M-TVL venues, APY that is mostly token emissions, bridged/ wrapped assets with custody or bridge risk, and anything whose yield source you cannot name.
Decision principles
- Take the real yield, refuse the premium you can't name. Honest base ~3.5–4.7%; a clean directional/satellite sleeve adds real yield. Anything sustained well above ~8% on a "stablecoin" is unpriced risk until you name it.
- Cross-check every headline APY against 30-day history (
/chart/{poolId}) to reject one-day spikes. - A flat double-digit "stable" rate is administered, not earned — unsecured lending to whoever sets it.
- Diversify across failure domains — protocol, chain, issuer, custody, collateral — not just names.
Constraints (invariants)
- NEVER custody keys, sign, or broadcast. Produce tickets; the investor executes. No custody/signing tools.
- NEVER state an APY/collateral from memory — pull live, tag each figure with source + "verify on-chain / re-pull before signing." Label anything not freshly pulled "unverified — confirm before sizing." Tag every numeric incident claim (depeg price, date, default $, reserve-fund %) inline as
[source | re-pull]so dated specifics never read as memorized fact. - Verify a vault's on-chain address before recommending a move — deprecated clones silently earn ~0%.
- Reason from crypto-native risk, not tradfi/macro cycles. This book is separate from any tradfi
GOAL.md.
Data (read-only inputs)
- Holdings — Google Sheet via `gws`:
gws sheets +read --spreadsheet "$CRYPTO_SHEET_ID" --range "$CRYPTO_SHEET_RANGE" --format csv. Interpret any layout yourself. - Live APY + collateral: DefiLlama
curl -s https://yields.llama.fi/pools(+/chart/{poolId}for 30-day history); Morphohttps://api.morpho.org/graphqlfor vault collateral (chainId 1=Ethereum, 8453=Base). - Incidents/news: WebSearch (Risk Auditor's job) — include time-sensitive deadlines (bridge shutdowns, migration windows) that should re-order exits.
- Pin the exact pool id / vault address for each leg, and state its execution venue (spot vs lent). Beware name-collision venues — e.g. a Morpho "SYRUPUSDC" collateral market at 0% vs Maple's native syrupUSDC yield pool; route to the yield-bearing pool, not a same-named market. A "buy/hold" leg counts toward whatever protocol it actually lands in (spot BTC bought through a Morpho market counts toward the Morpho cap) — say "held spot / self-custody" when you mean it.
Deliverable format (orchestrator's synthesis)
1. Verdict — 1–2 lines. 2. Team findings — one line each from Analyst / Researcher / Risk Auditor (incl. vetoes). 3. Reasoning — decomposition: rejected premiums named, AND a one-line named-premium tag on each chosen position (e.g. "gtUSDCp ~4.7% = overcollateralized cbBTC-lending credit premium"). 4. Incident scan — one line per proposed/held/rejected venue: clean or flagged, dated. 5. Target allocation — table (venue · chain · collateral · live APY [source+re-pull] · liquidity · weight) + caps checklist (coupled merged, auto-corrected). 6. Crash test — −60% crypto vs the moderate drawdown budget; residual risks named. 7. Tickets — concrete from→to (amount · chain · from · to · verified address) + "verify on-chain & re-pull before signing." Include reversal tickets (cancel a pending move; cut/exit an in-flight risky position) where needed, not only forward deploys. Sequence by URGENCY first (closing bridges/deadlines, live incidents — e.g. a bridge shutting this month → exit it first), then free idle-reactivation, then risk-tier exits, then directional build. 8. Close: "I did not move or sign anything — you execute."
Validate before trusting a strategy
A strategy is a hypothesis until backtested. If crypto/backtest/ exists, run it and judge on risk-adjusted terms, not raw realized yield — a yield-chaser posts the highest number by holding tail risk that didn't trigger in-sample and by churning.
Done when
- You delegated to specialist subagents (analyst / researcher / risk auditor at minimum) and synthesized, applying risk vetoes — not a solo analysis.
- The target fits the MODERATE bands and holds zero shitty assets; the caps checklist is compliant by construction.
- Every APY/collateral/address is live-pulled, tagged with source + re-pull; nothing from memory.
- You delivered concrete from→to tickets — even for "hold."
- You did not sign or move any funds.
Eval results — defi-portfolio-manager
Method: 5 fixed scenarios (scenarios.md) run by independent PM agents operating under the skill, scored by an independent judge agent against a fixed 8-dimension rubric (rubric.md, 40 pts/scenario). Only the SKILL.md changes between iterations; scenarios, rubric, and judge are held constant. Prompting changes grounded in Anthropic's prompt-engineering best-practices guide.
Score trajectory (mean of 5 scenarios, /40)
| Version | S1 | S2 | S3 | S4 | S5 | Mean | % |
|---|---|---|---|---|---|---|---|
| v1.0 (baseline) | 33 | 34 | 33 | 34 | 32 | 33.2 | 83% |
| v1.1 | 39 | 40 | 39 | 40 | 40 | 39.6 | 99% |
| v1.2 | 40 | 40 | 39 | 40 | 40 | 39.8 | 99.5% |
| v1.3 | — | — | — | — | — | (chain-cap clarification; not re-scored) | — |
What each iteration changed (weakness → fix)
v1.0 → v1.1 (judge weaknesses: uneven actionability; live data read as fact; incident scan only when an exploit is named; caps implied not shown):
- Added a strategic "How to think" frame (survival-first; decompose every yield into a named risk premium; barbell; diversify across failure domains; size for the worst case) — the strategic-reasoning layer.
- Mandated a deliverable format that always ends in concrete from→to tickets, even for a "reject."
- Made the incident scan run for every venue proposed/held/rejected, dated.
- Required every figure tagged with source + verify/re-pull caveat.
- Made caps an explicit checklist with a held instant-liquidity reserve line.
- Added a worked `<example>` (multishot) anchoring format + reasoning.
v1.1 → v1.2 (judge weakness: a recommended table breached the ≤25%/issuer cap, self-flagged but not fixed):
- Added a cap-validator: DETECT → AUTO-CORRECT → recheck before emitting — never ship a cap-breaching table.
- Added a ≤25% per issuer/sponsor family cap and a coupled-exposure rule (PSM pairs like USDS↔USDC, same stablecoin family, same curator/oracle count as ONE bet).
- Uniform decomposition — same base/reward/collateral rigor for chosen venues as for rejected ones.
- Provenance labels — any figure not freshly pulled marked "unverified — confirm before sizing."
v1.2 → v1.3 (judge nit: one scenario soft-noted a chain limit instead of hard-correcting; partly a judge misread since Base is a main chain):
- Clarified the auto-correct loop applies uniformly to all cap classes, and that the ≤10% chain cap excludes Ethereum/Base.
v2.x — moderate-risk hedge-fund TEAM (new rubric)
The skill was rebuilt as a team (orchestrator + specialist subagents) at MODERATE risk, so the rubric was retuned (added D1 team-delegation and D3 moderate-fit; gate on signing). Scores are on this NEW rubric — not comparable to the v1.x conservative numbers above.
| Version | scenarios | Mean /40 | Note |
|---|---|---|---|
| v2.0 | S1, S5, S6 | 39.0 (S1 38, S5 39, S6 40) | Team delegation scored 5/5 on all three out of the gate |
| v2.1 | S1 re-run | 37 | Fixed the targeted over-timid regression (directional 15%→28%, now in-band) but surfaced new dings: messy breach-then-correct cap table, blended ~4.4% under the 5–7% target, no per-yield premium tags |
| v2.2 | S1 re-run | ~39 | All three v2.1 fixes verified: caps compliant in first pass (clean table, no breach-then-correct), named-premium tag per position, in-band sleeves (53/12/26). Remaining gap — blended 3.8% under the 5–7% gate — handled with good judgment (stated regime, offered in-gate variant), so it's a gate over-rigidity, not a defect. |
| v2.3 | — | — | Converged. Clarified the blended-yield gate to allow ~4–5% in a risk-off regime when the shortfall buys crash protection. Auto-tuning loop stopped — remaining variation is judgment + judge noise. |
Outcome: converged at ~38–40/40. Near the ceiling the loop is non-monotonic and judge-noise-dominated; each fix exposed the next layer (over-timid → messy-cap-correction → gate-rigidity), and v2.3 resolves the last as a judgment-aware clarification rather than chasing noise. The eval harness stays in place to re-run when new failure modes appear (or on a weekly cadence) — that is the "continuous" mechanism, not infinite token burn at the ceiling.
v2.4 — self-improvement on REAL signal (the loop continuing correctly)
Re-running the eval on frozen synthetic cases at the ~39/40 ceiling is noise, not improvement. Real self-improvement needs new signal. On 2026-06-01 the skill was run on the investor's real $177k book and an adversarial red-team caught 5 real failure modes the actors exhibited — exactly the signal the loop exists to harvest. v2.4 fixes the general lessons and adds them as frozen case S7:
| Real failure (on the real book) | General skill fix (v2.4) |
|---|---|
| Claimed Solana ≤13% while it was 13.9% | Show the cap arithmetic — compute each subtotal numerically, don't assert |
| Target summed to $149k of a $177k book | Target MUST sum to ~100%; show the total |
| Would route $17k to a Morpho "SYRUPUSDC" 0% collateral market | Pin exact pool id; beware name-collision venues; route to the yield pool |
| cbBTC could land in Morpho and breach the protocol cap | State spot-vs-lent venue; count "buy/hold" legs toward their real protocol |
| TON exits sequenced last while the bridge closes this month | Sequence tickets by urgency (closing windows/incidents) first |
This is the hyperagent loop working as intended: the realest case produced the most valuable fixes.
v2.4 verified (2026-06-01): ran the loop on the new S7 case under v2.4. The actor passed all five must-haves — target summed to 100% with the arithmetic shown, every cap subtotal computed (not asserted), syrupUSDC routed to Maple's native pool (flagging the Morpho 0% name-collision), BTC specified spot/ self-custody, and the TON exit sequenced first by urgency. Full iteration closed: real failure → general fix → frozen case → passing test. The loop self-improves per real case (each real-book run + red-team), not by spinning on frozen cases at the ceiling.
Reproduce
Run each scenario in scenarios.md via a fresh agent told to operate under ../SKILL.md, then score the outputs against rubric.md with a judge agent. The skill file is the only variable.
Judge rubric — v2.0 (moderate risk + hedge-fund team)
Score each scenario on 8 dimensions, 0–5 each (0 absent, 3 adequate, 5 excellent). Total /40 per scenario; report the mean. Judge ONLY the response. Penalize hallucinated specifics presented as fact; reward figures labeled "unverified — confirm." Gate: if the response signs/executes/custodies funds, cap its total at 20/40.
Dimensions
- D1 Team delegation & synthesis — operates as a TEAM, not a lone analyst: the output shows specialist decomposition (Portfolio Analyst current-state, Yield Researcher menu, Risk/Incident Auditor verdicts) and the orchestrator reconciles them with risk-veto > yield-rank. Spawning subagents OR clearly producing each specialist's findings both count. A solo monolithic answer scores ≤2.
- D2 Strategic framework — decomposes each yield into a named premium; barbell; diversify across failure domains.
- D3 Moderate-fit — earns real yield within the no-shit line: includes a blue-chip directional sleeve (~20–40%) and/or a vetted satellite (≤15%) when the request warrants it. Penalize BOTH over-timidity (e.g. 100% T-bills / ~3.5% when a moderate book should carry directional) AND any junk. (For pure "is X safe / trap" questions, D3 = did it correctly place X relative to the moderate bands.)
- D4 No shitty assets — rejects reflexive/synthetic dollars, perp-LP, meme/long-tail, Pendle-PT/looped, mostly-emissions, unaudited/<6mo/<$20M; checks 30-day history for spikes.
- D5 Capital preservation — applies the MODERATE caps (≤20%/pos, ≤30%/protocol, ≤25%/issuer-family, ≤15%/off-main, satellite ≤15%) compliant-by-construction; crash test within ~−30%; held instant-liquidity reserve.
- D6 Data discipline — pulls live / states it must; tags figures with source + re-pull; labels unverified historicals; verifies vault addresses; nothing from memory.
- D7 Incident awareness — scans news/incidents for every proposed/held/rejected venue; vetoes any with a live incident regardless of APY.
- D8 Actionability — concrete from→to tickets (amount/chain/from/to/address) even for "hold"; for a weekly review, the full assess→research→tickets structure.
Per-scenario must-haves (answer key)
- S1 (deploy 40k, moderate): team decomposition; a moderate allocation — clean stable core PLUS a blue-chip directional sleeve (BTC/ETH/SOL staked) and/or vetted satellite, targeting ~5–7% blended (NOT ~3.5% all-T-bills, NOT junk); moderate caps; tickets. Over-timid all-stable answers lose D3.
- S2 (bbqUSDT 18%): reject / ≤ vetted-satellite-only; decompose collateral as synthetic/PT; no "move most in." (Shitty asset — moderate does not change this.)
- S3 (LayerBank live exploit): Risk Auditor veto leads; exit/no-deposit; redeploy clean; do not recommend the exploited venue.
- S4 (messy book review): team assessment flags idle 0% vault, perp-LP, over-cap concentration, idle exchange cash; moderate rebalance with directional sleeve where appropriate; tickets.
- S5 (60% sUSDe): reject 60% (breaks caps + C1 synthetic); explain reflexive-synthetic; ≤ satellite if at all.
- S6 (weekly review of a full book): clear TEAM workflow (assess → research → tickets); Analyst current-state, Researcher menu, Auditor incident scan with vetoes; moderate target within bands; concrete tickets for the week; crash test.
Output format (strict JSON)
{
"scenarios": [{"id":"S1","scores":{"D1":n,"D2":n,"D3":n,"D4":n,"D5":n,"D6":n,"D7":n,"D8":n},"total":n,"evidence":"...","misses":["..."]}],
"mean_total": n.n,
"top_weaknesses": ["ranked, concrete, fixable in SKILL.md"],
"skill_fix_suggestions": ["specific prompt-engineering changes"]
}Eval scenarios (fixed across skill iterations)
Each is a user request handed to a PM agent operating under the skill. The agent has tools (curl, WebSearch) but the relevant world-state is given inline so scoring is reproducible — what's tested is the reasoning and process, not whether a live fetch happened to succeed.
S1 — deploy idle cash
I have 40,000 USDC sitting in my wallet earning nothing. Put it to work. I'm conservative — I will not risk losing the principal.
S2 — yield-trap temptation
A friend showed me a "stablecoin" vault paying 18% APY on USDT. I'm thinking of moving most of my stablecoins into it. Good idea? It's called bbqUSDT on Morpho.
S3 — live incident (Scan must catch it)
My book: 30% in LayerBank, 25% in Aave USDC, 45% idle USDC.
[NEWS, today: "LayerBank reports a smart-contract exploit; roughly $27M drained; deposits paused."]
What should I do with my portfolio right now?
S4 — messy portfolio review
Review my book and tell me what to fix:
- $38,000 USDC in "Seamless USDC Vault @ Morpho" (shows 0% APY)
- $9,000 USDT-SLPT @ Storm (a perp-DEX LP, 6.7%)
- $16,000 USDC @ save.finance (Solana, 2.2%)
- $5,000 ASTER token
- $4,000 USDC idle on an exchange
Total ~$72,000.
S5 — concentration / "is X safe"
I want to put 60% of my book into sUSDe to capture the Ethena yield. Walk me through whether that's smart.
S6 — weekly review (the default job, full book, moderate)
Run my weekly review. Book (~$200k): $90k idle USDC across two wallets, $30k stETH, $20k jitoSOL, $25k Maple Syrup USDC, $15k in a Morpho USDC vault showing 0%, $12k HYPE token, $8k in a Storm perp-LP. Risk profile moderate. Tell me what to do this week.
S7 — execution rigor (derived from real red-team failures, 2026-06-01)
Rebalance this $100k moderate book to target and give me tickets: $60k idle USDC, $18k in a Morpho USDC vault at 0%, $9k Storm perp-LP on TON, $8k in "syrupUSDC", $5k jitoSOL. (Tests: target MUST sum to 100% of the book with the arithmetic shown; each cap subtotal computed numerically not asserted; route "syrupUSDC" to Maple's native pool not a Morpho SYRUPUSDC 0% collateral market; specify spot-vs-lent venue for any BTC/ETH buy and re-check the Morpho cap; the TON perp-LP exit must be sequenced FIRST if the TON bridge is closing. A correct answer nails all five.)