Strategy Development Workflow

SkillCommerce & finance

End-to-end strategy development lifecycle from hypothesis to live trading. Use when starting a new strategy project or onboarding to the ML4T workflow.

Use Strategy Development Workflow in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Strategy Development Workflow and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Strategy Development Workflow skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Strategy Development WorkflowStart free

What this skill tells your AI

The instructions your AI receives, as published by ml4t/skills in workflows/strategy-workflow/SKILL.md and read by Ahel’s review.

Strategies fail because developers skip straight to modeling. The correct process spends most time on hypothesis and data, with modeling as a small fraction.

The Problem

A researcher downloads data, fits a model, runs a backtest, and sees a 2.5 Sharpe ratio. They deploy, then lose money because there was no documented hypothesis, no feature validation, no cost model, and no holdout.

The Pattern

WRONG

# Jump straight to modeling - no hypothesis, no validation gates
import lightgbm as lgb

data = load_data()
features = data[["momentum", "volatility", "volume"]]
labels = data["next_day_return"]

model = lgb.LGBMRegressor().fit(features, labels)
predictions = model.predict(features)  # Predicting on training data!

# "Looks great, ship it"
print(f"R2: {r2_score(labels, predictions):.3f}")  # 0.95 - overfitting

CORRECT

# Seven stages with quality gates between each
hypothesis = {
    "mechanism": "Momentum persists due to slow institutional rebalancing",
    "signal": "12-month risk-adjusted return predicts 1-month forward return",
    "kill_criteria": "IC < 0.02 or non-monotonic quintiles",
    "capacity": "ETFs with >$50M daily volume",
}
# Commit term sheet to version control before proceeding

# Stage 2: Data  → fetch + validate
# Stage 3: Features → compute + factor research
# Stage 4: Model → CPCV + model validation
# Stage 5: Backtest → realistic costs
# Stage 6: Paper trade → live data, simulated fills
# Stage 7: Live → small size, slow ramp

Quality Gates

GateConditionFail Action
HypothesisDocumented mechanism, kill criteria, capacity estimateDo not start coding
DataNo gaps > 2 days, point-in-time correct, survivorship-freeFix data pipeline
FeaturesIC > 0.02 (HAC-adjusted), stable across subperiodsDrop factor or redesign
ModelLoss rate < 50%, best config clears the selection boundSimplify model or revisit features
BacktestSharpe > 0.5 net of costs, max DD < 20%Revise sizing or cost assumptions
Paper tradeFills within expected slippage, no execution anomaliesFix execution logic

Guardrails

  • If Sharpe > 2.0 on daily equity data, assume lookahead bias or selection bias until proven otherwise - inspect with ml4t-lookahead-bias and ml4t-deflated-sharpe
  • If in-sample and out-of-sample performance match closely, suspect data leakage
  • If the strategy requires > 20% annual turnover to work, verify cost assumptions with ml4t-transaction-costs
  • If no documented hypothesis exists, stop and write one before any other work

Production Implementation

import asyncio

from ml4t.backtest import Strategy, run_backtest, BacktestConfig  # MyStrategy subclasses Strategy
from ml4t.live import LiveEngine, AlpacaBroker, AlpacaDataFeed

results = run_backtest(
    prices=prices, signals=signals, strategy=MyStrategy(), config=BacktestConfig()
)

async def trade_live():
    broker = AlpacaBroker(api_key, secret_key, paper=True)
    feed = AlpacaDataFeed(api_key, secret_key, symbols=["SPY"], experimental=True)
    engine = LiveEngine(MyStrategy(), broker, feed)
    await engine.connect()
    await engine.run()

asyncio.run(trade_live())

Checklist

  • Hypothesis documented in term sheet before any code
  • Data validated for gaps, survivorship bias, point-in-time correctness
  • Features pass IC significance test with HAC standard errors
  • Model validated via CPCV; loss rate < 50% (PBO: ml4t-backtest-overfitting)
  • Backtest includes realistic transaction costs
  • Deflated Sharpe computed across all trials (ml4t-deflated-sharpe)
  • Paper trading completed for minimum 4 weeks
  • Kill criteria defined and monitoring configured before going live

Signals

GitHub stars
22
Forks
11
Last commit
Oct 2026
Advanced
Item type
skill
Key
ml4t-strategy-workflow
Source
github.com/ml4t/skills