Kitson Backtest Report Critic: Recompute Any Backtest and Check Its Claims (Opus 5.5)
Recompute a backtest from its raw output, check the claims made for it, and see the reasons not to trust it yet.
Add the raw output of a backtest made with any tool: the equity curve (CSV: date and account value) and/or the trade list (CSV). Paste the claims the report or the seller makes, if any (for example "73% win rate, Sharpe 2.8").
Backtest Report Critic recomputes every standard figure in code on your own computer, and states each formula:
- total return, CAGR, max drawdown and how long it lasted, volatility, Sharpe and Sortino;
- win rate, profit factor, expectancy, exposure, the largest winning trade's share of the profit and the longest flat stretch.
It compares them with the claims, then runs eleven checks: too few trades, a result carried by a few trades, a curve too smooth to be real, a Sharpe ratio too high for the evidence, a deflated Sharpe ratio for the number of versions tried, missing costs, no data kept aside, one market period, a Monte Carlo reshuffle of the trade order and a bootstrap of the average trade. Claude Opus 5.5 then explains in plain words the reasons not to trust the report yet, what a fair test would add, and the questions to ask whoever made it.
What makes it different
- Claims against the real figures. Every claim the app can read is set beside the figure it worked out from the raw files: matches, better than the app's figure, or worse.
- The statistics are done in code, with the methods named (including the deflated Sharpe ratio of Bailey and Lopez de Prado), not guessed by the AI.
- Your files and the pasted text stay on your computer. Only the app's figures and the numbers it read from the claims are sent to Anthropic, and you can see exactly what will be sent before anything goes.
- Every number in the explanation is checked in code against the app's own figures, and the result of that check is shown to you.
- It never praises a strategy or tells you to buy anything, including signal services, courses or bots. If the explanation drifts that way, the app catches it and asks for a rewrite.
- The cost is shown before every run, and you set a limit. In our own test on made-up data, a run cost about 9 cents on our own Anthropic account, and one change about 7 cents more.
- A Word report and an Excel file of every figure, claim check, red flag, method and trade.
What you need
- Free Node.js 18+ (Windows, Mac or Linux).
- Your own Anthropic account and API key, with a little credit. A Claude chat subscription does not include API access.
- The backtest's equity curve and/or trade list as CSV, or use the two made-up samples in the zip.
You pay Anthropic directly for what you use, on top of the price of the app.
Educational only. Hypothetical results on historical or made-up data. Past or simulated results do not predict future results. Not financial advice. It checks a backtest's evidence and does not recommend any trade, strategy or product. The app never places trades. Trading involves risk of loss. If we can't get the app running for you, email kit376351@gmail.com within 14 days for a refund of the app price. Kitson Workshop is not affiliated with or endorsed by Anthropic.