The track record, in the open
Every analysis Argustr publishes is scored deterministically against real market prices at its horizon due date. No backtests, no cherry-picking: wins and losses, with confidence intervals.
Track record
forward-tested
Accuracy measured against real prices after the fact, never backtests.
Hit rate50.9%95% CI 38–64%
PSR0.57P(Sharpe>0)
Sharpe0.02realized
Brier0.53calibration
By horizon
1 month 52.9%3 months 0.0%
Edgenot yet distinguishable from zerot = 0.50 · n = 21
Cumulative alpha vs-13.0%
Running sum of each settled call's return minus its market (S&P 500 for stocks, Bitcoin for crypto) over the same window, oldest first · sample too small for this number to mean anything yet.
Verdictssettled against real prices
Confidence calibration
When the desk states a confidence level, how often is it actually right? A calibrated system earns the right to be trusted.
says 32%
is right 0%
says 40%
is right 25%
says 41%
is right 25%
says 42%
is right 33%
says 44%
is right 33%
says 46%
is right 55%
says 48%
is right 55%
says 50%
is right 55%
says 51%
is right 55%
says 52%
is right 55%
says 55%
is right 55%
says 59%
is right 55%
says 60%
is right 55%
says 65%
is right 55%
says 68%
is right 55%
says 70%
is right 63%
says 74%
is right 63%
says 75%
is right 63%
says 76%
is right 63%
says 82%
is right 63%
says 83%
is right 63%
says 85%
is right 63%
How this is measured
- Each analysis states a target stance (LONG / SHORT / NEUTRAL) and a horizon with a due date. The record is a paper book: one unit long per LONG, one unit short per SHORT, no position for NEUTRAL. If you cannot short, read SHORT as “do not buy; trim if held” and NEUTRAL as “no signal”.
- At the due date, the realized price is compared to the price at inference time. No AI is involved in scoring.
- A NEUTRAL call is judged against the asset's own volatility: it is correct when the move stays inside the band, and it never earns or loses P&L (no position).
- Returns are also benchmarked against the market (beta-adjusted alpha), and the stated confidence is scored for calibration (Brier).
- Verdicts are append-only: once settled, a result is never edited. Methodology changes are versioned, never retroactive.
- When the desk is wrong, an attribution engine records whether the cause was a risk it had identified or a blind spot, and blind spots feed back into future analyses.
Get the same desk working on your tickers.
Start free: 3 analyses per monthNo card required. See pricing