Who this is for
- Operators comparing L1 / L2 / DeFi names before size
- Funds that need a shared score language across analysts
- Builders shipping factors who want IC + stability first
- Anyone tired of "trust me bro" research threads
Composite scores for protocols - so you know why a name looks strong before you commit size.
FIND A PROJECT
MARKET
Most WEB3 dashboards sell vibes: green candles, follower counts, TVL screenshots. EVALUATOR scores protocols as factors - measurable signals with holdout behavior, not as marketing pages. The output is a 0-100 composite plus a breakdown you can argue with.
SCORING MODEL
First version is a weighted composite across 4-5 dimensions you can compute separately. Show the total and the breakdown - otherwise you never know why a "good" factor died in live P&L.
IC + RankIC (PPS, beta = 0.5). IC Information Ratio = mean(IC) / std(IC) across periods. The alpha base - if this is weak, nothing else saves the score.
Rank correlation across adjacent windows + share of periods where IC sign holds. Unstable = high turnover and random P&L even when average IC looks fine.
Shift windows, add noise, drop top 1% extremes, change 1d to 5d horizon. If the score collapses - the factor is fragile and unfit for size.
Economic sense check. Start with LLM-as-judge + checklist: no look-ahead, no future leak, not pure noise. Later - expert rubric and human review.
Optional on MVP. Correlation vs benchmark factors (momentum, value, size). Penalize redundancy so the book does not stack the same bet five times.
Normalize each dimension (z-score or min-max) -> weighted sum -> clip to 0-100.
MVP shortcut: Predictive Power + Temporal Stability alone already beats "look at Sharpe after a backtest".
METHOD
The method is deliberately boring. Free or local data, liquid universe, fixed time splits, and a report that shows where the factor worked - and where it failed.
Agent / product eval (PDF tasks + rubric) is production-eval territory - not the first build.
Load prices, align calendars, drop illiquid names, mark look-ahead risks before any factor runs.
Run IC / RankIC / ICIR on validation, then freeze rules before you touch the holdout window.
Shift windows, add noise, change horizon. Keep only signals that survive without babysitting.
Publish 0-100 composite + breakdown + red flags. No score without a reason trail.
TOKEN
The token is designed to gate deep reports, batch scoring, and governance weight - not to invent a new yield narrative. Utility first. Speculation is optional and unsolicited.
Holders unlock full IC tables, robustness packs, and historical rescores for covered protocols.
Stake against a published score. If the holdout breaks the claim, stake gets slashed toward the dispute pool.
Token weight votes on model parameters: dimension weights, universe filters, and red-flag thresholds.
Pay bounties for new factors that clear IC + stability gates. Rejected factors stay public as negative examples.
Design rule: if a feature works without the token, keep it free. Token pays for depth, speed, and governance - not for looking at a homepage score.
No emission schedule theater in v1. Ship utility surfaces before emission charts.
GOVERNANCE
Governance is not a Discord poll about logo colors. It is control over scoring rules, dispute handling, and what counts as a red flag the product cannot soft-pedal.
If the goal is to test the idea fast - ship a factor Alpha Evaluator with IC + stability. Metrics are clear. Data is free. Results are checkable. Expand to quant depth or agent eval later.
START WITH IC + STABILITYROADMAP
Jupyter + Streamlit or CLI. No accounts. No paywall. No "AI platform".
Skip first: full portfolio backtest, custom ML scorer, alpha marketplace, 6-domain agent eval.
Landing + protocol carousel + hot deal flow scoring proxies.
Local evaluator CLI / Streamlit with IC tables and red-flag engine.
Batch compare, published model versions, token-gated deep reports.
Dispute staking + governance over weights - only after the score is trusted.
FAQ
No. A high score means the composite factors look strong on the published method. It does not price liquidity, legal risk, or your portfolio constraints.
Because a single number without IC / stability / robustness is how people ship factors that die in the first live month. Breakdown is the product.
Protocol carousel scores refresh when the model version changes. Hot Project deal scores refresh with the fundraising scrape (see updated timestamp on that page).
That is the MVP #1 path: local evaluator first. Public submissions come after we can score batches without turning the site into a spam sink.
They use investor / raise / sector proxies until IC history exists. Expect wider confidence bands and more red flags - that is honest, not broken.
After the score language is stable. Early governance without a trusted model is just costume politics.