TRUST LAYER FOR WEB3

Composite scores for protocols - so you know why a name looks strong before you commit size.

Predictive Power | Stability | Robustness | Logic | Diversity

MARKET

What EVALUATOR scores - and what it refuses to pretend.

Most WEB3 dashboards sell vibes: green candles, follower counts, TVL screenshots. EVALUATOR scores protocols as factors - measurable signals with holdout behavior, not as marketing pages. The output is a 0-100 composite plus a breakdown you can argue with.

Who this is for

  • Operators comparing L1 / L2 / DeFi names before size
  • Funds that need a shared score language across analysts
  • Builders shipping factors who want IC + stability first
  • Anyone tired of "trust me bro" research threads

What it is not

  • Not price prediction or financial advice
  • Not a replacement for legal / security diligence
  • Not a claim that high score = guaranteed alpha
  • Not a black-box AI that cannot show its work

SCORING MODEL

Composite score. Not a magic net.

First version is a weighted composite across 4-5 dimensions you can compute separately. Show the total and the breakdown - otherwise you never know why a "good" factor died in live P&L.

35-40%

Predictive Power

IC + RankIC (PPS, beta = 0.5). IC Information Ratio = mean(IC) / std(IC) across periods. The alpha base - if this is weak, nothing else saves the score.

20-25%

Temporal Stability

Rank correlation across adjacent windows + share of periods where IC sign holds. Unstable = high turnover and random P&L even when average IC looks fine.

15-20%

Robustness

Shift windows, add noise, drop top 1% extremes, change 1d to 5d horizon. If the score collapses - the factor is fragile and unfit for size.

10-15%

Logic / Interpretability

Economic sense check. Start with LLM-as-judge + checklist: no look-ahead, no future leak, not pure noise. Later - expert rubric and human review.

~10%

Diversity

Optional on MVP. Correlation vs benchmark factors (momentum, value, size). Penalize redundancy so the book does not stack the same bet five times.

Normalize each dimension (z-score or min-max) -> weighted sum -> clip to 0-100.

MVP shortcut: Predictive Power + Temporal Stability alone already beats "look at Sharpe after a backtest".

How to read a score

85-100 Strong composite. Still check red flags and sample length.
70-84 Usable with caveats. Watch stability and 2022-style holdouts.
55-69 Mixed. Fine for research, risky as a core position thesis.
below 55 Weak or noisy. Treat as exploratory until the signal improves.

METHOD

Cheap data. Hard splits. Honest holdouts.

The method is deliberately boring. Free or local data, liquid universe, fixed time splits, and a report that shows where the factor worked - and where it failed.

Data stack

  • Prices / volume - yfinance (US) or local CSV for crypto OHLCV
  • Universe - 100-300 liquid names (majors + liquid alts, no illiquid dust)
  • Frequency - daily as default; weekly for stress checks
  • Split - train 2015-2019 | validation 2020-2021 | holdout 2022-now
  • Target - forward return 1d / 5d / 20d
  • Cleaning - corporate actions, volume outliers, delistings, date look-ahead

How alphas enter the evaluator

  1. Manual formulas + simple Python functions
  2. Genetic search / LLM-generated formulas
  3. Agent signals (same logic, different data feeds)

Agent / product eval (PDF tasks + rubric) is production-eval territory - not the first build.

01

Ingest

Load prices, align calendars, drop illiquid names, mark look-ahead risks before any factor runs.

02

Compute

Run IC / RankIC / ICIR on validation, then freeze rules before you touch the holdout window.

03

Stress

Shift windows, add noise, change horizon. Keep only signals that survive without babysitting.

04

Report

Publish 0-100 composite + breakdown + red flags. No score without a reason trail.

TOKEN

Access layer - not a casino chip.

The token is designed to gate deep reports, batch scoring, and governance weight - not to invent a new yield narrative. Utility first. Speculation is optional and unsolicited.

ACCESS

Report unlocks

Holders unlock full IC tables, robustness packs, and historical rescores for covered protocols.

STAKE

Signal staking

Stake against a published score. If the holdout breaks the claim, stake gets slashed toward the dispute pool.

GOVERN

Weight votes

Token weight votes on model parameters: dimension weights, universe filters, and red-flag thresholds.

BUILD

Factor bounties

Pay bounties for new factors that clear IC + stability gates. Rejected factors stay public as negative examples.

Design rule: if a feature works without the token, keep it free. Token pays for depth, speed, and governance - not for looking at a homepage score.

No emission schedule theater in v1. Ship utility surfaces before emission charts.

GOVERNANCE

What the UI must shout - and who decides the rules.

Governance is not a Discord poll about logo colors. It is control over scoring rules, dispute handling, and what counts as a red flag the product cannot soft-pedal.

Hard red flags

high IC, unstable sign
strong correlation with momentum
broke on 2022 holdout
look-ahead / future leak risk
sample too short for stability
pure narrative, no measurable IC

On-chain / off-chain split

  • On-chain - parameter votes, stake, dispute outcomes
  • Off-chain - heavy compute, dataset builds, human review queues
  • Published - model version hash + score snapshot for every release

Dispute flow

  1. Anyone can flag a score with evidence (window, IC table, leak claim)
  2. Stakers bond on uphold vs revise
  3. Review window publishes a decision + changelog
  4. Losing side funds the next robustness run

If the goal is to test the idea fast - ship a factor Alpha Evaluator with IC + stability. Metrics are clear. Data is free. Results are checkable. Expand to quant depth or agent eval later.

START WITH IC + STABILITY

ROADMAP

Not a platform. One sharp tool - then layers.

MVP #1 - 1-2 weeks

Local Alpha Evaluator

  • Accept a factor (formula or value column)
  • Compute IC / RankIC / ICIR on holdout
  • Stability by year / quarter
  • Composite 0-100 + red flags

Jupyter + Streamlit or CLI. No accounts. No paywall. No "AI platform".

score + breakdown IC table by period factor distribution warnings
MVP #2 - only if #1 is used

Batch compare

  • Score 50-200 candidates
  • Factor vs factor compare
  • CSV export
  • 1-2 robustness tests

Skip first: full portfolio backtest, custom ML scorer, alpha marketplace, 6-domain agent eval.

Now

Landing + protocol carousel + hot deal flow scoring proxies.

Next

Local evaluator CLI / Streamlit with IC tables and red-flag engine.

Later

Batch compare, published model versions, token-gated deep reports.

Edge

Dispute staking + governance over weights - only after the score is trusted.

FAQ

Straight answers.

Is a high score a buy signal?

No. A high score means the composite factors look strong on the published method. It does not price liquidity, legal risk, or your portfolio constraints.

Why show the breakdown at all?

Because a single number without IC / stability / robustness is how people ship factors that die in the first live month. Breakdown is the product.

How often do scores refresh?

Protocol carousel scores refresh when the model version changes. Hot Project deal scores refresh with the fundraising scrape (see updated timestamp on that page).

Can I submit my own factor?

That is the MVP #1 path: local evaluator first. Public submissions come after we can score batches without turning the site into a spam sink.

What about early-stage / pre-token projects?

They use investor / raise / sector proxies until IC history exists. Expect wider confidence bands and more red flags - that is honest, not broken.

Where does governance start?

After the score language is stable. Early governance without a trusted model is just costume politics.