SCIENTER

Methodology

How a score is computed

Every number published on this site is derived from public data by code in this repository. This page is the specification: enough to recompute a scorecard independently and get the same answer, and enough to know when not to trust one.

Last updated: 2026-09-02

Methodology version: 1.0

1

What feeds the model

Two scoring systems run under this methodology, and they are different enough that conflating them would mislead. The track-record score is in production and is what every scorecard and leaderboard on this site displays. The token risk score is specified and is not in production; section 4 says so in detail.

Track-record score — inputs

Computed by ghostcopy/hl_ingest/ from public Hyperliquid data. No credential, no private feed, no order path.

Trader populationHyperliquid leaderboard snapshot. The entire scanned population, not the shortlist that received a scorecard.
Return seriesCumulative pnlHistory differences, not account-value changes — otherwise a deposit registers as a return.
Vault populationThe public vaults index, plus per-vault detail for anything judged.
BenchmarkBTC and ETH hourly candles.
Observation cadenceWhatever the venue reports, typically hourly. Every statistic is per-observation; the annualised figure on a scorecard is labelled and used for nothing.

Token risk score — specified inputs

The feature set the classifier is specified against: token supply distribution and mint authority, liquidity-pool concentration and lock state, deployer wallet age and prior deployment history, honeypot behaviour under simulated sell, and post-launch price and liquidity collapse patterns. These are the inputs. They are not evidence that the model works, and nothing on this site currently serves a score derived from them.

2

The statistics, and the one idea behind them

Sorting 44,000 traders by profit and publishing the top hundred is running 44,000 trials and reporting the luckiest. The maximum of N zero-skill trials grows like sqrt(2 ln N), so the top of any large leaderboard looks skilled whether or not anybody in it is. Every public crypto leaderboard sets that expected maximum to zero implicitly, by never mentioning it.

The Deflated Sharpe Ratio raises the benchmark to that expected maximum before asking whether a track record beats it.

Bailey, D. & López de Prado, M. (2014). “The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality.” Journal of Portfolio Management, 40(5), 94–107.

The Probabilistic Sharpe Ratio underneath it uses the Mertens standard error, which corrects the Sharpe estimator for non-normal returns. Kurtosis is passed as full kurtosis, not excess — passing excess understates the tail penalty by exactly 3 and flatters every fat-tailed record.

The expected-maximum term is computed against the entire scanned population, not the cohort that received a scorecard. Deflating against the shortlist would defeat the correction.

Implemented in ghostcopy/hl_ingest/stats.py (standard library only) and ported to apps/dashboard/lib/scorecard-math.ts. The two are pinned to each other by tests on both sides, so the number the browser renders and the number the engine computes cannot drift.

Two guards worth naming because they change results. Returns are floored at a $1,000 account value — a drained account otherwise divides by almost nothing and reports 4,000% — and the floor is flagged on the record when it binds. Gaps longer than three days split the series rather than being compressed into one enormous return across a period nobody observed.

3

What has actually been measured

This section exists to be checked. It reports only measurements that were run, and it names the ones that were not.

Measured

  • The first full judgment run: more than 23,000 traders scanned, 150 judged in depth, zero clearing even the loosest skill tier. That is the honest state of copy trading under this methodology, and the leaderboard reports it every time it loads.
  • Our own signal engine, 90 days, live: a Sharpe ratio of −0.14, fees included. Negative. Published in full at the post-mortem. The detector we built to find false positives in other people's records found one in ours.
  • Provider accuracy: third-party intelligence providers scored against ground truth on questions with a checkable answer. Published at provider accuracy, which labels its own data mode and publishes zero rows rather than a fabricated board when nothing has been measured.

Not measured

No precision, recall, F1 or confusion matrix has been computed for any Scienter output, including token risk. There is no labelled evaluation set in this repository against which such a figure could be produced. If you have seen an accuracy percentage for rug detection attributed to us — including in our own bot copy, where one appears — it is not sourced to a measurement and should not be relied on. Removing it from those surfaces is tracked work.

Whether our verdicts predict anything is also unmeasured. The study that would answer it — do records we rate highly go on to outperform records we do not — is blocked on retaining per-run determinations, which the engine does not yet do. It is the most valuable validation this product lacks and it is on the roadmap rather than quietly absent.

4

Limitations, edge cases, and what it cannot detect

A methodology page that lists only strengths is marketing. These are the conditions under which a Scienter number is weak, wrong, or absent.

Not in production

  • Token risk scoring. Specified, partly built on a separate workstream, not merged and not deployed. The API answers 503 with the reason rather than 404, because 404 would be a claim about the token when the truth is a claim about our deployment. No score is served.
  • Wallet clustering and MEV attribution. Same status. A scorecard describes an address, not the operator behind it, and cannot tell you that two addresses are one person.
  • On-chain proof of published scores. The AlphaLedger contract is written and covered by tests but has never been deployed. Wherever this site says a number was published, it means published — not verified on-chain.

Structural limits of the track-record score

  • It is a measurement of the past, not a forecast. The Deflated Sharpe asks whether a record is distinguishable from luck. A record that passes is not thereby expected to repeat.
  • Short records cannot be judged at all. Below the sample-size floor the honest output is no score. A three-trade record is refused rather than rated, which reads as a gap and is a correctness feature.
  • One venue, one book. Only positions visible on Hyperliquid are counted. A trader hedged on another venue, or trading size elsewhere, will be measured on a fragment of their real exposure and the score will be confidently wrong about them.
  • No wash-trading or manipulation detection. Fills are taken as reported by the venue. A record constructed by trading against oneself is not identified as such.
  • Deposits and withdrawals are handled, custody is not. Returns are deposit-proof, but nothing verifies that the account belongs to the person claiming it.
  • Upstream data is trusted. If the venue's history is wrong, so is the score. There is no independent reconstruction of the order book.
  • Published scores are not reproducible after the fact — yet. Each run currently overwrites the last snapshot, so a scorecard published last month cannot be regenerated today from retained inputs. This is a real gap, it is on the roadmap, and until it closes no claim of historical replay should be made on our behalf.

5

How evidence is cited

Every published judgment carries a link back to the surface that shows the underlying record. An assertion without one is not something we publish, and if you receive a Scienter claim through any channel — bot, alert, social post, API — with no evidence link, treat it as unverified.

SubjectThe address, vault or token the determination is about, in full. Never an abbreviated form alone.
Evidence linkA canonical URL on this host — /trader/<address>, /vault/<address> or /wallets/<address> — showing the record the judgment was computed from.
Methodology versionThe version this page carries at the time of publication (currently 1.0).
Data modeWhether the figures are measured or drawn from a replay fixture. Never defaulted to 'measured'; a demonstration that does not announce itself is a fabrication.
Observation windowThe period the record covers, and the sample size behind it.

To cite a section of this page back at us, use its anchor — this one is #citation. Anchors are named rather than numbered so that inserting a section above does not silently repoint an existing citation.

6

Update cadence

Scorecard recomputationOn each engine run. Cadence follows the deployment; see /status for whether the engine is currently reachable at all.
This documentVersioned, with an effective date. A change to how a number is computed is a version bump; a clarification to how it is described is not.
Never revised in placeA published version's meaning is fixed. Changing what a score meant after scores were attributed to it would rewrite history, and the same rule already governs compliance/DISCLAIMER.md.
RoadmapChanges to what is measured are announced on the roadmap.

7

Questions, and challenges

Methodology questions, disputes about a published determination, and reports that a number on this site is wrong all go to the same address. A challenge with a reproduction is the most useful thing anyone can send us, and it is answered.

[email protected]

Security vulnerabilities go to the bug bounty programme instead. The data behind these numbers is catalogued at data sources.