How Philidor Scores a Vault
This is a methods reference. It explains how Philidor computes a vault score. The score describes the evidence available at scoring time. It is not a forecast or an investment recommendation.
A Philidor score answers a narrow question. How strong was the evidence when we scored this vault?
The answer runs from 0 to 10. Higher is better. But a score of 8 does not mean an 80 percent chance of safety. The scale compares evidence across vaults. It does not state a probability.
Two rules shape the score.
First, one serious weakness can limit the whole result. Strong evidence in three areas cannot cancel unaudited code or weak collateral. Second, missing evidence never counts as good news. We use a cautious default, a neutral value, or a failed run. The choice depends on what is missing.
Why publish the method
A rating is more useful when readers can see what the number means. In its 2025 guide, S&P says its rating methods are transparent and free to read.1 Registered rating firms in the United States must also publish key parts of their procedures.2
By 2026, DeFi risk teams had taken several approaches to the same problem. DeFiSafety publishes an item-by-item process review.3 L2BEAT shows separate risk criteria instead of hiding them inside one number.4 Both approaches make the judgment visible. A reader can see which fact drove the result.
Economic-risk teams answer a different question. Chaos Labs publishes simulation methods for setting market parameters.5 That work connects market behavior to choices such as collateral limits. It is less concerned with producing one common reading across many vaults.
Prior work established how to publish review criteria, show distinct risk flags, and disclose economic models. This paper adds a single vault score with a fixed calculation. It combines current exposure, platform, control, and incident evidence. The public method lets a reader trace the result back to those inputs and the time of the scoring run.
The score in one line
The score starts with four components. We call them vectors. Each uses the same 0-to-10 scale.
| Vector | Weight | What it asks |
|---|---|---|
| Asset | 0.3 | What does the vault hold, and how strong is each asset? |
| Platform | 0.3 | How strong is the protocol, its code record, and its dependencies? |
| Control | 0.2 | Who can change the system, and how quickly can they act? |
| History | 0.2 | What serious events have affected this vault? |
The raw score is:
0.3 × Asset + 0.3 × Platform + 0.2 × Control + 0.2 × History
We round that result to two decimal places. We then apply the cap stack. A cap is a ceiling. It can pull the score down, but it can never push it up.
Asset: what the vault holds
The asset vector reads every holding reported by the vault adapter. An adapter is the connector that reads a protocol's positions. The scorer removes duplicate holdings by chain and address. It then scores each exposure through the asset registry.
Different assets need different tests. A native token is scored on liquidity and volatility. Each has a 50 percent weight. A bank-backed stablecoin needs a different test. Custody carries 25 percent. Reserve transparency carries 20 percent. Redeemability, peg stability, and liquidity each carry 15 percent. Governance carries the remaining 10 percent.
Private-credit tokens use another mix. Credit quality carries 20 percent. The default rate and custody each carry 15 percent. The rest covers reserve reporting, recovery, manager concentration, governance, liquidity, and redemption terms.
The point is simple. We do not grade every asset with one generic checklist.
The full category matrix is available through /v1/methodology.
Unknown and unreviewed assets
An asset with no registry entry scores 2.0. The same score applies when the
adapter cannot resolve its address. We do not guess which asset it meant.
Review status also sets a ceiling:
| Review status | Highest possible asset score |
|---|---|
| Reviewed | 10.0 |
| Provisional | 9.0 |
| Unreviewed | 7.9 |
A dated instrument gets special treatment after maturity. Pre-maturity market
evidence is no longer current. The scorer drops those stored dimensions and
observations. It also removes the lockup and yield-bearing adjustments. The
asset is then capped at 7.9. This is not a claim that the asset failed. It
means the old assessment has expired.
Hard-fail flags
Some facts are serious enough to set their own cap. The scorer calls them hard-fail flags.
| Flag | Asset cap |
|---|---|
| Active depeg | 1.0 |
| Confirmed sanctions exposure | 0.0 |
| Redemptions paused | 2.0 |
| Single-signer upgrade authority | 3.0 |
| High use of the protocol's own token as collateral | 4.0 |
| Missing proof of reserves | 4.0 |
| No recent reserve attestation | 5.0 |
| Unaudited token contract | 4.0 |
If several flags apply, the lowest cap wins.
Stablecoin backing
For a monitored stablecoin, Philidor can look through the coin to its backing.
The scorer grades each backing sleeve. It then weights the sleeve scores by
their share of the reserves. An unknown or unresolved sleeve scores 2.0.
That weighted backing score sets a ceiling for the stablecoin's asset score. The ceiling rises smoothly with backing quality:
| Backing score | Asset-score ceiling |
|---|---|
8.0 or higher | 10; no backing cap |
7.85 to below 8.0 | Rises from 7.9 to 10 |
5.0 to below 7.85 | Rises from 4.9 to 7.9 |
Below 5.0 | Rises from 0 to 4.9 |
Only the latest complete snapshot may bind. A partial or disputed snapshot cannot set the cap. The last complete snapshot remains in force. The snapshot must also cover most of the backing. If 50 percent or more is unknown or fail-closed, Philidor treats that as a monitoring gap and applies no backing cap.
A complete snapshot older than 24 hours gets an uncertainty drag. Its ceiling
falls by up to 0.5 points. The rule never raises a lower ceiling. When the
ceiling is already below 4.9, it stays below 4.9.
Old data and risky construction
Old evidence loses weight. A stale dimension is multiplied by 0.92. An
expired dimension is multiplied by 0.75. If more than half of the weighted
dimensions are stale, the asset score cannot exceed 8.0.
Loan-to-value settings also matter. Loan to value, or LTV, is the share of
collateral value that a borrower may borrow. High LTV leaves less room for a
price fall. At an LTV of 0.85, the asset score loses 2.5 percent. At 0.90,
it loses 5 percent. At 0.95, it loses 10 percent.
The scorer then looks at the portfolio. For a multi-asset vault, it uses the
Herfindahl-Hirschman Index, or HHI, to measure concentration. An HHI above
0.5 subtracts 0.5 points. An HHI above 0.8 subtracts 1.0 point.
Wrong-way risk subtracts another 1.0. Wrong-way risk means the exposure and
the system behind it may fail together. Category correlation is shown to
readers but does not change the score.
Platform: what the vault runs on
The platform vector starts with three readings. They cover age, audits, and strategy type. The scorer averages them.
Protocol age uses a curve called the Lindy score:
10 × (1 - e^(-days / 365))
The score rises quickly in the first years. It then flattens. New code has had less time to face real use. Older code has survived more conditions. Age alone does not prove safety, so it is only one part of the platform vector.
The audit score is 0 when no audit is recorded. With at least one audit, it
starts at 4. A public audit contest adds 2.0. An audit by an established
firm adds 1.5. A recognized firm adds 1.0. An unknown or unclassified firm
adds 0.5. The audit score cannot exceed 10.
This rule values audit quality, not just audit count. It does not grade the declared scope or findings. Those details are shown separately. Missing scope metadata does not change the audit score.
Protocol incidents
A recent major incident can cap the platform score. An unresolved incident sets these ceilings:
| Time since incident | Platform cap |
|---|---|
| Less than 30 days | 2 |
| 30 to 89 days | 5 |
| 90 to 179 days | 8 |
A resolved major incident is treated less harshly. The cap is 5 for the
first 30 days. It is 8 from day 30 through day 89. Minor incidents do not
change this vector.
Dependencies
A vault inherits risk from the services below it. The platform score therefore uses the weakest dependency.
A dependency score of at least 7.5 applies a 0.95 safety factor. A score
from 5.0 to 7.49 applies 0.8. A lower score applies 0.5. Each extra
dependency after the first adds a 3 percent discount. That discount stops at
0.85.
The internal exchange rate of a liquid-staking token is not treated as an
external service. The scorer excludes the dependency named exchange_rate.
Yield quality for money markets
Money-market vaults get one more platform check. It applies to Aave, Spark, Compound, and Morpho. It does not apply to liquidity pools or aggregators. Their reward data has a different meaning.
The check reads up to 14 days of net yield, base yield, and utilization. It needs at least five usable snapshots across three days. It uses an hourly, time-weighted median. Long-lived conditions therefore count more than brief spikes.
Two flags can lower the platform score:
- Thin exit liquidity. Median utilization above 95 percent caps the
platform score at
8.0, then subtracts1.0. Once active, the flag clears only when utilization falls to 90 percent or below. - Reward-dependent yield. A real-yield share below 50 percent caps the
platform score at
9.0, then subtracts0.5. The rule also requires at least $1 million of value and a net annual yield of at least 1 percent.
Low utilization is reported as idle capital. It does not lower the score.
Control: who can change the system
The control vector asks two questions. Who has authority? How much warning do depositors get before a change?
For EVM contracts, immutability scores 10. A liquidity pool with no pause
power scores 9. In that case, governance can change fees but cannot move user
funds.
Most other EVM contracts are scored by their timelock. A timelock is the delay between proposing a change and executing it. It gives users time to react.
| Timelock | Control score |
|---|---|
| At least 168 hours | 9 |
| At least 72 hours | 8.5 |
| At least 48 hours | 8 |
| At least 24 hours | 6 |
| At least 6 hours | 4 |
| More than 0 but less than 6 hours | 2 |
| No delay | 1 |
Missing permission data scores 5. An unknown timelock also scores 5. Five
means unknown here. It does not mean safe.
Solana uses a separate control test. An upgrade authority held by one ordinary
key scores 1. Revoked upgrade authority scores 7 when market parameters
can still change. It scores 8 when the observed market settings are also
immutable. We stop at 8 because emergency and global-admin paths may remain.
A multisig requires several approved signatures. An effective multisig can
score 3.5, 5, or 6.5. The result depends on its signature threshold. The
multisig must control every live program in the vault. Its own configuration
authority must also be disabled.
A native multisig delay can lift that base score. The scorer compares the base
with the matching EVM timelock score. It takes the higher value, adds 0.5,
and caps the result at 9.
A single-key market admin can cap the result at 4. A single-key emergency
signer can cap it at 6. Active emergency mode caps it at 3. Missing Solana
program evidence scores 5.
History: what has happened
The history vector starts at 10. It then subtracts penalties for published,
non-retracted events linked to the vault.
Only realized adverse event types count. They include incidents, bad debt,
debt socialization, emergency pauses, shutdowns, and proxy changes. Governance
and cap-management events do not count here. Their risk belongs in the control
and concentration readings. Philidor's own RatingChange events are also
excluded. A score must not react to its own output.
Each Critical event subtracts 1.5 points before time decay. The decay weight
is 1.0 inside 30 days. It falls to 0.6 through day 89, 0.4 through day
179, and 0.2 through day 364. Events older than one year leave the decaying
query.
A Critical event with confirmed loss also carries a permanent penalty of
1.0. These lifetime penalties stop growing at 4.0 points.
Warnings use the same time buckets. Their combined count passes through a
logarithm. This makes the first warnings matter more than the hundredth repeat
of the same pattern. The base multiplier is 2.0.
A vault earns a 0.5 clean-streak bonus when two conditions hold. It must have
no Critical event in the last 90 days. It must also have no lifetime Critical
loss event.
History can cap the whole vault score. A history score below 4 sets a total
cap of 4.9. A history score below 7 sets a total cap of 7.5.
The cap stack
The raw four-part average now passes through the cap stack. The order is fixed. Every step uses the lower of the current score and the new ceiling.
1. Asset-quality drag
The scorer first builds an asset-quality anchor. This is the weighted asset score before the market's LTV penalty. It answers a basic question: how strong is the collateral itself?
An anchor below 5.0 caps the vault at 4.9. An anchor from 5.0 to below
7.85 caps it at 7.9. From 7.85 to below 8.0, the ceiling rises smoothly
from 7.9 to 10. An anchor of at least 8.0 adds no cap.
The smooth final band prevents a tiny asset move from causing a large tier
jump. A change from 7.99 to 8.00 should not create a two-point rating move.
One contamination rule is stricter. Suppose one exposure is at least 5 percent
of the vault. If that exposure has a binding fail-safe cap, the whole vault is
capped at 4.9. A serious unknown cannot be averaged away by safer assets.
2. Unaudited protocol version
A protocol version with no audit history cannot score above 4.9. Strong
assets, controls, and history cannot lift unaudited code past that ceiling.
3. History
The history ceiling comes next. It is 7.5 or 4.9 when the history vector
crosses the thresholds described above.
4. Overrides and active incidents
A documented vault override can add another ceiling. The stored override uses a 0-to-100 scale. The scorer divides it by ten before applying it.
An unresolved, protocol-wide incident can also carry a ceiling. That value comes from the incident record. The code does not invent one. The incident must match the vault's protocol. A version or chain on the incident narrows the match. If several incident ceilings apply, the lowest one wins.
Curator conduct is disclosed separately. It does not change the production vault score. The methodology endpoint stated this rule on September 4, 2026.
How missing data is handled
There is no single default for every missing field. Each gap has its own safe treatment.
| Missing evidence | Treatment |
|---|---|
| Asset not found, or address unresolved | Asset score 2.0 |
| Protocol metadata absent | Platform score 0 |
| Audit history empty | Audit score 0, plus total cap 4.9 |
| Permission data absent | Control score 5 |
| Timelock unknown | Control score 5 |
| No qualifying event rows | History starts from the clean default of 10 |
| Stablecoin backing mostly unobserved | No backing ceiling binds |
| Curator identity absent | No score effect; coverage is shown as absent |
The history default needs care. A 10 means the query found no qualifying
events. It does not prove that no event happened outside Philidor's records.
The whole run can fail too. If there is no active asset methodology, the scorer
stops. It does not write zero scores. If the number of scored vaults differs
from the eligible count, the run is marked fail_closed. The run is marked
degraded when the asset cache is more than one hour old. The same happens
when more than half of scored vaults contain only provisional or unreviewed
exposures.
Shutdown and frozen vaults stay in the scoring set while they remain active. They may still hold user funds. Their evidence must therefore keep changing with the record.
From a score to a public rating
The raw score maps to three public tiers:
| Score | Tier |
|---|---|
8.0 to 10.0 | Prime |
5.0 to 7.99 | Core |
Below 5.0 | Edge |
The database can store a small score movement without publishing a new tier. This separation keeps the score honest while limiting visible noise.
An upgrade must hold for 6 hours before the public tier changes. Most
downgrades must hold for 24 hours. A deeper fall moves at once. The immediate
floor is below 7.75 when leaving Prime. It is below 4.75 when leaving Core.
Certain risk events also move at once. They include a change to an incident cap or hard-fail flag. A methodology or override change also moves at once. A reversal during the waiting period resets the clock.
Material changes enter rating_history. A tier change always qualifies. So do
incident caps, hard-fail flags, and methodology changes. A within-tier score
move qualifies at 0.5 points.
The history row and the database notification are written in the same transaction as the score update. They either succeed together or fail together. Public feed events are created only after that transaction commits.
A downgrade into Edge becomes Critical. Other downgrades become Warning.
Upgrades become Info. A new incident cap or hard-fail flag becomes Critical
when it lowers the score. A within-tier move reaches the public feed only at
1.0 point. At most 50 rating events are projected from one scoring run.
How to check a score
The scorer has no random branch. Fixed inputs and a fixed clock produce the same result. That is what deterministic means here.
Each scoring run records its start time and status. It records the code version and the asset-methodology version. It also records a hash of the vault scoring method. On September 4, 2026, that method was version 6. Its hash covers the four weights, yield-quality rules, asset-quality drag, history event set, matured-asset rule, and stablecoin-backing policy. The hash also records whether the backing ceiling was enforced.
The run also records input fingerprints. One covers the vault set and asset registry summary. It includes the exact backing snapshots and ceilings used by the run. Another covers the yield-history rows and their as-of time. These fingerprints tie a score to a run. They are not a copy of every source row.
A full replay still needs the original database rows. It also needs the price feed state, run clock, code version, and methodology version. The method gives the recipe. The recorded run identifies the ingredients.
- GET /v1/methodology/vaultproduction
- GET /v1/methodologyproduction
- Vault composite:
0.3 × Asset + 0.3 × Platform + 0.2 × Control + 0.2 × History
Limits of the reading
The score does not predict the next exploit. It does not promise liquidity. It does not estimate a probability of loss. It does not replace due diligence.
Its purpose is narrower. It turns a defined set of evidence into one repeatable reading. The four vectors show where the score came from. The caps stop strong averages from hiding serious weaknesses. The run record lets a reader check which method produced the result.
That is the standard: score risk, show the evidence, and keep uncertainty visible.
References
-
S&P Global Ratings (2025). Understanding Credit Ratings. Accessed 2026-07-15. ↩
-
U.S. Securities and Exchange Commission. 17 CFR 240.17g-1. Electronic Code of Federal Regulations. Accessed 2026-07-15. ↩
-
DeFiSafety (2026). 0.9 Process Documentation. Accessed 2026-07-15. ↩
-
L2BEAT (2026). Risk Analysis. Accessed 2026-07-15. ↩
-
Haimovich, Y. & Goldberg, O. (2023). Aave V3 Risk Parameter Methodology. Chaos Labs. Accessed 2026-07-15. ↩