How Philidor Scores a Vault

This is a methods reference. It explains how Philidor computes a vault score. The score describes the evidence available at scoring time. It is not a forecast or an investment recommendation.

A Philidor score answers a narrow question. How strong was the evidence when we scored this vault?

The answer runs from 0 to 10. Higher is better. But a score of 8 does not mean an 80 percent chance of safety. The scale compares evidence across vaults. It does not state a probability.

Two rules shape the score.

First, one serious weakness can limit the whole result. Strong evidence in three areas cannot cancel unaudited code or weak collateral. Second, missing evidence never counts as good news. We use a cautious default, a neutral value, or a failed run. The choice depends on what is missing.

Figure 1Four readings combine into the vault score
The Philidor vault score combines Asset and Platform at 30 percent each with Control and History at 20 percent each. Asset-level caps apply first. Vault-level caps can then lower the weighted result.
Source: The production composite and cap stack, methodology v6; dataset committed at /data/how-philidor-scores-a-vault/board_scores_vault.json

Why publish the method

A rating is more useful when readers can see what the number means. In its 2025 guide, S&P says its rating methods are transparent and free to read.1 Registered rating firms in the United States must also publish key parts of their procedures.2

By 2026, DeFi risk teams had taken several approaches to the same problem. DeFiSafety publishes an item-by-item process review.3 L2BEAT shows separate risk criteria instead of hiding them inside one number.4 Both approaches make the judgment visible. A reader can see which fact drove the result.

Economic-risk teams answer a different question. Chaos Labs publishes simulation methods for setting market parameters.5 That work connects market behavior to choices such as collateral limits. It is less concerned with producing one common reading across many vaults.

Prior work established how to publish review criteria, show distinct risk flags, and disclose economic models. This paper adds a single vault score with a fixed calculation. It combines current exposure, platform, control, and incident evidence. The public method lets a reader trace the result back to those inputs and the time of the scoring run.

The score in one line

The score starts with four components. We call them vectors. Each uses the same 0-to-10 scale.

VectorWeightWhat it asks
Asset0.3What does the vault hold, and how strong is each asset?
Platform0.3How strong is the protocol, its code record, and its dependencies?
Control0.2Who can change the system, and how quickly can they act?
History0.2What serious events have affected this vault?

The raw score is:

0.3 × Asset + 0.3 × Platform + 0.2 × Control + 0.2 × History

We round that result to two decimal places. We then apply the cap stack. A cap is a ceiling. It can pull the score down, but it can never push it up.

Asset: what the vault holds

The asset vector reads every holding reported by the vault adapter. An adapter is the connector that reads a protocol's positions. The scorer removes duplicate holdings by chain and address. It then scores each exposure through the asset registry.

Different assets need different tests. A native token is scored on liquidity and volatility. Each has a 50 percent weight. A bank-backed stablecoin needs a different test. Custody carries 25 percent. Reserve transparency carries 20 percent. Redeemability, peg stability, and liquidity each carry 15 percent. Governance carries the remaining 10 percent.

Private-credit tokens use another mix. Credit quality carries 20 percent. The default rate and custody each carry 15 percent. The rest covers reserve reporting, recovery, manager concentration, governance, liquidity, and redemption terms.

The point is simple. We do not grade every asset with one generic checklist. The full category matrix is available through /v1/methodology.

Unknown and unreviewed assets

An asset with no registry entry scores 2.0. The same score applies when the adapter cannot resolve its address. We do not guess which asset it meant.

Review status also sets a ceiling:

Review statusHighest possible asset score
Reviewed10.0
Provisional9.0
Unreviewed7.9

A dated instrument gets special treatment after maturity. Pre-maturity market evidence is no longer current. The scorer drops those stored dimensions and observations. It also removes the lockup and yield-bearing adjustments. The asset is then capped at 7.9. This is not a claim that the asset failed. It means the old assessment has expired.

Hard-fail flags

Some facts are serious enough to set their own cap. The scorer calls them hard-fail flags.

FlagAsset cap
Active depeg1.0
Confirmed sanctions exposure0.0
Redemptions paused2.0
Single-signer upgrade authority3.0
High use of the protocol's own token as collateral4.0
Missing proof of reserves4.0
No recent reserve attestation5.0
Unaudited token contract4.0

If several flags apply, the lowest cap wins.

Stablecoin backing

For a monitored stablecoin, Philidor can look through the coin to its backing. The scorer grades each backing sleeve. It then weights the sleeve scores by their share of the reserves. An unknown or unresolved sleeve scores 2.0.

That weighted backing score sets a ceiling for the stablecoin's asset score. The ceiling rises smoothly with backing quality:

Backing scoreAsset-score ceiling
8.0 or higher10; no backing cap
7.85 to below 8.0Rises from 7.9 to 10
5.0 to below 7.85Rises from 4.9 to 7.9
Below 5.0Rises from 0 to 4.9

Only the latest complete snapshot may bind. A partial or disputed snapshot cannot set the cap. The last complete snapshot remains in force. The snapshot must also cover most of the backing. If 50 percent or more is unknown or fail-closed, Philidor treats that as a monitoring gap and applies no backing cap.

A complete snapshot older than 24 hours gets an uncertainty drag. Its ceiling falls by up to 0.5 points. The rule never raises a lower ceiling. When the ceiling is already below 4.9, it stays below 4.9.

Old data and risky construction

Old evidence loses weight. A stale dimension is multiplied by 0.92. An expired dimension is multiplied by 0.75. If more than half of the weighted dimensions are stale, the asset score cannot exceed 8.0.

Loan-to-value settings also matter. Loan to value, or LTV, is the share of collateral value that a borrower may borrow. High LTV leaves less room for a price fall. At an LTV of 0.85, the asset score loses 2.5 percent. At 0.90, it loses 5 percent. At 0.95, it loses 10 percent.

The scorer then looks at the portfolio. For a multi-asset vault, it uses the Herfindahl-Hirschman Index, or HHI, to measure concentration. An HHI above 0.5 subtracts 0.5 points. An HHI above 0.8 subtracts 1.0 point. Wrong-way risk subtracts another 1.0. Wrong-way risk means the exposure and the system behind it may fail together. Category correlation is shown to readers but does not change the score.

Platform: what the vault runs on

The platform vector starts with three readings. They cover age, audits, and strategy type. The scorer averages them.

Protocol age uses a curve called the Lindy score:

10 × (1 - e^(-days / 365))

The score rises quickly in the first years. It then flattens. New code has had less time to face real use. Older code has survived more conditions. Age alone does not prove safety, so it is only one part of the platform vector.

The audit score is 0 when no audit is recorded. With at least one audit, it starts at 4. A public audit contest adds 2.0. An audit by an established firm adds 1.5. A recognized firm adds 1.0. An unknown or unclassified firm adds 0.5. The audit score cannot exceed 10.

This rule values audit quality, not just audit count. It does not grade the declared scope or findings. Those details are shown separately. Missing scope metadata does not change the audit score.

Protocol incidents

A recent major incident can cap the platform score. An unresolved incident sets these ceilings:

Time since incidentPlatform cap
Less than 30 days2
30 to 89 days5
90 to 179 days8

A resolved major incident is treated less harshly. The cap is 5 for the first 30 days. It is 8 from day 30 through day 89. Minor incidents do not change this vector.

Dependencies

A vault inherits risk from the services below it. The platform score therefore uses the weakest dependency.

A dependency score of at least 7.5 applies a 0.95 safety factor. A score from 5.0 to 7.49 applies 0.8. A lower score applies 0.5. Each extra dependency after the first adds a 3 percent discount. That discount stops at 0.85.

The internal exchange rate of a liquid-staking token is not treated as an external service. The scorer excludes the dependency named exchange_rate.

Yield quality for money markets

Money-market vaults get one more platform check. It applies to Aave, Spark, Compound, and Morpho. It does not apply to liquidity pools or aggregators. Their reward data has a different meaning.

The check reads up to 14 days of net yield, base yield, and utilization. It needs at least five usable snapshots across three days. It uses an hourly, time-weighted median. Long-lived conditions therefore count more than brief spikes.

Two flags can lower the platform score:

  • Thin exit liquidity. Median utilization above 95 percent caps the platform score at 8.0, then subtracts 1.0. Once active, the flag clears only when utilization falls to 90 percent or below.
  • Reward-dependent yield. A real-yield share below 50 percent caps the platform score at 9.0, then subtracts 0.5. The rule also requires at least $1 million of value and a net annual yield of at least 1 percent.

Low utilization is reported as idle capital. It does not lower the score.

Control: who can change the system

The control vector asks two questions. Who has authority? How much warning do depositors get before a change?

For EVM contracts, immutability scores 10. A liquidity pool with no pause power scores 9. In that case, governance can change fees but cannot move user funds.

Most other EVM contracts are scored by their timelock. A timelock is the delay between proposing a change and executing it. It gives users time to react.

TimelockControl score
At least 168 hours9
At least 72 hours8.5
At least 48 hours8
At least 24 hours6
At least 6 hours4
More than 0 but less than 6 hours2
No delay1

Missing permission data scores 5. An unknown timelock also scores 5. Five means unknown here. It does not mean safe.

Solana uses a separate control test. An upgrade authority held by one ordinary key scores 1. Revoked upgrade authority scores 7 when market parameters can still change. It scores 8 when the observed market settings are also immutable. We stop at 8 because emergency and global-admin paths may remain.

A multisig requires several approved signatures. An effective multisig can score 3.5, 5, or 6.5. The result depends on its signature threshold. The multisig must control every live program in the vault. Its own configuration authority must also be disabled.

A native multisig delay can lift that base score. The scorer compares the base with the matching EVM timelock score. It takes the higher value, adds 0.5, and caps the result at 9.

A single-key market admin can cap the result at 4. A single-key emergency signer can cap it at 6. Active emergency mode caps it at 3. Missing Solana program evidence scores 5.

History: what has happened

The history vector starts at 10. It then subtracts penalties for published, non-retracted events linked to the vault.

Only realized adverse event types count. They include incidents, bad debt, debt socialization, emergency pauses, shutdowns, and proxy changes. Governance and cap-management events do not count here. Their risk belongs in the control and concentration readings. Philidor's own RatingChange events are also excluded. A score must not react to its own output.

Each Critical event subtracts 1.5 points before time decay. The decay weight is 1.0 inside 30 days. It falls to 0.6 through day 89, 0.4 through day 179, and 0.2 through day 364. Events older than one year leave the decaying query.

A Critical event with confirmed loss also carries a permanent penalty of 1.0. These lifetime penalties stop growing at 4.0 points.

Warnings use the same time buckets. Their combined count passes through a logarithm. This makes the first warnings matter more than the hundredth repeat of the same pattern. The base multiplier is 2.0.

A vault earns a 0.5 clean-streak bonus when two conditions hold. It must have no Critical event in the last 90 days. It must also have no lifetime Critical loss event.

History can cap the whole vault score. A history score below 4 sets a total cap of 4.9. A history score below 7 sets a total cap of 7.5.

The cap stack

The raw four-part average now passes through the cap stack. The order is fixed. Every step uses the lower of the current score and the new ceiling.

1. Asset-quality drag

The scorer first builds an asset-quality anchor. This is the weighted asset score before the market's LTV penalty. It answers a basic question: how strong is the collateral itself?

An anchor below 5.0 caps the vault at 4.9. An anchor from 5.0 to below 7.85 caps it at 7.9. From 7.85 to below 8.0, the ceiling rises smoothly from 7.9 to 10. An anchor of at least 8.0 adds no cap.

The smooth final band prevents a tiny asset move from causing a large tier jump. A change from 7.99 to 8.00 should not create a two-point rating move.

One contamination rule is stricter. Suppose one exposure is at least 5 percent of the vault. If that exposure has a binding fail-safe cap, the whole vault is capped at 4.9. A serious unknown cannot be averaged away by safer assets.

2. Unaudited protocol version

A protocol version with no audit history cannot score above 4.9. Strong assets, controls, and history cannot lift unaudited code past that ceiling.

3. History

The history ceiling comes next. It is 7.5 or 4.9 when the history vector crosses the thresholds described above.

4. Overrides and active incidents

A documented vault override can add another ceiling. The stored override uses a 0-to-100 scale. The scorer divides it by ten before applying it.

An unresolved, protocol-wide incident can also carry a ceiling. That value comes from the incident record. The code does not invent one. The incident must match the vault's protocol. A version or chain on the incident narrows the match. If several incident ceilings apply, the lowest one wins.

Curator conduct is disclosed separately. It does not change the production vault score. The methodology endpoint stated this rule on September 4, 2026.

How missing data is handled

There is no single default for every missing field. Each gap has its own safe treatment.

Missing evidenceTreatment
Asset not found, or address unresolvedAsset score 2.0
Protocol metadata absentPlatform score 0
Audit history emptyAudit score 0, plus total cap 4.9
Permission data absentControl score 5
Timelock unknownControl score 5
No qualifying event rowsHistory starts from the clean default of 10
Stablecoin backing mostly unobservedNo backing ceiling binds
Curator identity absentNo score effect; coverage is shown as absent

The history default needs care. A 10 means the query found no qualifying events. It does not prove that no event happened outside Philidor's records.

The whole run can fail too. If there is no active asset methodology, the scorer stops. It does not write zero scores. If the number of scored vaults differs from the eligible count, the run is marked fail_closed. The run is marked degraded when the asset cache is more than one hour old. The same happens when more than half of scored vaults contain only provisional or unreviewed exposures.

Shutdown and frozen vaults stay in the scoring set while they remain active. They may still hold user funds. Their evidence must therefore keep changing with the record.

From a score to a public rating

The raw score maps to three public tiers:

ScoreTier
8.0 to 10.0Prime
5.0 to 7.99Core
Below 5.0Edge

The database can store a small score movement without publishing a new tier. This separation keeps the score honest while limiting visible noise.

An upgrade must hold for 6 hours before the public tier changes. Most downgrades must hold for 24 hours. A deeper fall moves at once. The immediate floor is below 7.75 when leaving Prime. It is below 4.75 when leaving Core.

Certain risk events also move at once. They include a change to an incident cap or hard-fail flag. A methodology or override change also moves at once. A reversal during the waiting period resets the clock.

Material changes enter rating_history. A tier change always qualifies. So do incident caps, hard-fail flags, and methodology changes. A within-tier score move qualifies at 0.5 points.

The history row and the database notification are written in the same transaction as the score update. They either succeed together or fail together. Public feed events are created only after that transaction commits.

A downgrade into Edge becomes Critical. Other downgrades become Warning. Upgrades become Info. A new incident cap or hard-fail flag becomes Critical when it lowers the score. A within-tier move reaches the public feed only at 1.0 point. At most 50 rating events are projected from one scoring run.

How to check a score

The scorer has no random branch. Fixed inputs and a fixed clock produce the same result. That is what deterministic means here.

Each scoring run records its start time and status. It records the code version and the asset-methodology version. It also records a hash of the vault scoring method. On September 4, 2026, that method was version 6. Its hash covers the four weights, yield-quality rules, asset-quality drag, history event set, matured-asset rule, and stablecoin-backing policy. The hash also records whether the backing ceiling was enforced.

The run also records input fingerprints. One covers the vault set and asset registry summary. It includes the exact backing snapshots and ceilings used by the run. Another covers the yield-history rows and their as-of time. These fingerprints tie a score to a run. They are not a copy of every source row.

A full replay still needs the original database rows. It also needs the price feed state, run clock, code version, and methodology version. The method gives the recipe. The recorded run identifies the ingredients.

ReproducibilityData as of
Formula
  • Vault composite: 0.3 × Asset + 0.3 × Platform + 0.2 × Control + 0.2 × History

Limits of the reading

The score does not predict the next exploit. It does not promise liquidity. It does not estimate a probability of loss. It does not replace due diligence.

Its purpose is narrower. It turns a defined set of evidence into one repeatable reading. The four vectors show where the score came from. The caps stop strong averages from hiding serious weaknesses. The run record lets a reader check which method produced the result.

That is the standard: score risk, show the evidence, and keep uncertainty visible.

References

  1. S&P Global Ratings (2025). Understanding Credit Ratings. Accessed 2026-07-15. ↩

  2. U.S. Securities and Exchange Commission. 17 CFR 240.17g-1. Electronic Code of Federal Regulations. Accessed 2026-07-15. ↩

  3. DeFiSafety (2026). 0.9 Process Documentation. Accessed 2026-07-15. ↩

  4. L2BEAT (2026). Risk Analysis. Accessed 2026-07-15. ↩

  5. Haimovich, Y. & Goldberg, O. (2023). Aave V3 Risk Parameter Methodology. Chaos Labs. Accessed 2026-07-15. ↩