Pricing off a published benchmark does not put a number in your contract. It incorporates somebody else's rulebook - a definition of which barrels count, where and when, owned by an administrator entitled to revise it. The number is an output of that rulebook, and so is your exposure.
Summary
Nearly every physical crude contract prices off a published benchmark. What the contract incorporates is not a number. It is somebody else's rulebook — a written definition of which barrels count, delivered where, assessed over which hours of which days, maintained by an administrator who is entitled to revise that definition and who has no contract with you. The number is an output of the rulebook. The exposure is to the rulebook.
This article sets out what a crude benchmark is as a contractual object rather than as a market quote: who owns the definition and what they are allowed to do with it, why one benchmark name usually covers several different instruments that do not print the same number, why a benchmark's definition has to keep changing in order to stay honest, and what a price reference has to say before two parties can be sure they mean the same thing by it. It is written for the people who inherit the consequences — traders choosing which marker to price and hedge against, the risk and middle office who value the result, and the settlement desks who discover at invoice stage that both sides read the same three words differently.
If you want the mechanism, start with The benchmark you reference is a definition, not a quote. If benchmarks are already your daily work and you want the part that goes wrong silently, start with A methodology change reprices your book without a trade.
The benchmark you reference is a definition, not a quote
A published benchmark is the output of a written methodology owned by its administrator. Pricing off the benchmark means contracting on that methodology, including the parts of it that have not been written yet.
Two kinds of number get used as crude benchmarks, and both are defined before they are observed. An exchange settlement price is defined by a contract specification: which grades are deliverable, at which point, for which delivery period, and by what procedure the settlement is determined at the end of the day. A price reporting agency's assessment is defined by a published methodology: which transactions are eligible, of what minimum size, on what delivery basis, transacted or bid within which window, and how that evidence is turned into a single printed figure.
In neither case does anyone trade the benchmark itself. People trade cargoes and contracts that fall inside a definition, and an administrator applies the rules to what it observed. The number arrives at the end of that process, and it means exactly what the rules say it means — no more.
The consequence is worth stating plainly, because it is not how the reference is usually read. When a contract says "priced off" a marker, it has incorporated by reference a document that neither party wrote, most likely has not read in full, and cannot amend. That document is maintained by a third party who is entitled to change it, who will consult the market before doing so, and whose obligation runs to the integrity of the benchmark rather than to your open position.
This is not a scandal. It is the design, and the alternative — a benchmark whose owner may not adjust it — is worse, for reasons the next two sections set out. But it does mean that a benchmark reference is a live obligation to track rather than a static input, and very few trade capture systems treat it as one.
One benchmark name, several different instruments
A benchmark's name is not by itself a sufficient price reference. The same name customarily covers a dated physical assessment, a forward for a delivery month, an exchange-traded futures contract and a cleared swap — and those four do not settle at the same number.
Each of the four is a legitimate use of the name, and each exists because someone needed a different thing. The dated assessment answers what cargoes loading in the near term are worth right now. The forward answers what a cargo in a named month is worth. The futures contract answers the same question in a standardised, cleared and margined form, which is what makes it hedgeable. The swap converts between them, which is only necessary because they differ.
They differ for structural reasons that do not go away: they are defined over different periods, they clear or do not clear, and the physical thing each is defined against is not identical. A contract that names the benchmark without naming which of the four it means has an ambiguity in its price term, and that ambiguity is invisible for as long as nobody has a reason to argue about it. The moment it becomes visible is the moment when one reading is worth more than the other, which is also the moment when neither party is inclined to concede.
The practical form of this problem is duller than a dispute and far more common. A hedge is placed against one of the four, a cargo is priced against another, and the residual between them is reported as an unexplained variance in the profit and loss — month after month, by a middle office that has correctly matched the volumes and correctly matched the dates, and is comparing two numbers that were never the same number.
A benchmark is anchored to a physical basket, and the basket keeps moving
A benchmark stays meaningful only while enough physical barrels actually trade under its definition. When that volume falls, the definition has to change or the benchmark stops measuring the thing it is named after.
Fields deplete. A marker defined on the output of a particular stream thins out as that stream declines, and a thin marker is both easier for a single participant to influence and less representative of the market that uses it. The administrator's options are limited and well known: widen the basket by admitting further grades, move or add a delivery point, adjust the assessment window or the quality adjustments that make dissimilar barrels comparable — or accept that the marker is drifting away from the market it purports to describe. The history of crude benchmarks is largely a history of the first two.
The conclusion that follows is the opposite of the intuitive one. Methodology change is not an anomaly in a benchmark; it is the maintenance that keeps it honest. A marker that has not changed in a market whose production has changed is not stable. It is losing its grip quietly, and losing it in a way that produces no error message: it still prints every day, at a plausible level, and the drift shows up as your differential doing something inexplicable.
What this asks of a trading business is modest and almost never done: know which markers your book depends on, know who administers each of them, and know that the definition behind each is a document with a version.
A methodology change reprices your book without a trade
A change to a benchmark's definition changes the value of open contracts that reference it, and it arrives with no trade, no counterparty and no entry into the trading system.
This is the load-bearing claim of the article, and the reason is structural. A position system learns that prices have changed by receiving prices. A definitional change is not a price — it is a change in what the price means — and there is no field on a trade for it. Everything downstream of the definition moves anyway:
A hedge relationship rests on an assumption about co-movement, and the definition is half of that assumption. If the basket underlying a marker is widened to admit barrels from another basin, the marker now responds to a wider set of things. It may become a better hedge for one book and a worse one for another. Nothing about either book changed.
A differential quoted against the old definition is not comparable to one quoted against the new. Historical differentials, negotiating precedents and the internal benchmarks people carry in their heads all reset, and they reset at different speeds for different people.
A model or a curve calibrated on the published history is now calibrated on a series that changed meaning partway along it. Whether the administrator restates the history, and how far back, is the administrator's decision and not usually a decision the model knows about.
And the notice arrives somewhere other than where the exposure is held. Administrators consult, publish proposals and give lead time — often generous lead time. That warning lands in a market data mailbox, or in a subscriber notice, or in a consultation paper. It very rarely lands in front of the person holding the position that the change will reprice.
The useful test takes an afternoon. For each marker your open contracts reference, ask three questions: who in this company would receive the change notice; which contracts would need to be repapered if the definition moved; and would anyone find out from us, or from the counterparty's proposal to amend.
Choosing a benchmark is choosing whose local problems you inherit
A benchmark carries the logistics of the place it is defined on. When those logistics break, the price you are priced against moves for reasons that have nothing to do with your cargo.
A marker defined on delivery at an inland pipeline hub prices, among other things, the cost and the possibility of getting barrels out of that hub. A marker defined on waterborne cargoes prices freight, loading windows and the ability of a vessel to be where it said it would be. These are not defects. They are what makes each marker a faithful description of its own market.
But they are inherited. Two markers for the same commodity can diverge sharply on a week when nothing whatever has happened to global oil: a pipeline outage, storage filling toward its limit, a maintenance season, a river or a canal doing what rivers and canals do. If a waterborne cargo is hedged against an inland marker, that inland constraint has become a position — one nobody chose, that will not appear in the position report as a constraint, and that will be discussed afterwards as if it were a market move.
So the benchmark choice is usually posed backwards. The question is not which marker gives the best price or the most liquidity. It is which marker's local failures you are willing to own, given where your barrels actually load and discharge — and, where the answer is "not these", whether the basis between the marker you must price against and the marker you can hedge in is being carried as an explicit position or as a residual.
The clause that matters is the one about the day the number is missing
A price reference has to say what happens when there is no number. That clause gets written while both parties believe it is theoretical, and read for the first time on the day it decides the money.
Publication fails in ordinary ways: a holiday the two parties classified differently, a technical failure, a day the administrator declines to assess because it judged the evidence insufficient. It fails in rarer and more serious ways too — a material change to the methodology that the contract does not describe, or the discontinuation of the quotation altogether.
The contract's fallback governs, and the available fallbacks are not equivalent to each other. Rolling to the last published day is neutral but arbitrary. Naming a substitute publication is precise but assumes the substitute still exists. "As agreed between the parties" is a negotiation held at the worst possible moment. A determination by one party is a decision handed to whoever drafted that line — and it is worth noticing which side that usually is.
Each of those allocates a decision to somebody, and they do not allocate it to the same somebody. The asymmetry is the whole point: a fallback clause is cheap to negotiate when nobody expects to use it, and expensive to be on the wrong side of once it is in force. It deserves the attention that the differential gets, and it does not receive a tenth of it.
Where this leaves the middle and the back office
None of the above is unusual. It is ordinary benchmark pricing, and it has been ordinary for a long time. What makes it expensive is where the governing facts are kept and how long they survive.
The trade capture system holds a price reference as a short string. The methodology behind that string lives on the administrator's website. The change notice lives in a mailbox. The published series lives in a market data feed, and often in a spreadsheet beside it. The licence that determines whether the firm is even entitled to store and pass on those values lives with whoever signed it. When an invoice is queried nine months later, the question is not what the marker is worth now, but what the administrator published for a specific past date, under the definition then in force.
That gives the practical version of everything above, and it is the sentence worth taking away: a price you can see is not the same as a price you can reproduce. A screen shows the current value. Settling and defending an invoice requires the value as published for a date that has passed, together with enough of its context to show which quotation it was and which definition produced it. Displaying is not keeping, and the difference between them is discovered at the point where it is too late to start keeping.
Four things a benchmark name can mean
| Dated physical assessment | Forward for a delivery month | Exchange futures contract | Cleared swap | |
|---|---|---|---|---|
| What it answers | What cargoes loading in the near term are worth now | What a cargo in a named month is worth | The same question, standardised and cleared | The difference between two of the others |
| Who defines it | The price reporting agency's published methodology | Market convention and the parties' terms | The exchange's contract specification | The clearing house rules and the swap terms |
| What it is mainly used for | Pricing physical cargoes | Trading a month before it becomes current | Hedging and curve construction | Converting one basis into another |
| Where its value comes from | Transactions and bids observed in an assessment window | Bilateral and brokered dealing in that month | Exchange trading, margined daily | Settlement against an average of one of the others |
| How a contract goes wrong on it | Names the marker but not which assessment | Assumes it equals the futures month | Assumes it equals the physical assessment | Assumes both legs share a definition |
The bottom row is the one to sit with. Each of these four failures is a different sentence in a contract, and each of them produces the same symptom: a residual that middle office cannot explain and cannot make go away by hedging more.
What a price reference has to state before both parties can be sure they agree
| Field | Why it is load-bearing |
|---|---|
| The administrator, and the exact published name of the quotation | The name of the commodity is not the name of a quotation, and administrators publish several per commodity |
| Which of the four instruments above is meant | This is the ambiguity that stays invisible until one reading is worth more than the other |
| The unit, the currency and the delivery basis of the quotation | Two of these can match while the third quietly does not |
| The rule that turns published values into the contract's number, and who calculates it | Both sides computing independently from the same series is how identical inputs produce different invoices |
| Rounding and decimal convention | The smallest field on this list, and the one that produces the most disputes per unit of importance |
| What happens on a day the quotation is not published | Without it, a routine holiday becomes a negotiation |
| What happens if the methodology changes materially | Silence here means the change applies to your open contracts by default |
| What happens if the quotation is discontinued | The only item on this list that cannot be fixed after the fact |
| Whether the firm is licensed to store and redistribute the values it prices on | Storing a benchmark series is a licensed activity, and the invoice depends on being able to store it |
A price term carrying these nine items can be valued and settled by someone who was not in the room. One missing several of them can be handled by whoever remembers the negotiation — which works, until the administrator publishes a consultation and that person is on holiday.
Questions people ask about this
Who decides what a crude benchmark actually measures?
Its administrator — the exchange that lists the contract, or the price reporting agency that publishes the assessment. The definition sits in a public document: a contract specification for a futures contract, a published methodology for an assessment. Market participants are consulted and their views carry real weight, but the decision is the administrator's, and it is made in the interest of keeping the benchmark representative rather than in the interest of any particular open position.
Our contract just says the benchmark's name. Is that a problem?
It is an unresolved ambiguity rather than an immediate one. A benchmark name typically covers a dated physical assessment, a forward month, a futures contract and swaps that settle against them, and those instruments do not print the same number. In most months the reading both sides assume is the same one and nothing happens. The cost surfaces when the two readings diverge enough to matter, which is precisely when neither side wants to concede the point. Naming the administrator, the exact quotation and the instrument costs a line of drafting and removes the question permanently.
What happens to our open trades if the benchmark's methodology changes?
Unless the contract says otherwise, they continue to price against the benchmark as newly defined, because what the contract referenced was the benchmark rather than a snapshot of its rules. The commercial effects arrive without any trade taking place: the hedge relationship you assumed may be better or worse than it was, differentials negotiated under the old definition stop being comparable, and models calibrated on the published history are calibrated on a series whose meaning changed partway along. Administrators generally give substantial notice — the difficulty is that the notice arrives in a market data mailbox rather than in front of the position.
Why do two benchmarks for the same commodity move apart?
Because each one is defined at a place, and places have their own constraints. A marker defined at an inland hub carries the cost of moving barrels out of that hub; a waterborne marker carries freight and loading. When one of those constraints binds — an outage, storage filling, a maintenance season — that marker moves and the other does not, without anything having happened to the commodity globally. If you are priced against one and hedged in the other, that local constraint is now a position on your book.
How should we choose which benchmark to price against?
Start from where your barrels physically load and discharge, and ask which marker's local failures you are prepared to own, rather than which marker looks like the best price. Then check the second question separately: whether the marker you must price against is one you can actually hedge in. Where it is not, the gap between the two is a real exposure and the useful thing is to carry it explicitly, before deciding whether it can be covered at all.
What if the benchmark is not published on a day we need it?
Whatever your contract's fallback provision says — and the fallbacks differ from each other in who they hand the decision to. Rolling to the last published value is neutral. Naming a substitute publication is precise but depends on the substitute existing. Leaving it to agreement between the parties schedules a negotiation for a bad day. A determination by one party gives the pen to whoever drafted the clause. This is worth reading before a cargo is fixed, because it is far easier to agree a rule than to agree an outcome.
Do we really need to keep our own copy of published benchmark prices?
If you settle invoices against them, yes, and the reason is not analysis but defence. Pricing a past cargo requires the value the administrator published for that specific past date, under the definition then in force — not the value on a screen today, and not a restated series. That means capturing the series as published, with enough context to identify which quotation it was, and keeping it for as long as invoices can be queried. It also means checking what your market data licence permits, because storing and passing on those values is a licensed activity rather than a technical one.
What should we check first across the contracts we already have?
Two passes, both narrow enough to finish. First, list the distinct price references your open contracts actually use and see how many are ambiguous about which quotation is meant — this is usually a shorter list and a worse result than people expect. Second, for each administrator on that list, find out who in the firm receives its notices. Those two answers tell you where a definitional change would land and who would see it coming, which is the whole of the problem in this article reduced to something a person can go and do.
Where this lands in a trading system
Nothing above is an argument for buying software. It is an argument about which facts a price reference depends on, where they live, and how long they have to survive. The conclusions do map onto system capabilities in places, and it is worth being specific about which places, and about where the mapping stops.
The first conclusion is that a hedge and the thing it hedges have to be visible together before anyone can tell whether they share a definition. That is the job of a CTRM/ETRM system providing complete trading lifecycle management for commodity traders, integrating physical trade, financial hedging, risk control, and settlement in one platform; Time Dynamics' Fusion is one, and the part of it that bears on this article is that complete physical trade lifecycle management and execution and advanced trading management for futures and derivatives markets sit in the same place, which is the condition under which a residual between two benchmark legs can be attributed rather than merely observed.
The second conclusion is that benchmark values arrive from outside the firm and have to be kept, not merely displayed. That is a different problem, and it is what X-Ray addresses: a non-invasive data processing and analysis platform designed specifically for enterprise clients, whose collection toolkit XDK does non-invasive automated data collection from databases, Excel files, and web interfaces, held in a time series store offering unified storage with unlimited scalability for structured and unstructured data. The relevant property of that pairing is reproducibility rather than volume — a published series kept as published is what lets an invoice be defended on a date long past. The second relevant property is that it will collect data without disrupting existing systems or workflows, which matters here because the spreadsheets holding your marker history are working, and a plan that begins by replacing them will not survive its first pricing week.
The third conclusion is that exposure has to be re-cut by marker and by quotation faster than a person can rebuild the view, particularly in the weeks around a consultation. X-Sheet does one-click generation of personalized reports with Excel-like interface.
No system on this page reads a methodology document, and none of them will tell you a definition is about to change — that is a subscription and a named person, not a feature, and pretending otherwise would be the most useful-sounding false claim in this article. Nor will a system choose your benchmark, judge whether a fallback clause is fair, negotiate the licence that lets you store a series, or make a hedge fit a marker that was the wrong one to pick. What it can settle is narrower and duller: whether the number that priced an invoice is the number that was published for that date, whether the instrument behind it is recorded rather than assumed, and whether the person who needs either can see them without asking someone who was in the room. That is a real part of the gap. It is not the whole of it, and the difference is worth knowing before you buy anything.
A note on the evidence in this article
There are no statistics here, and that is a choice rather than an oversight. The observations behind the article — which parts of a price term are commonly left ambiguous, where change notices land, what a settlement desk is missing when an old invoice is queried — come from material we are not in a position to publish as a citable figure. A number nobody can check would make the argument read as stronger while making it worth less.
So the argument runs on mechanism instead. Each claim is made from how the thing is built, and each is stated so that you can test it against your own contracts rather than accept it on our authority. The three tests proposed above are the intended way to use it: which of our price references name a quotation rather than a commodity, who here receives an administrator's notice, and can we reproduce the value that priced last year's invoice. If your own book answers those differently from what is written here, your book is the better evidence.