Methodology

How AI-matched sourcing works: a buyer's guide to Claude-reasoned supplier discovery

How the marketplace's discovery engine reasons over verified lots to rank them against a buyer's plain-language brief, with a rationale per match.

Methodology · 12 min read · 2026-08-01

A sourcing manager with a specific brief — a washed lot from a particular altitude band, a certification, a volume that fits a container — usually starts the same way: a spreadsheet of supplier contacts, a round of cold outreach, and a wait. Replies trickle in over days, unevenly, and half of them don't actually match what was asked for. The brief was clear. The search process wasn't built to read it.

There's a second cost buried in that wait that's easy to underweight: it's a poor use of a sourcing manager's actual expertise. The judgment calls that matter — whether a lot's quality profile suits a specific roast, whether a producer's volume covers a season, whether the compliance evidence is solid enough to build a filing on — are exactly the calls a buyer should be making. Reading through a long list of supplier replies just to find which ones are even in the running isn't that kind of work. It's clerical filtering wearing the shape of sourcing decisions.

That gap — a precise buyer intent versus an undifferentiated list of suppliers — is the problem AI supplier matching is built to close. Not by replacing judgment about who to buy from, but by doing the first pass of reading a long list of lots against a plain-language brief, faster and more consistently than a manual keyword filter can.

The problem with keyword search on a supply list

A conventional filter interface asks a buyer to translate their brief into fields: product, region, certification, minimum cupping score. That works when the brief is that simple. It works less well when the brief has trade-offs baked into it — "organic preferred but not required if the cupping score clears 85", or "volume matters more than origin this quarter." A dropdown can't hold that. A buyer either narrows the filters until almost nothing matches, or widens them until the results are a wall of undifferentiated listings that all technically qualify and have to be read one by one anyway.

The other cost is time. Reading through fifty listings' attributes to judge fit against a multi-part brief is exactly the kind of repetitive comparison work that doesn't need a person doing it fifty times in a row — it needs doing once, well, against every candidate, with the reasoning shown.

There's also a consistency problem that a manual pass over a long list runs into regardless of how careful the reader is. Judging fifty listings against a multi-part brief means holding several conditions in mind at once, across every row, without drifting. A person reading quickly toward the end of a long list applies the brief a little differently than they did at the start — not from carelessness, just from the ordinary fatigue of repetitive comparison. A single pass that applies the same stated criteria to every candidate, in the same way, every time, removes that particular source of drift. It doesn't remove the need for the buyer's own second read of whatever comes out the other end.

What the discovery engine actually does

The marketplace's discovery engine takes a buyer's natural-language query — written the way a person would actually describe what they want, not structured as a filter — and reasons over the pool of active, verified listings to rank them against it.

Mechanically, it works like this:

  1. 1The engine loads active, non-deleted listings from across the marketplace — not scoped to one producer or region, so a query can surface a fit the buyer wouldn't have thought to search for by name.
  2. 2Each candidate listing is reduced to the attributes a match actually depends on: product, origin region, quantity available, certifications, cupping score, and EUDR compliance signal. Producer and company identity are not part of what the model sees — the ranking runs on lot evidence, not on who is selling it.
  3. 3A single call to Claude (Anthropic's Claude Sonnet model) reads the buyer's query alongside that candidate set and returns a ranked list: which listings fit, a 0–100 fit score, and one sentence explaining why each one was ranked where it was.
  4. 4The response is schema-validated before it reaches a buyer. Anything the model returns that doesn't match the expected shape — or that references a listing ID that doesn't exist in the candidate set sent to it — is discarded rather than surfaced.
  5. 5Results are mapped back to the real, anonymized listing records and returned in ranked order, each carrying its fit score and its one-line rationale.

A few boundaries are worth being specific about, because they shape what a buyer should expect from a result set. The candidate pool sent to the model is capped, so a query is reasoned over a bounded, recent slice of active listings rather than the entire historical catalog. Listings below a low fit threshold are dropped rather than force-ranked, so a genuinely poor match doesn't show up padded into position forty. And the model is instructed to use only the listing IDs it was actually given — a safeguard against the familiar failure mode of a language model inventing a plausible-looking reference to something that isn't there.

What a fit explanation is, and what it isn't

Each result in a ranked list carries a short, plain-English reason: why this particular lot was placed where it was against this particular query. That explanation is generated by the model reading the same attributes a buyer could read themselves — region, certification, quantity, quality signal, compliance status — and stating which of them lined up with the brief.

It is a reasoned match, not a guarantee. The fit score is the model's estimate of relevance given the evidence it was shown, not a certified compatibility rating, and not a substitute for a buyer's own review of the lot's full evidence before an RFQ. A high fit score means the listing's recorded attributes line up well with what was asked for. It does not mean the lot has been independently vetted beyond what the underlying verification record already establishes, and it does not mean the buyer's own diligence — on quality, on compliance, on terms — is no longer needed.

It also isn't a running score that updates itself. A fit ranking is computed for the query a buyer actually asked, against the pool of listings active at that moment. Change the wording of the query, or come back to search again after new lots have gone up, and the ranking is computed fresh — not adjusted from the last result set. That's a reasonable thing to expect from a system built around a single reasoning pass over current evidence rather than a persistent buyer profile it's tracking over time.

A fit score and its explanation are a reasoned match against the evidence on file — not an audit, and not a guarantee that a lot is right for a buyer's order.

Reading a ranked result list well

The ranking is most useful treated as a fast first pass over a wide pool, not as a final decision. A few habits make that distinction concrete:

  • Read the fit explanation as a pointer to which attributes mattered, then verify those attributes against the listing's own evidence — cupping score, certification, EUDR status, region — rather than taking the sentence as the finding itself.
  • Treat a lower fit score on an otherwise interesting listing as a prompt to check what didn't line up, not as an automatic disqualifier. A lot can score lower on a query because of one mismatched attribute — quantity, say — while still being worth a look if that attribute is flexible.
  • Use the ranked list to narrow a long pool to a short one worth reading in full, then do the reading. The engine's job is to save the time spent scanning listings that were never going to fit; it isn't built to replace the review a buyer does on the handful that do.
  • Remember that identity isn't part of the ranking or the result. A buyer sees evidence and a fit rationale before an RFQ; the producer behind a matched lot is revealed only once that RFQ is sent and accepted, the same sequencing that applies to every listing on the marketplace regardless of how it was found.

Why this is framed as a ranking, not a recommendation engine

"Recommendation" implies the system is telling a buyer what to buy. That's not what's happening here. The engine is answering a narrower question — given this pool of verified lots and this stated brief, which ones read as a fit, and why — and leaving the buying decision, and everything that depends on it, with the buyer. That's a deliberate framing, not a hedge: a language model reasoning over structured attributes is well suited to comparison and ranking against a stated brief, and poorly suited to standing in for a buyer's commercial judgment about price, relationship, or risk appetite.

How this fits with the rest of a lot's evidence

A fit score sits alongside, not on top of, the other signals attached to a listing. The EUDR tier on a matched lot still reflects the same batch-level compliance status and risk level it would if the buyer had found the lot by browsing directly. A cupping score is still a quality reading, not something the matching process adjusts. A trust-vouch summary, where one exists on a lot, is a separate AI-generated summary of the evidence on file — also produced using Claude — and is written independently of how a buyer arrived at that listing. Discovery changes how a buyer finds a lot. It doesn't change what the lot's own evidence says once they're looking at it.

That separation matters for a simple reason: a strong fit score describes alignment with a query, and a strong EUDR tier or cupping score describes something else entirely about the lot itself. Neither one substitutes for the other, and a buyer's own review still has to hold both in view before an RFQ goes out.

Where discovery sits in the buyer's workflow

It helps to be concrete about where this sits relative to the rest of a sourcing process, because it's easy to overstate what one search does. Discovery replaces the first, widest pass — the one where a buyer is trying to figure out which lots, out of everything active, are even worth a closer look. It doesn't replace the RFQ conversation, the sample request, the price negotiation, or the contracting step that follow once a shortlist exists. Those stay exactly where they were: between the buyer and the producer, once identity is exchanged.

In practice, a query stands in for the manual equivalent of scanning a supply list against a written brief — the part of sourcing that is comparison work, not judgment work. A buyer still decides which of the ranked results to pursue, still requests samples, still reads the full evidence on a lot before committing to anything. What changes is how much of the list has to be read by hand before that real work starts.

The practical starting point

For a sourcing manager working a brief with more than one condition attached to it, the difference this makes is mostly one of speed and coverage: one written query stands in for a round of manual filtering across a pool wider than a single set of supplier contacts, and comes back with a ranked shortlist and a stated reason for each entry, instead of a folder of replies to read one at a time over several days. What it doesn't do is replace the reading, the verification, or the decision. It narrows what has to be read.

The underlying discipline is the same one that governs every AI-generated surface on this marketplace: the model reads real evidence and states its reasoning, a buyer checks that reasoning against the evidence itself, and nothing is presented as more settled than the record behind it. A fit score is exactly as good as the attributes it was computed from — which is precisely why those attributes stay visible on every matched listing, not just the sentence explaining the match.

Buyers ready to try a brief against the live pool of verified lots can browse the marketplace and search in plain language now; buyers who want a standing account before sending an RFQ can request buyer access.

Put this into practice

Browse EUDR-pilot lots matched to your buying profile.