Scaled customer success · Part 2

What did we know about this customer in March?

Part 1 argued that the hard part of enabling AI is not the AI. It is getting the context ready for it. This one is about the context that has a date on it, which turns out to be most of it.

A customer success manager inherits an account in August and reads that its health dropped in March. The obvious question is what changed. The question underneath it, and the one that decides whether the account is recoverable, is what anyone knew at the time.

That second question is where most retrieval systems quietly stop working.

Point a frontier model at a search engine and ask it about a customer. It will answer well, and it will answer about the world as it is now. When a page is rewritten, the previous version is gone, so the model has no source for what the page used to say. I use exactly that combination inside my own system as a last-resort fetch, so this is a boundary I have tested. It is very good at finding a page. It has no memory of what that page said last spring.

A graph can hold that memory. Most graphs do not, and the reason is usually that the design treated time as a property to store instead of as a thing to traverse.

Three clocks, routinely collapsed into one

Every fact worth keeping carries three timestamps, and they are all different.

When the fact was true. The quarter the revenue was earned. The day the champion started the job. The date the licences were activated.

When the source said it. The publication date of the press release, the filing date of the report. Frequently months after the fact was true, and occasionally before, in the case of anything announced in advance.

When you learned it. The moment your pipeline fetched the page and wrote the fact down.

Ask "what is this customer's revenue" and only the first clock matters. Ask "what did we know about this customer in March" and you need the third, because the honest answer includes facts that were already true in March but that nobody in your company had yet read.

Most systems store the first, sometimes the second, and almost never the third. The collapse is invisible until someone asks a question that needs two of them held apart, and by then the data to separate them was never written.

In my system the third clock lives on the provenance edge, not on the fact. Every node links back to the chunk it came from, and that link carries the extraction timestamp and the extracting model. The retrieval timestamp sits on the source document and is set only when the document is first created, so re-ingesting a page does not reset it. That placement was not deliberate at first. It turned out to matter, because it means the question "when did we learn this" and the question "how do we know this" are answered by the same traversal. Provenance and time are closer to being one problem than two.

The identity key is the temporal model

Here is the failure that taught me this, and it is worth the detail because the shape of it is general.

I ingest structured financial data from regulatory filings. That data labels every fact with the fiscal year and period of the filing it was disclosed in, not of the fact itself. An annual report carries the prior quarters as comparatives right next to the annual figures and stamps the whole set with the annual label.

Trusting that label gave me eight rows near $1.4 billion, every one of them marked as a full year's revenue for a company whose actual revenue that year was just under $6 billion.

The worse half is what happened next. The identifier for each stored figure was built from the company, the metric, and the end date. A company's fourth quarter and its full year end on the same day. So the annual figure and its own Q4 figure resolved to the same identifier, the second write landed on top of the first, and the annual revenue was silently overwritten by a quarter of itself. The graph did not hold a wrong annual revenue for that company. It held none at all.

The fix was simple. Derive the period from the fact's own start and end span rather than from the filing's label, and put that period into the identifier.

If your identity key does not carry the time dimension, you do not have versioning. You have overwriting with extra columns.

Make it consistent. You can store asOfDate on every node in the graph and still lose history on every run, because the write that destroys the old value never consulted the field you were so careful to populate.

The same rule catches a subtler case. Qualifiers belong in the key too. A figure reported on a standard basis and the same figure reported on an adjusted basis are two facts about one period, and if only the period is in the key they will overwrite each other exactly the way the quarter overwrote the year.

Three ways to put time in a graph

There are broadly three, they are not interchangeable, and the useful discovery for me was that a real system needs all three. The pattern should follow the shape of the fact.

01A versioned fact node per period

The fact gets its own node, and the period is part of its identity. The entity keeps one node and accumulates many facts hanging off it.

MATCH (c:Company {companyId: $id})-[:HAS_METRIC]->(f:FinancialSnapshot)
MATCH (f)-[prov:EXTRACTED_FROM]->(:TextChunk)
WHERE f.asOfDate <= $asOf          // true by then
  AND prov.extractedAt <= $asOf    // and we had read it by then
WITH f.metricType AS metric, f ORDER BY f.asOfDate DESC
RETURN metric, head(collect(f)) AS latest

Two filters, two different clocks. Drop the second one and you have answered what was true in March, which is an easier question and a different one. Use this pattern for anything measured repeatedly: revenue, headcount, usage, a health score. Re-ingestion accumulates a genuine time series with no separate versioning layer anywhere in the system, because every run either writes a new period or updates the one it belongs to. A trend question is then a sort over rows you already hold.

The cost is node count. This pattern is why a retention policy stops being optional, which I come back to at the end.

02An interval on the relationship

The fact is the relationship itself, and it has a start and an end. One edge, two dates, and an open end meaning still true.

MATCH (c:Company {companyId: $id})-[r:USES_VENDOR]->(v:Company)
WHERE r.periodStartDate <= $asOf
  AND (r.periodEndDate IS NULL OR r.periodEndDate > $asOf)
RETURN v.name, r.relationshipType

Use this when the fact has a duration and does not recur: a vendor relationship, an entitlement term, a contract-free notion of "was using this between these dates." The traversal reads as the English question does, which is a good sign. Note that the open end has to be a null and not a far-future date, because a far-future date is a claim you cannot support and it will eventually be wrong in a way nothing flags.

03A state flag alongside the interval

The entity carries a boolean saying which version is live. isCurrent on a role. isActive on a subscription.

MATCH (c:Company {companyId: $id})<-[:AT]-(r:Role)<-[:HOLDS_ROLE]-(p:Person)
WHERE r.isCurrent = true
RETURN p.name, r.title

This is the fastest of the three and by far the most common, because "who is the CEO now" gets asked a hundred times for every "who was the CEO in 2019." The flag is a cache of a temporal computation.

And a cache of a temporal computation is a temporal answer with the date already baked in. Which means the query above has no as-of form. There is no parameter you can pass it to ask about March. You have to fall back to the interval:

WHERE r.startDate <= $asOf
  AND (r.endDate IS NULL OR r.endDate > $asOf)

So the flag is fine as an index. It is not fine as the truth. Keep the interval underneath it, always, and treat the flag as something derived that you are allowed to rebuild.

What goes wrong when the flag is the truth

I learned that last paragraph twice, from two bugs pointing in opposite directions.

The first: a role extracted from a five-year-old press release is written with isCurrent true, because the press release said so at the time and nothing in the sentence knows it is old. Nothing ever sets it false. Ask who runs the company and you get everyone who has ever run it, all of them equally current, sorted by nothing meaningful.

The obvious fix is to close out prior holders. When a new person is written into an office, retroactively end any other still-open role at that company for the same office.

Which produced the second bug. A board has many directors holding the title "Director" at the same time. Every newly ingested director retroactively marked all the previously ingested ones as no longer current, so the board collapsed to whichever director happened to be processed last, and "last" meant last in chunk order rather than last in time. I found this against a real company's board and it had eaten the entire list.

The exclusivity assumption was doing the damage. Some offices hold one person, some hold many, and the mechanism has to know which before it is allowed to close anything. There is now a second guard as well: if two people with the same title turn up in the same day's ingestion, neither retires the other, because two same-titled people surfacing together out of one pass are far more likely to be colleagues the title rules did not anticipate than a same-day replacement. When the choice is between destroying data and keeping an ambiguity, keep the ambiguity.

The denominator has a date too

This one did not come from the graph work. It comes from every adoption metric I have ever had to defend in a review.

What percentage of your customers were on the latest release last quarter? Ask that question about today and there is no problem at all, because both halves of the fraction are sitting in front of you. The trouble starts the moment the question carries a date.

The numerator is obviously a point-in-time fact, and teams treat it as one: the number of accounts running that version on that day. The denominator is a point-in-time fact as well, and that is the half that slips past you, because the current customer list is always sitting right there ready to use. What the question actually needs is the number of accounts entitled to run the product on that day, and that is a different number from the count entitled today.

Join a historical numerator to a live denominator and you produce a figure that was never true on any date. Not on the date it claims to describe, and not on the date you ran it.

Left panel shows a ratio: the numerator is the count of accounts on the latest version as of March, and the denominator is the count of accounts entitled to the product as of today. Right panel shows two adoption curves over twelve months. The solid line, what actually happened, climbs from forty per cent to seventy per cent. The dashed line, what the report showed, falls from eighty per cent to seventy per cent. The two lines meet only at today.
The customer base shrank, so every earlier month was divided by today's smaller list. A real climb renders as a decline, and the two curves agree only at today.

The error has a direction, and the direction flips depending on which way your customer base moved. Every past month gets divided by today's population, so if the base grew, the past is divided by a larger world than it had and comes out too low. The improvement since then looks bigger than it was, and a program that moved very little collects credit for a climb. If the base shrank, the same query inflates the past instead, and a genuine improvement renders on the chart as a decline. That second one is the expensive case, because the number is telling you to shut down the thing that is working.

The distortion is also largest in the oldest periods and zero in the newest, since today's denominator is correct for today. So the two curves, the real one and the reported one, always agree at the right-hand edge. Whatever the chart is wrong about, it is wrong about it in the past, which is the half nobody re-checks.

The symptom is easy to recognise once you know it. Re-run last quarter's number this quarter and it has moved, with no change to the code and no correction to the data. Teams usually blame the pipeline for that. The pipeline is fine. The question was underspecified.

So the rule is that a ratio is two facts and both of them carry a date, which means the population has to be versioned exactly as carefully as the behaviour you are measuring against it. That is the part people skip, and the reason is that the entitled population does not feel like a fact. It feels like a table. It is the current customer list, it lives in a system somebody else owns, and it is always available in its current form, which is precisely what makes reaching for it so easy.

The universal set is a fact with a date on it. If you cannot reconstruct who was entitled on a given day, every trend line you have drawn compares periods measured against different worlds.

Tense is a retrieval concern

Storing time correctly does not get you an answer that is correct about time.

A rollout scheduled for 2027 is a plan. It sits in the graph with a date, correctly, and the date is in the future. Retrieve it without saying so and hand the rows to a language model, and the model will write "200 licences were activated on 1 January 2027." Past tense, about a date that has not happened. I saw it in two of four runs before the retrieval layer started marking future-dated rows explicitly.

Nothing was wrong with the data. The row was right, the date was right, the model read a date next to a number and did what reading a date next to a number usually implies. Tense is not a property of a fact. It is a relationship between a fact and the moment you are asking, so it can only be established at retrieval time, and it has to be stated explicitly.

That is why the temporal boundary belongs in the question rather than in the store. My retrieval layer takes it as an explicit parameter with two settings: as of today, which filters out anything not yet true, and including planned, which keeps future-dated rows and labels them. An unrecognised value falls back to the narrower of the two, because a boundary is a scoping choice and the safe direction to fail is the one that claims less.

This is also where Part 3 of this series connects back. The argument there is that a graph lets a system say "they do not have it" instead of "I could not find it," which is a far stronger claim than text retrieval can make. That claim is only sound if the world you closed is a world you can describe. Closed as of when is part of describing it. An entitlement answer that is complete by construction as of today is not complete as of March, and a system that cannot tell the two apart should not be making either claim.

What I would actually build

Merge every recurring fact on a composite key that carries the period and every qualifier that distinguishes one reading from another. Put intervals on relationships that have a duration. Keep state flags for the questions asked constantly, and keep the interval underneath every flag so the flag can be rebuilt and never has to be believed. Version the population with the same care as the facts, because a rate is only as reconstructable as its denominator.

Then prune. Re-ingestion only adds and updates, so an early unbounded run left me with nineteen years of quarterly history for companies where nothing downstream asks about a quarter from 2008. Note that not all history is equally recoverable: figures pulled from a structured regulatory API can be rebuilt on demand, while a fact extracted from a press release lives on a page that may not be fetchable next year. Those deserve different retention rules, and a single age cutoff applied to both will throw away the half you cannot get back.

And one thing is still open, which I would rather say than leave for a reader to find. A future-dated appointment is stored as future so that "who is the next CEO" is answerable in the same structured way as "who is the current CEO." Nothing flips it to current when the effective date arrives. That needs a scheduled reconciliation pass over the whole graph, and I have not built it. Every temporal design has a moment where the world moves and the database does not notice, and the useful question about a system is not whether that moment exists. It is whether the people running it can tell you where.

Next in this series

Part 3 looks at the other half of the decision: your knowledge graph has a schema, and that is only half of it. Where the schema comes from and how the query gets written are two independent choices, and only one of them is about the schema.

I hope this helps. If you are building one of these, I would like to hear which of the three clocks your system is missing. Reach me at [email protected].