Jev shadow mode still sends your data
In Jev shadow mode your local ranking is kept and Jev's scores are only recorded. The query and candidate memories still go out to TypeSafe AI.
By Heartwood MemoryPublished . Dates are Pacific time.
In Jev shadow mode, a memory tool sends each eligible recall to Jev, records Jev's scores, and returns your results in the order your local ranker chose. A trial changes nothing about what your agent sees. It sends exactly what the judge sends when it is fully on: the recall query and the text of the candidate memories, to TypeSafe AI. Only the ordering stays local. Heartwood Memory's trial works this way, and it needs the same written opt-in as turning the judge on.
As of October 1, 2026, Heartwood Memory's Jev judge for recall ranking is available to hosted Team and Professional organizations that opt in; it is off by default.
This post is about method: how to try a judge model on real recall, and what the trial itself costs you in data. It has no results in it, ours or anyone's. What is in each request, and what Heartwood holds back, is covered in What Jev sees when it ranks agent memory and on the Jev integration page.
What is Jev shadow mode?
It is the middle of three settings. With the judge off, only the local ranker runs. In a trial, both rankers score the same recall, the caller gets the local order, and Jev's scores go into the record. With the judge on, Jev's order is used whenever Jev answers in time.
"Shadow" describes the effect on your results. It says nothing about your data.
Does shadow mode send my data to TypeSafe AI?
Yes. A trial sends the same request as the live setting, to the same place. Jev can't score a memory it hasn't read, so there is no way to run a real trial without sending real candidates.
The same checks run first. A recall with a candidate above your organization's egress ceiling, a candidate labeled as personal data, or anything shaped like a secret is held back in a trial exactly as it would be with the judge on. Nothing is sent for that recall, and the record says why.
So treat a trial as a production data flow, because it is one. In Heartwood the admission check for a trial is the same check as for the live setting: a Team or Professional plan, a written opt-in, and the processor disclosure on record. A trial is not a lighter agreement.
What stays local during a trial?
- The order of the results your agent receives.
- The name on each result of the ranker that ordered it, which is the local ranker.
- Every recall that can't be sent. Those never leave, in a trial or after it.
What does not stay local: the query and the text of the candidates, for every recall that passes the checks.
Can a trial slow recall down?
It can add waiting time. Heartwood runs the two rankers at the same moment, and the recall waits for Jev's answer or for a short fixed deadline, whichever comes first. That is true in a trial too, even though a trial doesn't use the answer for ordering. It needs the answer for the record.
We haven't published speed figures for the judge, and this post has none. Each trial record holds how long each ranker took, so you can measure it on your own traffic.
What does a trial record?
For each recall that was judged, the record shows that it ran as a trial, that the local order was the one used, a probability for each memory Jev scored, the local ranker's score for each candidate, how long each ranker took, the egress decision, and a hash of the request. The judge's record holds no query or memory text. Scores are listed by memory ID.
For a recall that was held back, the record gives the reason. For a recall where Jev was slow or returned an error, it gives that reason, and only the local scores.
Jev scores the leading candidates from Heartwood's own retrieval. The local ranker scores all of them. Keep that in mind when you compare. What the record is for, and what it can't prove, is the subject of What to record when a model ranks memory.
How do I compare the two rankers?
- Write the rule down first. Before the trial starts, decide what result would make you turn the judge on and what would make you stop. A rule chosen after you have seen the numbers is not a test.
- Compare recall by recall. Each trial record holds both rankers' scores for the same query and the same candidates. Compare the two orders on the same recall. Don't compare an average from one week with an average from another.
- Bring your own answer key. The record tells you how each ranker ordered the candidates. It doesn't tell you which order was right. Because it holds IDs and scores and no text, you look the memories up in your own store and judge them yourself, or use questions whose right answer you already know.
- Compare like with like. Only the leading candidates have a Jev score. Compare the two rankers over the candidates both of them scored.
- Count what couldn't be sent. Add up the held-back recalls by reason. If a large share of your recalls is over the ceiling or carries a personal-data label, the judge will rank fewer of them, however good its ordering is.
- Say how many. Report how many paired recalls your conclusion rests on. A trial on light traffic can produce too few to decide anything.
We haven't published what our own trials showed. When we publish figures for the judge, they will come from Heartwood's pre-registered recall test and will sit next to how each one was measured, on our measurement receipts.
How does a trial end?
An operator changes the setting, to on or to off. The setting is operator configuration that the service reads at startup. A recall request can't change it, and neither can an agent.
A trial never promotes itself. Nothing moves an organization from trial to on because a number looked good. Someone decides, and the decision is yours to make with the rule you wrote down in step one.
Ending a trial stops new sends. It can't pull back what was already sent. That data stays under TypeSafe's terms, which you can read at TypeSafe AI legal.
What should I ask before any shadow trial?
These apply to any tool that offers a shadow or evaluation mode, ours included.
- Does the trial send the same data as the live setting?
- Does it need the same sign-off, or is it treated as a lighter step?
- Do the same hold-back checks run before a trial request goes out?
- Does the trial add waiting time to recall?
- What does it record, and does the record hold text?
- Who can end the trial or promote it, and can anything promote it automatically?
If your organization is on a hosted plan and wants to start with a trial, ask to turn on the Jev judge. If none of your memory may leave your environment, self-host instead: the quickstart runs governed agent memory on your own machine with the local ranker and no outside calls. Checking AI-written memories with Jev is planned and not available yet.
Jev is a model from TypeSafe AI, Inc. Heartwood Memory is made by Edukas Solutions LLC and is not affiliated with or endorsed by TypeSafe AI.
Questions
Does Jev shadow mode send data to TypeSafe AI?
Yes. In a trial, Heartwood Memory sends the recall query and the candidate memories to TypeSafe AI exactly as it does when the judge is fully on. Only the ordering of your results stays local.
Does a shadow trial change what my agent retrieves?
No. During a trial the results come back in the order Heartwood Memory's local ranker chose. Jev's scores are recorded and not used for ordering.
Does a trial need a separate opt-in?
It needs the same one. Heartwood Memory runs the same admission check for a trial as for the live setting: a Team or Professional plan, a written opt-in, and the processor disclosure on record.
Can a trial switch itself to the live setting?
No. The setting is operator configuration read at startup. Nothing promotes a trial automatically, and a recall request can't change it.