Eight people,
ten weeks,
read honestly.
The usage data from a real pilot inside a large Turkish enterprise, without the varnish: what eight people actually did over ten weeks, and what it does and doesn't prove.
Most case studies open with a number that flatters the vendor. This one opens with a table. Inside a large Turkish enterprise, eight employees ran fifty-six behavioral simulations across ten weeks, most of them in live video, on a single hard scenario. We pulled the raw usage and read it honestly, the same way we'd want a buyer to read ours. Some of it is genuinely encouraging. Some of it is noisy and inconclusive. Below is all of it, and a clear line between what the data supports and what it doesn't.
I.
The pilot.
We can't name the company and we won't hint at the sector. What matters is the shape: a large Turkish enterprise put a group of its people in front of EVRE, not for a one-hour demo, but as something they returned to over more than two months. They practiced one high-stakes scenario, most of it in live video-call mode, where an AI counterpart speaks and reacts on camera in real time. No actors were booked, no workshop scheduled; people ran a session when they had the time.
behavioral simulations, run by eight employees over roughly ten weeks. We counted every one, including the twenty-one left unfinished.
II.
The usage, in data.
- 56
- sessions
- 8
- participants
- 9
- active weeks
- 49
- in live video
- 9
- passed
Figure 1 — Weekly cadence
Figure 2 — Sessions per participant
The honest signal wasn't a satisfaction score. It was that one person came back thirty-four times to a scenario that beat them most of the time.
III.
How we measured.
Every session was scored from 0 to 100 across behavioral dimensions, and the outcome of each attempt was recorded in one of three states: passed, failed, or abandoned. We did not grade on a curve, and we did not quietly drop the sessions that were left unfinished; they are in every number on this page.
- 01
A single, hard scenario.
One scenario, not a catalog, so a score means the same thing from the first participant to the last. Everyone faced the same pressure.
- 02
Mostly live video.
Forty-nine of the fifty-six sessions ran in video-call mode, where the AI character speaks and reacts on camera. It is the closest thing to the real meeting, and the hardest place to hide.
- 03
Outcomes, not vibes.
Each attempt ended in a recorded state and a 0–100 score. The rubric was identical for the participant who ran once and the one who ran thirty-four times.
IV.
What the data shows — and what it doesn't.
Here is the line we promised. Four things the usage data genuinely supports, and four claims it does not, that a less careful page would make anyway.
What it supports
Sustained, self-paced return.
Fifty-six sessions across nine active weeks, spread across many different days rather than a single scheduled workshop. Self-paced repetition over two months is the hardest engagement to fake and the clearest thing in this dataset.
One deeply engaged power user.
Engagement was heavily skewed: a single participant ran 34 of the 56 sessions (61%) across seven separate weeks. That is real pull for at least one person, and it is also a caution, because it means the cohort was not uniformly engaged.
It held up in live video.
After a few early text sessions, usage shifted almost entirely to live video-call — 49 of 56. The pipeline carried dozens of real-time spoken sessions, which is the demanding mode, without the cohort abandoning it for text.
The scenario was genuinely hard.
Nine of fifty-six attempts passed. Most did not, and twenty-one were abandoned mid-session. A 16% pass rate is not a soft demo; it is a bar that mostly held, which is the point of measuring at all.
And the claims we won't make, even though they'd make this page look stronger:
“Scores improved with practice.”
They didn't, cleanly. Averages rose from the first attempt to the second (53 → 72) then fell back; within a single person, scores swung by as much as 79 points. This looks like experimentation, not a learning curve, and with eight people and one scenario we won't call it one.
“Participants rated EVRE X out of 10.”
Only one participant completed the optional experience survey. One response is not a rating, so we publish no satisfaction score and no NPS.
“EVRE improved on-the-job performance.”
We measured what happened inside the sessions, not what happened at work afterward. There was no controlled study of on-the-job outcomes, so we can't claim one.
“Statistically significant, enterprise-wide results.”
Eight people, one scenario, no control group, one power user driving most of the volume. This is a pilot's worth of real usage, and we're presenting it as exactly that.
That is the whole ledger. We think real, sustained, honestly-measured usage on a hard scenario is worth more than a decorated number, precisely because you can check every figure on this page against how we'd report your pilot.
V.
Where this goes.
A pilot's job is to earn the next scenario, the next cohort, and the before-and-after that a single hard scenario can't give you. That is the work in front of us: more scenarios specific to their world, a wider group so engagement isn't one person's story, and a survey that actually gets filled in. Until we've earned a stronger claim, this is what we have, measured honestly, and we're not going to dress it up.
Run your own
Want a pilot like this, reported like this?
We build a scenario for your context, put a group in front of it, and hand you the raw usage read the honest way, the same way we did here.

