What calibration means in MemGym
The core idea in one sentence
Calibration is the gap between how confident you feel and how well you actually perform — and MemGym is built to measure that gap, show it to you, and shrink it over time.
Why predict-then-test, not just test

Every card in MemGym asks for something before the reveal: a confidence rating. On a simple 1–5 scale, you commit to how sure you are that you know this, believe this, or can execute this — before you see the answer or attempt the behaviour.
That one extra step changes everything the app collects.
Without a prediction, the app knows whether you got it right. With a prediction, the app knows whether you knew you would get it right. Those are different quantities, and the second one is far more useful.
The reason: you make decisions all day based on your sense of what you know, not based on objective knowledge tests. If you feel 90% confident about something and you're right 60% of the time, the problem isn't that you got 40% wrong — it's that you're walking into situations overestimating your readiness. The overconfidence is the vulnerability, not the gaps themselves.
This is not what Anki does
Anki (and all spaced-repetition systems like it) optimise one question: did you remember this fact? They schedule cards based on whether you got them right, with more forgetting = more review. That's a solved problem, and Anki solves it well.
MemGym optimises a different question: do you know what you know?
To be concrete:
- Anki tells you that you forgot something.
- MemGym tells you where your confidence is lying to you.
An Anki user with a lot of forgotten cards knows they need to review more. A MemGym user who's overconfident on sales objection-handling knows they're walking into calls with a false sense of readiness — even if their recall rate looks fine.
You can have near-perfect recall and be badly miscalibrated. You can also be well-calibrated with mediocre recall — and that's a much better position, because at least you know what you don't know and can prepare accordingly. MemGym targets calibration. Recall is a byproduct, not the point.
The gap is the point, not the fact
The most important number MemGym surfaces is not your score. It's your calibration error: the average distance between your predicted confidence and your actual performance, computed over all your cards and across time.
When that gap is large, it usually takes one of two shapes:
Overconfidence — you predicted high, performed lower. You felt more ready than you were. This is the more common failure mode once you know a domain well enough to feel certain: familiarity reads as competence, so you stop checking. You don't know what you're missing because you're not looking.
Underconfidence — you predicted low, performed higher. You knew more than you thought. This is common in high performers who set extremely high standards for themselves, or in new domains where they don't trust their own pattern recognition yet. The cost is hesitation, excessive preparation, or underselling.
Neither gap is "bad" in a moral sense. Both are costly in a practical sense. A founder overconfident about their understanding of a customer segment will miss signals. An exec underconfident about their communication skills will under-prepare and over-disclaim in board settings. The gap costs you in the specific context where it lives.
The science behind it (briefly and honestly)
The research concept here is metacognitive monitoring — your ability to accurately assess your own knowledge and performance states. Decades of cognitive science show that metacognitive accuracy is a trainable skill: people who regularly predict their performance before feedback improve their calibration accuracy over time.
The prediction-then-test loop is the training mechanism. It works because it forces you to commit to a belief before you get the answer, making the gap between belief and reality viscerally felt rather than abstractly acknowledged. That felt gap is what drives learning to update.
The honest caveat: most calibration research has been done in academic and clinical settings. MemGym applies this to a broader set of performance domains — sales calls, board meetings, deep work, behavioural intentions. The mechanism has strong theoretical support; the specific application to high-performer professional contexts is less studied. We are tracking this carefully rather than overclaiming.
How to read your calibration data

MemGym breaks down calibration in three ways:
Overall calibration error. A single number representing your average prediction-vs-actual gap across all cards. Smaller = better calibrated. This is your headline metric: it tells you whether your self-model is tight.
Directional bias. Are you systematically overconfident or underconfident? Your calibration screen shows which direction your errors run. Most people have a consistent directional lean, and knowing which way you lean is as valuable as knowing the magnitude. If you know you tend to overestimate your readiness, you can build a habit of adding preparation time before high-stakes situations.
Per-context breakdown. This is where calibration becomes actionable. You might be perfectly calibrated on conceptual understanding and badly overconfident on in-the-moment behavioural execution. You might be well-calibrated in solo deep-work contexts and overconfident in adversarial ones (presentations, negotiation, conflict). The context tags on your cards make this visible. The question to ask: where is my confidence lying to me? That's the context to pay most attention to, because it's the context where you're making the worst decisions about your own readiness.

A brief example
You have a card: "I stay calm and direct when someone challenges my reasoning in a group setting." You predict 4/5 confidence. You attempt it in your next meeting and rate the actual performance a 2. That's a gap of 2 — you were significantly overconfident.
Over ten such attempts, your average gap in interpersonal-challenge scenarios might be +1.8, while your average in solo execution contexts is +0.3. The calibration view makes this legible: your self-model is accurate in solo work and meaningfully off in high-stakes social contexts. You can now deliberately practise in exactly that domain, and track whether the gap closes.
That's the loop. Predict. Test. See the gap. Practise the gap. Measure again.
What calibration is not
It is not a score to optimise by gaming the prediction. If you predict 1/5 on everything, your calibration error will look great — and you will have learned nothing. The prediction is a commitment to honesty about your self-assessment; the value comes entirely from making it real before you know the outcome.
It is not a measure of how much you know. You can know a great deal and be poorly calibrated. You can know less and be precisely calibrated. The target is accuracy of self-knowledge, not volume of knowledge.
It is not a judgement about your competence. A large calibration error means your self-model is miscalibrated for the domain — that's a navigational problem, not a character problem. The whole point is to surface it so you can correct it.
The gap is a signal. MemGym's job is to make it visible, keep tracking it, and turn it into the next thing you practise.