Skip to main content
TEEI in Practice
Engineering note · AI & Automation · 1 October 2026 How Learners Get Sentence Feedback in Practice Companion

TEEI in Practice

Engineering noteAI & Automation

How Learners Get Sentence Feedback in Practice Companion

Engineering note on the TEEI Language writing room: one sentence, one piece of advice, and a correction checked against her words before it counts.

On this page
  1. Key facts
  2. The desk and the four rooms
  3. One sentence, end to end
  4. Anchoring: a correction is evidence only when it points at her words
  5. The taxonomy travels with the request
  6. What we record, and what we refuse to infer
  7. Steering, not drilling
  8. The level conversation
  9. Say it: a contract, not an improvised chatbot
  10. House rules
  11. Verify this
  12. Frequently asked questions
  13. Notes and sources
A writing exercise in Practice Companion: the learner's sentence with one word marked, and a short correction beneath it.

Practice Companion is the part of TEEI Language a learner uses between her conversations with a volunteer. This note follows one written sentence through it, and explains the decisions on the way: why she gets one piece of advice rather than a chat, why a correction has to point at her own words before we count it, why progress is a direction and not a number, and why the whole thing has no face.

Key facts

  • One companion with one memory: her level, her situation, the shelf of lines she has kept, and one way to a person from every room.1
  • Four rooms around one desk: Write, Listen, Say it, and the level conversation. A line kept in one room lies on the shelf in the others.
  • In the writing room she writes a sentence and gets one piece of advice. The model answers in a fixed structure, a guardrail judges that structure before she reads it, and the correction is checked against her exact words before it can count as learning evidence.
  • Nothing scores her. There are no counters, streaks or percentages. The level calibrates how the companion speaks to her and is never a requirement or a certificate.
  • We look for a corrected form in her own later sentences, in four windows after the correction. Progress between two level conversations is reported as a direction, never a number.

The desk and the four rooms

The companion is one desk with four rooms around it. In Write she types a sentence about her own situation and gets one piece of advice back. In Listen she opens a catalogue of short English dialogues about everyday situations, booking a doctor, a job interview, registering an address, a flat viewing, written and reviewed by people and recorded once. She listens first; the words are hidden until she asks for them. Three questions are answered against the recording and never scored, and a glossary gives each hard word in English and Ukrainian. In Say it she opens a situation, a receptionist holding two appointment slots, a short job interview, a new colleague on her first day, or a situation of her own built from the goal in her profile, and talks it through with a voice on the other end, speaking or writing, with as much support as she wants. Meaning counts and grammar does not: “No, I working morning. Can after three?” is a complete, correct turn. In the level conversation she answers six written tasks and is read, not taught.

The rooms share one shelf. A line she keeps in any of them is there in the others, and she can take it into her next booked session, where it is the only thing the volunteer sees. Every room has one way out to a conversation with a person.

One sentence, end to end

Take the writing room. She types a sentence, and this is what happens before the reply appears.

  1. Her browser sends the sentence to the TEEI Language Worker. The Worker checks her session, checks that her programme opens the writing room, applies the rate limits and the length cap, and loads her last turns and the lines on her shelf.
  2. The Worker signs a short-lived token and calls the practice function in TEEI’s AWS account. The Worker holds no model keys; the signed bridge is the only way across.
  3. The function calls Claude on Amazon Bedrock with tool use forced: the model must answer in one schema, a reply, at most one correction that quotes her words, and any vocabulary it offered. The taxonomy of correction categories travels with the request, so the model labels from a fixed list rather than inventing a label.
  4. A Bedrock guardrail with a PII policy judges the structured answer before anything returns.
  5. Back in the Worker, the correction is checked against the sentence she wrote. Anchored corrections can count as learning evidence; unanchored corrections are kept for model-quality checks.
  6. The Worker sends the completed reply to her browser as a paced stream of newline-delimited JSON.

The stream is a handful of event types. The following illustrative stream shows feedback on the sentence “After the night shift I am boring and I go home.”:

jsonl
{"type":"delta","text":"That sounds like a long night. One small thing: "}
{"type":"delta","text":"after a shift you are bored, or more likely tired."}
{"type":"correction","original":"I am boring","suggested":"I am bored","oneLineExplanation":"Boring describes the thing; bored describes how you feel.","startOffset":22,"endOffset":33,"category":"word-choice"}
{"type":"vocabulary","word":"shift","definitionInContext":"The hours you work in one go, for example a night shift."}
{"type":"done","conversationId":"c_01HZ..."}

The offsets on the correction are what let the browser put a copper mark under the right words. If the anchor did not hold, the correction event still arrives, because the advice is still real help, but it carries no offsets and the browser shows the repair without drawing a mark rather than drawing one in the wrong place. A mark_confirmed event follows later in the session when a corrected form appears again in her own sentence; the evidence row records how much help was in front of her.

Anchoring: a correction is evidence only when it points at her words

The model is told to quote her verbatim. When it does not, the correction is a claim about words nobody typed, and counting that as learning evidence would be counting a hallucination as a measurement.

So the Worker, not the model, decides where a correction sits. For an anchored correction, the stored quotation is the slice of her own sentence at the resolved offsets, never the string the model handed back. If anchoring fails, the Worker keeps the model’s quotation with its failure status for model-quality checks, and it does not count as learning evidence. That makes “the original text matches the learner’s message exactly” true by construction rather than by assertion.

ts
export type AnchorStatus = 'valid' | 'ambiguous' | 'not_found' | 'invalid_payload';

export function anchorCorrection(learnerTurn: string, quoted: unknown): AnchorResult {
  if (typeof learnerTurn !== 'string' || learnerTurn.length === 0) return INVALID;
  if (typeof quoted !== 'string') return INVALID;

  const needle = quoted.trim();
  // An empty quotation, or one longer than the message it claims to come from, is a
  // malformed payload rather than a failed search: there is nothing to look for.
  if (needle.length === 0 || needle.length > learnerTurn.length) return INVALID;

  const exact = occurrences(learnerTurn, needle);
  if (exact.length === 1) return span(learnerTurn, exact[0]!, needle.length);
  if (exact.length > 1) return AMBIGUOUS;

  // Case is the one difference that still leaves the span exact: the model often re-cases a
  // quotation it lifts from the start of a sentence.
  const lowerTurn = learnerTurn.toLowerCase();
  const lowerNeedle = needle.toLowerCase();
  // Case folding can change length in some scripts, which would make the offsets point at
  // the wrong characters. Refusing to guess is the whole point of this module.
  if (lowerTurn.length !== learnerTurn.length || lowerNeedle.length !== needle.length) {
    return NOT_FOUND;
  }

  const insensitive = occurrences(lowerTurn, lowerNeedle);
  if (insensitive.length === 1) return span(learnerTurn, insensitive[0]!, needle.length);
  if (insensitive.length > 1) return AMBIGUOUS;

  return NOT_FOUND;
}

Three decisions are visible in that function. A quotation that appears twice in her sentence is ambiguous, not “probably the first one”: the model gives no position, and picking an occurrence would be inventing the one fact needed to make the row valid. The matching stops at case: anything fuzzier, stemming, collapsing punctuation, nearest match, would start guessing which words she meant, and a guess that lands on the wrong span is worse than no evidence at all. And nothing is dropped. An earlier version ran the check inside the AWS function and discarded whatever failed it, which destroyed the one signal that tells us how often the model invents a quotation. Now the anchor returns a status, and the status decides what the row may be counted in.

The taxonomy travels with the request

Trackable corrections use a versioned taxonomy of 27 correction categories, each with a one-line gloss the model reads: a real word used with the wrong meaning (“I am boring”), a word borrowed from her own language that means something else here, and so on. The list goes with the request, so the model chooses from it rather than inventing labels, and the version is stored on every item. Two corrections months apart are comparable because they were labelled from the same list, and when the list changes, the version on the row says which one applied. Unclassified corrections are kept for model-quality checks without becoming learning items.

What we record, and what we refuse to infer

The evidence layer is a set of nine tables: her turns, the anchored corrections, the learning items they belong to, the evidence for each item, the interventions we made, the observation opportunities we opened, the evaluation runs, the observations a person made in a live session, and the links that take a kept line into a booked session.

The constraints on those tables are where the method lives. Evidence about something she produced must carry her exact span; a row without one cannot be written. “Correct without support” cannot coexist with any support above none, because a form she produced after a hint is not independent use. A presented_at timestamp means we steered the conversation towards that form; it is recorded when the steering happened, and it is anchored on when her reply arrived, not when it was processed. A chance that expired, or that was never presented, is never a failure. And any rate on the board returns null under its minimum sample, which the board shows as “not enough data” rather than as a confident zero.

Steering, not drilling

A stored correction opens four windows in which we would like to see the form again.

Table 1
WindowOpensStays open for
Same sessionAt once2 hours
Day 11 day after the correction3 days
Day 77 days after the correction14 days
Day 3030 days after the correction60 days
The four observation windows that follow a correction in the writing room, read from the scheduling source on 1 October 2026.

When a window is due and she is writing, the conversation steers towards a moment where that form comes up naturally, without naming it, at most once per session. She is never handed a drill. We record every reuse with how much help was in front of her. In the same session the correction is still on her screen, so reuse there is recorded as produced with the answer shown. In a later window, reuse in a turn we steered is recorded as produced with a hint, and reuse that comes up with no steering is recorded as produced without support. If the moment never comes, nothing is recorded against her. Over the four windows this gives a reading of same-session uptake, delayed retention and the time it took to reach independent use, each one over a counted sample, and each one null until the sample is large enough to mean anything.

The level conversation

The level conversation is the companion’s other mode. Practice teaches, with help; the level conversation measures, with no teaching. She answers six written tasks. An interviewer with its own prompt asks them, and a judge reads her answers against a task plan that travels with the request, the same way the taxonomy does, and anchors every observation in a quotation from her own answer. The reading is stored with its source marked as a conversation, and provenance is kept with it.

What the level does is calibrate: it sets how the companion writes to her, from the first reply on, and the dialogues in the listening room are grouped around it as “At your level”, “A step up” and “Easier”. What it never does is gate. It is not a requirement for any room, not a right to anything, and not a certificate. When she takes a second conversation later, the two readings are compared, and the result is a direction, never a number. A level she measured on the web test before signing up flows into the companion at her first login, so the first reply is already written at her level.

Say it: a contract, not an improvised chatbot

The three fixed situations in the speaking room are built on a scenario contract rather than on an open conversation. Each of them declares its goal, its facts, its lines, its help, its rules and the intents the model may return, with a pure step function that moves the conversation on. The facts belong to the server: a line names a fact by its id, and the engine refuses an id the contract does not carry. In those situations the model varies nothing she hears. It returns an intent, only the intents the contract declares survive, and when the model’s reading contradicts what she actually said, the engine refuses the reading and asks her again rather than booking a time she never chose.

Two rules hold throughout. The other person in a situation never asks for her data, and help never invents her answer: it gives her a line she can say, not a story about her. A situation of her own has no contract to pick a line from, so the model writes the other person’s line behind the Bedrock guardrail, and the room refuses to open that situation without one. A personal detail the guardrail detects in her line is masked, and the masked form is what the model reads on every later turn. ElevenLabs provides the voices. TEEI stores the speaking-room transcript as text, not audio.

House rules

  • Never a persona. The companion has no name, no face and no avatar. Its presence is a copper mark in the margin. The waiting state is the mark breathing, never the word “Thinking”.
  • Nothing is seeded into her turns. What she wrote is what she wrote.
  • No counters, XP, streaks or mastery percentages appear in the learner’s rooms. A rate under its minimum sample reads “not enough data”.
  • Colour separates feedback from later use. A mark in the action colour reads as something to press; a mark in the earned colour reads as a corrected form she used later.
  • One way to a person in every room. The way out of the companion is always a conversation with a volunteer.

Verify this

Frequently asked questions

Is Practice Companion a chatbot?

No. A learner writes a sentence and gets one piece of advice back, with at most one correction. A correction is marked in her sentence only when its quotation can be matched to her words. There is no persona, no name and no face. The companion exists for the days between her conversations with a volunteer, and every room has one way to a person.

Does it replace the volunteer?

No. The conversation with a volunteer is the point of TEEI Language. She practises with the companion on the days in between, and a line she keeps in any room can travel with her into her next booked session, where it is the only thing the volunteer sees.

What does it store about a learner?

Her written turns, correction records and their anchor status, the lines she chose to keep, her level readings and, in the speaking room, the transcript of what was said. The speaking room stores no audio. A Bedrock guardrail checks for personal details; a value it detects in her own spoken situation stays masked in every later request to the model. The other person in a situation never asks for her data.

Which AI models does it use?

Anthropic's Claude models on Amazon Bedrock, inside TEEI's own AWS account, for the writing room and the level conversation, judged by a Bedrock guardrail before anything reaches her. ElevenLabs for the voices. The listening dialogues are written and reviewed by people and recorded once; nothing in that room is generated while she listens.

Does the level count as a certificate?

No. The level calibrates how the companion speaks to her. It is never a requirement, a right or a certificate, and the companion never shows her a score, a percentage or a streak.

Notes and sources

  1. TEEI Language programme page, theeducationalequalityinstitute.org/programmes/education/teei-language/, read 1 October 2026. ↩

More in AI & Automation