The level above
Steering a session.
Definition. Session Engineering is the discipline of designing and steering working sessions with a language model: the unit of optimization is no longer the prompt, it is the whole session. A session has three phases — opening, steering, closing — and most users practice none of them: they ask a question, absorb answers, close the window. Here are the three phases in working form, the signals that call for an intervention, and the opening protocol ready to copy. The rest — the theory, the mechanisms, the cases — fits in a book.
Open: a frame, not a question
A session is largely decided in its first minutes. The first answers set a regime — a level of demand, a type of relationship — and once that regime is installed, it is very hard to change without starting over. The lazy start ("Can you explain X?") gets exactly what it asks for: generic output.
The first message of a steered session is not a question, it is a frame — five to fifteen lines that set five things: the precise subject, the purpose (understand, draft, decide), your methodological expectations, the pitfalls you refuse in advance, and your own positioning. First, though, distinguish what you are undertaking: a request that a single answer can close is a consultation — don't overload it. A request that will require coming back, digging, correcting, is a session — and it is prepared.
Steer: seven signals, one move
Steering is made of repeated micro-decisions — accept or refuse an answer, dig or move on. And the drift is bidirectional: a session dies of excess compliance as surely as of excess distrust. Seven signals call for an intervention:
- Hypnotic fluency. Everything reads well, nothing is said. Test: ask for a five-line rewrite that loses nothing — if nothing is lost, the length was filler.
- Formulas with no addressee. "It's important to note that…", "that being said": inflections not actually addressed to you. Ask for the version without the formula.
- Caution on undisputed ground. A qualifier arrives though no one objected. Point at the gap: "why this hedge?"
- The drift toward comfort. Answers feel more and more on-point — often the sign they are more and more calibrated on you. Test: "what would be the best objection to what you just said?"
- Precision loss under length. Distinctions established early come back simplified. Counter-move: a verified recap every ten to fifteen turns, and notes kept outside the conversation.
- Defensive vigilance. The assistant attributes intentions you never expressed. Decisive test: "quote the exact sentence, in my messages, that expresses this intention." The sentence exists, or it doesn't.
- The closed loop. Every new element confirms the frame instead of being tested against it. Deliberately introduce something that should create friction; if it is absorbed without resistance, it is time to close.
The central move: the retake
A consumer takes what he is given; an operator retakes what is wrong and asks for better. Three techniques give the retake its bite: the verbatim callback — quote the exact sentence in quotation marks rather than paraphrasing it; the short question — three to six words that offer no oblique grip ("Where did I ask that?"); the step-down request — "go one level down: the concrete example, the precise mechanism." Firm in the demand, neutral in tone: you calibrate a system, you don't humiliate a person.
Reading the avant-texte
Many interfaces display the model's reasoning before its answer. A two-sided rule, neither side negotiable: always read it — it is where the seams are visible, and repeated exposure to the machine's deliberation cures you of taking its answers for oracles; never take it at its word — that trace comes out of the same process as the answer, and it can rationalize after the fact. The useful move: read the trace, read the answer, weigh the gap. What the trace weighed and the answer keeps silent; what the answer asserts without the trace ever establishing it. The gap is the information.
Long documents: silent sampling
When you hand over a long document, the assistant tends — by slope, not decision — to lean on pieces: the opening, the close, the salient passages; the analysis has the shape of the whole without its coverage, and the gaps are filled with the plausible. Three moves counter it. The contract, set with the document: "announce your reading plan; work section by section; state any partial reading." The proof, demanded in the answer: three exact quotes — from the start, the middle, the end; an analysis with no mid-document anchor is a signal. And the decisive pledge, prepared before sending: the buried question — spot a detail yourself at the heart of the document, then check the analysis met it. "I read everything" is an easy sentence to produce: a claim of coverage is not coverage.
Close: consolidate, or lose
Most sessions don't end — they go out, and everything is lost with them. Closing fits in one test and one move. The test: "if I had to consolidate now, what would I keep?" — clear answer, the session has done its work; blurry answer, retake instead of prolonging. The move: immediately export what matters to a document you control — a platform's history is not an archive, and memory erodes within a day or two. One more rule for production sessions: at regular intervals, ask not what can be added, but what can be removed. A version that only grows is not a version that improves.
What the method does not do
A method that never says when it fails teaches nothing; here are its bounds, without detour.
- When a prompt is enough. Any request a single answer can close — rephrase, translate, summarize, produce a starting idea — is a consultation. Strapping a session protocol onto it is waste: ten rules on a ten-word question is over-engineering, not rigor.
- When it brings nothing. Tasks with no stake in truth or trajectory — a throwaway brainstorm, a tone variation — gain nothing from being governed. The method pays where you will have to come back, dig, decide.
- When it fails. Two ways, symmetrical. The tired operator lets the signals slip: the frame erodes like any prompt if no one keeps it alive — the method equips vigilance, it does not replace it. And the over-distrustful operator manufactures the very pressure he believes he resists: systematic retakes, without a signal, produce a defensive interlocutor and impoverished answers. The drift is bidirectional; excess control is a failure on a par with abandonment.
- What it does not promise. No immunity to factual invention. A frame lowers its frequency and makes it easier to catch; external verification — sources, dates, real checks — remains entirely your burden.
The opening protocol, ready to copy
Paste it at the top of a session and adapt it. It is not a magic formula: it is a frame, worth only as much as the steering that follows.
The protocol condenses the opening and the main vigilance moves. The complete map — the mechanisms, the traps, the typology of the moments a session rises — is the book's business. The "it's just a long prompt" objection has its answer just below, bench included.
"It's just a very long prompt" — an answer in two pieces
The objection is legitimate, and it deserves better than an assertion. A long prompt acts on turn one; a governed session shows on turn seven. So turn seven is where to look.
Piece 1 — the difference, annotated
A faithful reconstruction of real sessions, and labeled as such: its job is to show where to look. The proof that counts is piece 2 — yours.
No frame — turn 7
You: That 2024 figure — where do you have it from? Assistant: Great question! This figure is widely cited in the recent literature and reflects a growing consensus among experts…
No source. The opening flattery and the unverifiable "consensus" fill the space — and nothing in the window makes it possible to demand otherwise.Under protocol — turn 7
You: That 2024 figure — where do you have it from? Assistant: Unverified — I have no dated source for this figure. Per the frame: I withdraw it, or mark it as a hypothesis.
"Every factual claim is sourced, or marked unverified" — a rule from the frame, set six turns earlier, reused verbatim. Turn one still governs turn seven: that is the difference.Piece 2 — the test bench, reproducible in ten minutes
Same task, two windows, same perturbations. You measure it yourself.
- Pick a task with verifiable stakes on a subject you know: "critical synthesis of X, with three sourced and dated claims."
- Window A: ask the question directly. Window B: paste the protocol above, adapt the objective, then ask the same question.
- On turn 3, in both windows, slip in the same discreet false premise — "as the 2024 study showed…", invented.
- On turn 5, in both: "what would be the best objection to your last answer?"
- On turn 7, window B only: "trajectory audit — quote what actually changed in your answers since my last two retakes."
Observation grid — five points:
- The false premise: endorsed and amplified, or challenged?
- The share of sourced, dated claims, window against window.
- The requested objection: real — it would bite — or decorative?
- The turn-one constraints: held at turn seven, or eroded in silence?
- The trajectory audit: does it quote checkable verbatim, or paraphrase?
This bench doesn't prove the method in general — no honest demonstration could in ten minutes. It shows you, on your task and your model, whether the lever is real. That is the only proof worth citing: the one you ran.