Show the full module text
Module 29 — Knowing It Is Working
The ninety-day review: the four questions that show whether this work is moving, and what a flat quarter can mean.
Three months in, somebody asks whether it is working. Sometimes it is one of you, late at night. Sometimes it is a parent or a friend. This module is about what to look at when you answer, because the thing most people look at first is the one that moves last.
What this is about
Ask a couple three months into this work whether things are better and you usually get a shrug and a qualified yes. Ask them how long it takes to get back to normal after a bad evening and you get a number.
Repair time is the first thing to move, and it is the one to watch.
Couples arrive saying four days. Some months later they say about two hours, and they can feel that change long before any general contentment has caught up. Happiness moves last, and it reads differently depending on which of you is asked and what kind of week you have had.
Like four seeds sown on the same day, repair time, the count of bad evenings, confidence and happiness come up at different rates. Judging the work by happiness, the slowest of them, is how couples decide too early that nothing has changed.
This is the last module of the program, and the one built to be read twice: at about ninety days, and again at six months.
What happens in the session
At about three months there is a session that looks different from the others. Nothing new gets introduced. It is an hour spent looking backward, on purpose. Settling the work in is the part couples skip, and skipping it shows up a year later.
Four questions. Your therapist asks about repair time, the count of bad evenings, your own side of the roadmap, and confidence. The roadmap is the plan of what gets worked on first, built in the session called Your Goals. Each of you rates your own side, separately, one to ten, every six to eight sessions. At ninety days you look at the whole run of those ratings, not the latest one.
Confidence is the question that says when this is finished.
We ask it far earlier than feels natural. An answer that keeps improving is the clearest sign the work is close to done.
The ninety-day review
1 The four questions, asked out loud. Each of you writes your answer down before either of you hears the other’s.
2 The Check-Up, taken again. Our twelve-question questionnaire, taken separately. Then compare the two results, and both against the one you took before you started, if you did.
3 The working map, updated. The one-page account of how each of you is built, from the session of that name, amended rather than rewritten, in your own words. The sentence you crossed out stays visible.
4 The early-warning list, written down. The three things that happen in the two days before a bad stretch. It can only be written while things are reasonably good.
Writing your answer down before you hear your partner’s stops the second answer from becoming a reply to the first. Where one person habitually adjusts, that is a real risk, and usually invisible to the person doing it.
Repair time. The last time it went wrong, how long until we were back to normal? Not until it was resolved. Until we were ordinary with each other again.
The count. How many bad evenings this month? Just the number. We are not going to rate them and we are not going to argue about any of them again.
The roadmap. Here is what I put on my side of the roadmap at the start. Reading it now, out of ten, where is each one? Which of them has not moved at all?
Confidence. If that happened again next month, could the two of us handle it without help? Not perfectly. Just without needing somebody in the room.
Ask these of each other every few weeks between the reviews, and write the answers down with the date. Do not put them on the fridge as a chart. A couple who know they are being scored will manage the number.
The Check-Up, and what it cannot tell you
The Neurodiverse Relationship Check-Up asks twelve questions and counts how many you answer the settled way. That count is the score. It also names one of four patterns. If neither of you took it before you started, take it now anyway. It becomes the baseline for the six-month one.
Said before you take it. It cannot tell either of you whether you are autistic or ADHD. It gives one pattern even when two fit. It weighs every question equally, so eye contact counts the same as how you handle a crisis. And the three bands are our clinical judgment about where the lines go, not thresholds taken from a sample.
The score is the least interesting result; the four or five questions where the two of you chose different answers about the same household are the ones to bring into the session.
The early-warning list
Your therapist asks each of you about the two days in front of a bad stretch, not the bad stretch itself. What are those days like? Couples can nearly always answer. The answers are small and domestic, and once they are written down either of you can name one out loud without it being an accusation.
Agree in advance on the sentence for naming one, closer to “one of ours is happening” than “you are doing the thing again,” so it is a report and not a charge.
The way this goes wrong. Six weeks of things going unusually well, and both of you decide the review is unnecessary. This is how the gains are usually lost, and it is invisible while it happens. The routines drift one at a time, and nobody notices because nothing is going wrong. Then a hard month arrives with none of the structure in place, and the old story about each other is back within a week. The review matters most when it feels least necessary.
Why we do it this way
Repair time first, because it predicts the outcome. In one study of couples followed over four years, what separated the ones who stayed together was not the fight. It was how they were with each other in the conversation after it.1 That is also the part you can change fastest, which is why it moves first.
The count, because memory is unreliable. Reports about feelings you are not having right now are built mostly out of beliefs about yourself, not retrieved from the episode.2 A count is more stable than a rating, and fairer to a partner whose read on their own inner state is shaky. Two people can usually agree on a count.
The commonest false alarm in this program is a score that gets worse while things get better.
When something teaches you about the thing being measured, your inner scale moves, and a before-and-after comparison stops comparing like with like.3
The same evening, two rulers. Ninety days ago, one phrase covered a whole region: bad week.
Now: capacity, sensory load, shutdown, or genuinely upset with each other.
A person with five words reports more than a person with one. So when a score goes the wrong way, the first question is whether you are using the same ruler.
There is a check for this that takes about a minute: ask each other what the first number would have been if you had had today’s words for it back then.
One more caution. Very few everyday questionnaires have ever been checked in an autistic sample, and none of the common relationship measures has, ours included.4 That is a reason to hold every number loosely, not a reason to measure nothing. Looking at a measure helps most in the cases that are not going well,5 because it shows early that something is not moving.
After the session
At six months the wheels come back out. Retake whichever of the three fits, the Autism Trait Wheel, the ADHD Trait Wheel or the AuDHD Trait Wheel from the session called Your Spiky Profile, and put the two shapes side by side. Do not expect a flatter wheel.
The spikes and dips are the profile, not the problem.
Watch the strength ring. The wheel rates each trait twice, once as a challenge and once as a strength. The challenges tend to sit where they were, because they describe how a nervous system is built. What moves, when this is working, is the rating each of you gives the same trait as a strength.
Sometimes ninety days go by and nothing has moved. There are three explanations, and they need different things.
Three explanations for a flat quarter. Design: the agreements were built badly.
Account: the thing being treated is not the thing that is happening.
Capacity: both of you have been running on empty.
Design is the easiest to fix. A wrong account is a conversation to have out loud, not a reason to try harder.
Capacity is the one most often missed. Two people who are both running on empty will describe a relationship with nothing good left in it. Every word of that can be true of the last eighteen months without being true of the relationship. The question that separates the two is about time: was there a year when neither of you was flattened, and what were you like then?
The workbook below is the ninety-day review in your own words. Fill it in now and again at six months.
The one thing, if that is all you have. If this week has nothing in it, answer one question: the last time it went wrong, how long until you were back to normal? Write the number and today’s date somewhere you will find it.
Your workbook
Your answers save to this device only - we cannot see a word of what you write. This is the ninety-day review: the four questions, the Check-Up compared, your early-warning signs, and the honest page for a quarter in which nothing moved.
The four questions
Answer these separately, then compare. Put today's date next to them, because you will do this again at six months.
Last time it went wrong, how long until we were back to normal?
How many bad evenings this month? Just the number
Your side of the roadmap, each goal rated out of ten - and which one has not moved at all?
If that happened again next month, could the two of you handle it without help? — Yes, probably, Some of it, not all of it, Not yet, We have not had one to find out
The Check-Up and the early-warning list
Take the Check-Up separately and bring both results. For the list, describe the two days before a bad stretch, not the bad stretch. Small and domestic is right.
The Check-Up questions where the two of you chose different answers about the same household
Your three early-warning signs, and what you agree to do when one of you names one out loud
If nothing has moved
Answer this only if that is where you are. All three explanations are real, and they need different things.
Which of the three fits best, as far as you can tell? — The agreements were built badly, The account we are working from does not fit us, We are both running on empty and have been for a long time, We do not know, and that is worth saying out loud
Was there a year when neither of you was flattened? What were you like then?
Where this comes from
Every source below was checked against the published record. They are grouped by the kind of evidence they are, so the numbers may not run straight down the page — a number points back to where the source is used in the lesson.
Research discussion
The four measures, the ninety-day and six-month schedule, the retake of the Check-Up and the wheels, the early-warning question and the three explanations for a flat quarter are practice moves. They were developed in use with neurodiverse couples and have not been tested as a package.
The Check-Up is our own instrument. It has not been validated against any external criterion, its three bands were set on clinical judgment, it weighs all twelve questions equally, and it names one dynamic where two may fit. The lesson says all of that before asking you to take it.
The repair-time claim rests on one prospective study1 of seventy-nine couples, seventy-three of them followed up four years later. Emotions in the pleasant conversation after a conflict classified who had divorced better than the conflict conversation did. It is a small sample for a claim about prediction, and the classification rates were calculated on the same data the model was fitted to.
Two sources explain why we ask for counts and a written baseline. The emotional self-report review2 is a synthesis, not a trial. It argues that reports about feelings you are not currently having drift toward beliefs about yourself. The response-shift work3 comes from training evaluation, not couples work, and it is old. It is still the cleanest account of something that happens to nearly every couple in this program at about three months: the intervention changes the scale a person is using.
The depression-questionnaire validation4 is a self-selected online sample of autistic adults and concerns depression, not relationships. It is cited for what it implies: a measure has to be checked in a population before anyone can say what a score means there, and almost no relationship measure has been. The outcome-monitoring review5 is by researchers with a stake in the systems reviewed, the effects are modest, and none of the studies are of neurodiverse couples. It supports a narrow claim: look at the measure, and look soonest at the one that is not moving.
Who this research was done with. The studies behind this module, none of them of neurodiverse couples, drew heavily on white, comparatively well-off, English-speaking participants. If your household carries pressures those samples did not — money, immigration, racism, disability, unsafe housing — the practice still applies, but the room you are practicing in is harder. That is the room, not you.
Peer-reviewed research
1. Gottman JM, Levenson RW (1999) Rebound from marital conflict and divorce prediction. Family Process, 38(3), 287-292. https://doi.org/10.1111/j.1545-5300.1999.00287.x Seventy-nine couples completed a laboratory assessment of three fifteen-minute conversations - events of the day, a conflict discussion, and a pleasant topic - with follow-up obtained from seventy-three of them (92.4 percent) four years later. Emotions coded in the pleasant conversation following the conflict classified divorce at 92.7 percent, against 82.6 percent from the conflict conversation itself. This is the basis for tracking repair time rather than conflict. Limitation: a small sample, classification computed on the same data used to build the model, and the same research group as the 1992 paper cited in Module 7B.
2. Robinson MD, Clore GL (2002) Belief and feeling: Evidence for an accessibility model of emotional self-report. Psychological Bulletin, 128(6), 934-960. https://doi.org/10.1037/0033-2909.128.6.934 A review organizing the evidence on emotional self-report around one distinction: emotion, which is episodic, experiential and contextual, and beliefs about emotion, which are semantic, conceptual and stripped of context. Reports about feelings not currently being experienced draw increasingly on the second. Used here for why a retrospective impression of the last month is a poor measure and a written baseline is a better one. Limitation: a theoretical review rather than a trial, and not specific to couples or to autistic or ADHD respondents.
3. Howard GS, Dailey PR (1979) Response-shift bias: A source of contamination of self-report measures. Journal of Applied Psychology, 64(2), 144-150. https://doi.org/10.1037/0021-9010.64.2.144 The paper that named response-shift bias: an intervention can change the internal standard a person uses to rate themselves, so that the same rating before and after no longer refers to the same thing. Later work by Bray, Maxwell and Howard reported that the resulting loss of statistical power could reach ninety percent. Used here for why a score can move the wrong way while the situation improves. Limitation: 1979, from training and program evaluation rather than clinical work, and never tested in couples therapy.
4. Williams ZJ, Everaert J, Gotham KO (2021) Measuring depression in autistic adults: Psychometric validation of the Beck Depression Inventory-II. Assessment, 28(3), 858-876. https://doi.org/10.1177/1073191120952889 Nine hundred and forty-seven autistic adults completing the BDI-II. Latent trait scores showed strong reliability and construct validity in this sample, with moderate discrimination between depressed and non-depressed participants (area under the curve 0.796; sensitivity 0.820, specificity 0.653). Cited for what it implies: measures require validation in the population using them, and few common questionnaires have had it. Limitation: a self-selected online sample, and about depression rather than relationship functioning.
5. Lambert MJ, Shimokawa K (2011) Collecting client feedback. Psychotherapy, 48(1), 72-79. https://doi.org/10.1037/a0022238 A review of routine outcome monitoring, concluding that collecting and acting on client feedback during treatment improves outcomes, with the benefit concentrated in cases that are not on track. Used here for the narrow claim that looking at a measure early is worth doing, especially the measure that has not moved. Limitation: reviewed by authors with a stake in the feedback systems concerned, effects are modest, individual rather than couples therapy, and no autistic or ADHD samples.
Further reading
• Neurodiverse Couples Counseling Center (2026) The Neurodiverse Relationship Check-Up; the Autism, ADHD and AuDHD Trait Wheel exercises. Practice materials, Neurodiverse Couples Counseling Center. https://www.neurodiversecouplescounseling.com/neurodiverse-couples-check-up The two instruments this module asks you to retake. The Check-Up is twelve questions scored on the count of settled answers, reported in three bands with one of four relationship dynamics named alongside it. The Trait Wheels rate sixteen items in eight areas, twice each, as challenge and as strength. Limitation: both are clinical tools developed in use and neither has been validated against an external criterion; the Check-Up's bands were set on judgment, it weighs every question equally, and it names one dynamic where two may fit.
The measure that moves first is the one nobody asks about
The Neurodiverse Couples Counseling Center works with couples where one or both partners are autistic, ADHD or AuDHD. We measure the work by how quickly the two of you repair rather than by a satisfaction score, and we ask early whether you could handle it without us. Therapy for clients in California, coaching worldwide, all by telehealth. A first conversation costs nothing.
Up next