ONESTEP SCRIBE

PHQ-9 and GAD-7 documentation: scoring ranges, interpretation, and note examples that hold up

By Selam Koch, PMHNP-BC, founder of OneStep Scribe and practicing psychiatric nurse practitioner. Published July 2026.

Open ten psychiatric progress notes and most will handle standardized scales the same way: a PHQ-9 or GAD-7 total on a line by itself, connected to nothing. The score gets charted, the box gets checked, and the note moves on. That is the easy half, and it is the half that does not count. The score is not the work; the interpretation is. A number without a prior number, a severity band, and a resulting decision is trivia. A number with all three is measurement-based care, and the difference is about one extra sentence. This guide covers the scoring bands for both instruments, the trajectory conventions that give the numbers meaning, and the documentation patterns that make these scales worth the thirty seconds they take to review.

Scoring at a glance

The PHQ-9 is a nine-item depression measure. Each item is scored 0 to 3 over a two-week recall window, for a total of 0 to 27. The GAD-7 is its seven-item anxiety counterpart, also scored 0 to 3 per item over two weeks, with a total of 0 to 21. Both instruments were developed by Drs. Kroenke, Spitzer, Williams and colleagues, and both are free for clinical use with no permission required.

PHQ-9 totalDepression severity
0 to 4Minimal
5 to 9Mild
10 to 14Moderate
15 to 19Moderately severe
20 to 27Severe
GAD-7 totalAnxiety severity
0 to 4Minimal
5 to 9Mild
10 to 14Moderate
15 to 21Severe

For the GAD-7, a score of 10 or higher is the common convention for the threshold that prompts a fuller anxiety evaluation. It is a flag for closer attention, not a verdict.

The conventions that make scores meaningful

None of what follows is a rule. These are widely used research conventions, and they are worth adopting because they are what a knowledgeable reviewer, a collaborating physician, or the next clinician reading your chart will recognize.

Below those margins, the honest word is stable. A PHQ-9 that drifts from 14 to 12 is not "improving"; it is within the noise until it clears the margin, and charting every small wobble as better or worse builds a trajectory your own numbers will eventually contradict.

One thing stated plainly: these are severity and monitoring tools, not diagnostic instruments. Diagnosis is a clinical judgment made against DSM-5-TR criteria, informed by history, examination, and course. A PHQ-9 of 18 supports a severity impression; it does not establish major depressive disorder, and a score never diagnoses anyone.

Item 9 is different

PHQ-9 item 9 asks about thoughts of death or self-harm, and it is the one item where the total score does not matter. A patient can total 6 overall and still endorse item 9. Any positive response, even a 1, changes what the visit owes the chart.

Any positive response on item 9 requires a documented risk assessment the same visit: what you asked, what the patient said, your clinical assessment of risk, and the plan. An item 9 positive sitting in the chart with no risk documentation in the note is among the most serious findings in a malpractice or board review context, because the screening form itself documents that you were on notice.

Keep the perspective, though. A positive item 9 is common in depression care, and it is a prompt for assessment, not automatically an emergency. Passive thoughts without intent, plan, or preparatory behavior, in a patient with intact protective factors, is a frequent and defensible finding. The problem is never the positive response. The problem is the silence after it.

Documenting scores so they mean something

The pattern that works is four moves in one or two sentences: score, prior score, interpretation, action. Here is what that looks like at three points in treatment. The patients are fictional.

Intake baseline

PHQ-9 today: 17, moderately severe range. GAD-7: 11, moderate. Scores establish a measurement-based baseline prior to initiating sertraline 50 mg daily. Will readminister at each visit during titration.

Why this works: a baseline is only a baseline if the note says so. This snippet names the severity band and ties the number to the treatment being started, so every later score has something concrete to be measured against.

Follow-up with improvement

PHQ-9 today: 9, down from 17 at intake eight weeks ago. The 8-point drop exceeds the commonly cited 5-point margin for meaningful change and is approaching the 50 percent reduction convention for response. Continue sertraline 100 mg at current dose; recheck PHQ-9 in 4 weeks.

Why this works: score, prior score, interpretation, action, in one breath. A reviewer can see both the trajectory and the reason the regimen was left alone.

Worsening that drives a plan change

GAD-7 today: 15, up from 8 six weeks ago, past the moderate threshold of 10 and now in the severe band. Worsening is consistent with the escalating work stressors reported above. Increasing buspirone from 15 mg to 30 mg daily given clear symptomatic worsening on the current dose; recheck GAD-7 in 2 to 4 weeks.

Why this works: the score does not just sit next to the dose change, it justifies it. The interpretation names the band crossing, agrees with the narrative, and the plan responds to both.

Where scales fit in the note and the billing picture

Scores belong in the note body, next to your interpretation, with the trajectory carried into the plan. If you want to see where they sit inside a full note, the psychiatric progress note example walks through the whole structure, and the 90791 vs 90792 guide covers baseline scores at the intake visit, where they matter most.

On coding: many coding educators advise against counting standardized questionnaires as data-element tests when selecting an E/M level, so the safer posture is to let problems addressed and risk carry the level and treat the scales as clinical monitoring. The 99214 with 90833 billing guide covers how the level actually gets selected.

Five scale documentation failures

Frequently asked questions

How often should I administer them?

Common patterns: every visit during active treatment or dose changes, then spaced out once the patient is in maintenance. Payer quality programs have their own expectations for measurement frequency, so check the ones you participate in.

Can patients complete them before the visit?

Yes. Both are self-administered by design, so the portal or the waiting room works fine. The step that cannot be delegated is reviewing the result with the patient and interpreting it in the note.

What if the score contradicts the presentation?

Document the discrepancy and your clinical reasoning. "PHQ-9 of 4 despite reported low mood; patient acknowledges minimizing on forms" is the kind of documentation that holds up well under review, because it shows you looked at the number instead of filing it.

Do I need the patient's exact score in the note?

Yes. Chart the number and the prior number. "PHQ-9 administered" with no result is not measurement; it is a claim that measurement happened.

Are there versions for adolescents?

The PHQ-A exists for adolescents, and validated tools vary by age and setting. Choose the instrument validated for the population in front of you rather than defaulting to the adult forms.

The score-prior-interpretation-action pattern above is what OneStep Scribe drafts for you.

OneStep Scribe includes standardized scales built in, with scores flowing straight into the drafted note alongside their trajectory, for psychiatrists, PMHNPs, therapists, psychologists, and other mental health providers. Every account is NPI-verified, and every note is yours to review and sign.

Start a 14-day free trial

This article is educational only and is not legal, billing, or clinical advice for any specific patient, and it does not guarantee reimbursement. The PHQ-9 and GAD-7 were developed by Drs. Kroenke, Spitzer, Williams and colleagues and are free for clinical use. CPT is a registered trademark of the American Medical Association. Severity bands and conventions are summarized from the published literature; verify against current guidance and your payers. The patients described are fictional.