Dr Neil Paul

Marking my own homework: a full review of the Usual GP Continuity Analyser

Drafted with Claude — edits and comments welcome to .

In July I released a free Usual GP Continuity Analyser. It checks a practice’s Usual GP field against who patients actually see. Practices are now using it, so this month I asked the question you should ask of any tool that touches patient data: is it actually right?

I set four AI reviewers loose on it in parallel, each with a different brief. One checked the maths. One checked how files are loaded. One checked security and privacy. The last checked whether the method and wording stand up clinically. Then I checked what they found against the code and the real data.

The good news. The core engine was sound. Both of my benchmark datasets reproduced exactly, and nothing leaked data or let a malicious file run code. The offline file really did make no network requests.

The bad news. There were real errors, and some were the quiet kind that give plausible-looking wrong answers:

  • Excel files had their dates scrambled. Opening the export in Excel and saving it as .xlsx turned 4 March into 3 April and dropped any date after the 12th. The results still looked believable.
  • One typo could empty the analysis. A single date like “01/01/49” was read as 2049. That moved the 12-month window into the future, and almost every consultation fell outside it.
  • It could suggest recoding patients to the wrong people. About one in five suggestions pointed at someone who holds no list, often a registrar. Anyone patients mostly saw could be suggested, even a GP who had left.
  • Some wording claimed more than the numbers supported. “The top of this list is the safest” wasn’t true: the ranking reflects volume, not certainty. Two different “correct” percentages sat on the same page with no explanation of the difference.
  • The independent check wasn’t independent any more. My Python copy of the algorithm had drifted from the real tool, and it couldn’t read the current EMIS export at all.

What v1.5.0 fixes. There is a new clinician setting, Active, never suggest, for registrars, locums and nurses. Their consultations still count, but patients are never recoded to them. Ties between two equally-seen GPs are flagged for the practice to decide. Excel dates are read correctly, and future-dated rows or typos are skipped with a warning. Labels now say exactly what each figure is a percentage of, and you can sort suggestions by certainty. The page now tells the browser to block every network connection, so “no data leaves this computer” is enforced, not just promised. The spreadsheet library has been upgraded to fix two known security flaws. The Python check has been rebuilt, and both benchmarks still match exactly.

The lesson. None of these bugs would have shown up in a demo. They show up with real exports, real staffing and real edge cases, and AI-built tools are no exception. What worked was treating review as part of the build: several independent reviewers, every claim checked against the code, and fixed benchmark numbers that every release must hit. It took an afternoon.

There’s still more to do: a printable PCN summary, filtering by appointment type for triage-heavy practices, and a standard continuity index so practices can compare themselves with published figures. If you use the tool and something looks wrong, please tell me. That’s how the last lot were found.