Inside Our Methodology: How We Review AI

A closer look at the public administration approach behind every HCI Ethical Review — and how a review actually unfolds, step by step.

Our Our Methodology page explains the core idea: for AI operating inside or alongside public institutions, the right standard isn’t a new AI ethics checklist — it’s the body of thought public administration has used for over a century to ask whether a rule-bound, discretion-exercising institution has earned the trust of the people subject to its decisions. This page goes one level deeper: what each of the four ideas actually asks, and what happens, step by step, when we apply them to a real system.

The complete scoring instrument we use internally is not published here — it’s the working tool our reviewers use case by case, and keeping it out of general circulation is part of what keeps a review meaningful rather than something a system can be built to pass. What follows is a genuine, substantive look at how the methodology works, illustrated with real (if partial) examples.

The four ideas, in practice

Legitimate authority. Before we ask whether a system works well, we ask whether it was ever properly authorized to make the decision it’s making. That means: a real, documented decision by a body with the standing to make it — not a procurement choice that quietly became policy. A defined scope that current use hasn’t drifted beyond. Rules that are written down and inspectable, not just “whatever the model learned to do.” A system can be accurate and still fail this test, if the institution running it never actually granted it the authority it’s now exercising.

Real accountability. A log file is not accountability. Accountability means an external forum — a court, an appeals board, an oversight committee, an affected community — can ask a specific question about a specific decision and get an answer that can actually change something. We check five kinds of forum: political, legal, administrative, professional, and social. Most institutions we review have built a strong administrative trail and very little else. The gap that matters most is usually legal: can the person affected by a decision actually contest it, in a forum with the power to overturn it — or does the appeals process quietly route back to the same unaccountable process that produced the original result?

Bounded discretion. Every AI system that shapes an individual outcome is, in effect, exercising discretion that used to belong to a person — a caseworker, an officer, a reviewer. The difference is that the system’s “judgment” is frozen at training time rather than exercised case by case. We ask two things: where, precisely, does the system’s output determine an outcome rather than merely suggest one? And at each of those points, can a human genuinely intervene — with the time, authority, and institutional backing to actually do so — or is “human oversight” a formality that, in practice, rubber-stamps whatever the system decided?

Public value. Efficiency for an institution and value for the people it serves are not the same thing, even though they’re often treated as interchangeable. A system can cut processing time and staff hours — real gains — without anyone checking whether those gains are shared evenly, or concentrated among the easiest cases while the hardest ones wait longer than ever. We ask whose value is actually being created, and whether that value is something the people affected would recognize and endorse, not just something that looks good in an internal report.

How a review actually unfolds

  1. Scoping. We confirm the system materially affects access to a government service, benefit, legal process, or an equivalent quasi-public function before applying the full methodology — not every AI deployment needs this depth of review.
  2. Evidence gathering. We review governance and delegation records, model documentation, and audit history, and we talk to both the technical team and the frontline staff who work with the system’s outputs every day — and, wherever possible, to people affected by its decisions.
  3. Scoring. Each of the four areas is assessed against a defined set of criteria, each scored on a five-point scale — from no evidence at all to fully documented, evidenced, and independently checked. We score every criterion against a specific document, log, or corroborated account, not general impressions.
  4. Pattern-checking. Because these four ideas are connected, not independent, we specifically check for the failure patterns that tend to travel together — for instance, a system whose authority was never properly delegated is almost always also a system with no external forum able to question it.
  5. Reporting. Every review produces a report structured around the four lenses — never collapsed into a single score, because doing so would hide exactly the pattern the methodology is built to surface — a clear list of critical findings, a prioritized remediation roadmap, and a plain-language public value statement the client is welcome to publish.

What one criterion looks like

To make this concrete rather than abstract, here’s a real (partial) example from the accountability assessment — one of roughly twenty criteria across the four areas.

Legal forum access. Can someone affected by a specific decision access a legal or quasi-legal forum — a tribunal, an appeals board, an ombudsman with binding authority — empowered to review and, where warranted, overturn that decision?

A system scores lowest when no such forum exists at all, and highest when an accessible forum exists, has clear binding authority, and has a documented record of actually being used.

That’s one criterion, in one of four areas. The full instrument — every criterion, every scoring anchor, and how the four areas combine into a final assessment — is what our reviewers use in every engagement, and it’s available in full to review clients and qualifying partners as part of that engagement.

Why we work this way

None of this replaces the technical and legal review every credible AI assessment already includes — bias testing, data protection compliance, and documented limitations remain part of every review we do. What it adds is the institutional question technical checks can’t answer on their own: not just does this system work, but does the institution running it have the right to make this decision, answer for it, and account for who it actually serves. For the government agencies, courts, public benefits systems, and humanitarian organizations we most often work with, that’s not a new bar. It’s the one their institutions have always been expected to meet — we’ve simply built the instrument that lets AI be held to it too.

Start an Ethical Review →

For the full academic treatment of the theoretical framework — including its relationship to the EU AI Act, OECD, and UNESCO frameworks — see our published paper, “Borrowing Legitimacy: A Public-Administration Framework for the Ethical Review of Artificial Intelligence in Public-Facing Institutions.”