Responsible Innovation

Responsible AI: A Tool, Not a Replacement for Human Judgment

Artificial intelligence can compress the distance between a question and a usable first draft. It cannot take responsibility for whether that draft is accurate, lawful, original, or appropriate. Knowing where that line sits is the entire discipline.

What these tools are genuinely good at

It is worth being specific, because vague enthusiasm and vague suspicion are equally useless. In our experience the strongest uses are the ones where a person remains the reader and the decider.

Research support, where a model helps map a subject, surface the questions worth asking, and identify terminology that makes further searching more effective. Drafting, where getting from a blank page to a rough structure is the expensive part and a first draft is easier to improve than to originate. Summarizing and comparing, where the volume of material is the obstacle. Analysis, where laying out options and their trade-offs helps a person think. Prototyping, where a working sketch answers a design question faster than a discussion. And documentation, which almost everyone underinvests in and which these tools make markedly less tedious.

The common thread is that each of these produces something a person then evaluates. The tool moves work to the point where judgment can be applied. It does not apply the judgment.

Why human review is not a formality

The reason review matters so much is that the failure mode is invisible. A model does not produce output that looks uncertain when it is uncertain. Fluent prose reads as authoritative whether or not the content behind it is sound, which means the usual signals a reader uses to detect weak work are absent.

This inverts a normal editorial assumption. Reviewing human work, you can often trust that confident writing reflects confident knowledge. Reviewing generated work, confidence tells you nothing, and every checkable claim has to be checked. Review that consists of reading for tone and coherence will pass material that is wrong, because tone and coherence are precisely what the tool is best at.

Accuracy and fabrication

Language models produce plausible text, and plausibility and truth overlap without being the same thing. The result is confident invention: details, figures, citations, quotations, and sources that do not exist but are formatted exactly as real ones would be.

Our practice is that anything checkable gets checked against a primary source before publication. Names, numbers, dates, quotations, technical specifics, and legal or policy claims are all verified independently, not confirmed by asking the model again. If a claim cannot be verified, it comes out. The paragraph is usually better for the removal, because an unverifiable specific was doing less work than it appeared to.

Privacy and confidential information

What goes into a tool is a decision with consequences, and it is easy to make carelessly because pasting is frictionless. Before information is entered, the question is whether the tool is appropriate for that category of data and whether we have the right to share it at all.

Personal information belonging to other people is the clearest case: it is not ours to circulate for convenience. Material covered by an obligation of confidence is similar. Where AI assistance is genuinely useful for sensitive work, the answer is usually to describe the problem in general terms rather than to paste the specifics, which gives up surprisingly little.

Copyright and originality

A model trained on existing text can reproduce it, sometimes closely, sometimes in a way that is hard to detect without checking. Output that arrives without attribution is not therefore original, and treating it as original is how a company ends up publishing someone else's work under its own name.

So we treat generated text as a draft to be rewritten rather than copy to be published, check distinctive phrasing when a passage seems unusually polished, and credit sources where the substance came from somewhere identifiable. The obligation is unchanged from before these tools existed. Publishing is a claim of authorship, and the claim should be true.

Bias and missing context

These systems reflect patterns in the material they learned from, including patterns nobody would defend on purpose. They also have no access to the context that makes an answer suitable: who the audience is, what has already been promised, which constraints apply, and what would be inappropriate in this particular situation.

The practical consequence is that generic-sounding advice deserves suspicion rather than trust. When output reads as though it would be equally applicable to any company, it usually is, which means it is not applicable to this one. Supplying the missing context, and checking whose perspective the answer has quietly assumed, is work only a person can do.

Keeping a record of how work was produced

We keep our own notes on where AI assistance was used and what verification was performed. This is not ceremony. When a question arises months later about why something says what it says, the record is what makes the answer recoverable, and it also makes patterns visible: which uses consistently save time, and which consistently produce work that needs rewriting anyway.

The record is internal, and it does not replace judgment about disclosure. Where a reader's understanding of a piece of work would reasonably depend on knowing how it was produced, that is a disclosure question, and the answer should be decided honestly rather than by default.

Accountability does not transfer

All of this reduces to one principle. A tool cannot be accountable. It has no stake in the outcome, cannot be asked to explain itself in any binding sense, and cannot make a commitment. Accountability therefore stays where it already was: with the person who published the work, approved the decision, or made the promise.

The standard we hold ourselves to is that "the model produced it" is never an explanation for an error. If something inaccurate is published, the failure is in our review. If confidential information is mishandled, the failure is in our decision to enter it. Used this way, these tools make a small company considerably more capable. Used as a substitute for judgment, they mostly make it faster at being wrong.

More from OPM

Insights collects our writing on building and operating digital work. The About page explains how these principles shape the company's decisions.

All articles About OPM