See exactly what we produce, before you decide anything.
This is the Specimen: a worked example of what a Genesis engagement produces, written out in full so that you can judge the quality of the thinking for yourself. It is one workflow in one invented firm, taken from the problem the audit found to the decision about where a tool belongs and where it does not. You can read all of it here, and it costs you nothing to look.
You are looking at the finished thinking, not a brochure about it. Read it and judge for yourself.
The same shape, whatever you do.
When we decide where AI belongs in a business, we write the decision down in a fixed form. It covers one job, from the problem it starts with to the check that tells you whether it is still earning its place, and it shows you where we stopped the tool and put a person back in. The form is the same for a plumbing firm and a law firm, so you can read how a decision is made whatever your sector.
Every record answers the same questions: the job and how it is done today; the source behind the decision and the date it was checked; who owns it; where the tool must stop; the test it had to pass; the human check that stays in place; and the honest result, including when that result is not to use AI at all.
Now the same record, filled in.
What you have just read is the method. What follows is one record filled in, for an invented heating and plumbing firm, so you can see the questions carrying real weight. It is a trades and field-service example. When we work with you, we produce one for your own business.
The invented firm. Kestrel Building Services, an eleven-person heating, plumbing and mechanical maintenance firm. The job recorded is checking supplier invoices against what was ordered and what was delivered, before the weekly payment run.
Where a tool belongs, and where it must stop.
The operating record is how we write a decision down. It covers one job in fourteen fields: the problem and the baseline, the ruling on whether to use AI at all, the boundary the tool does not cross, the data allowed nowhere near it, the named owner, the sources and their dates, the tests and what they returned (including the one the tool was caught failing), the runbook and the fallback, the measure that tells you it is still earning its place, and the next decision already queued. It is the working record, with the test results in it as they came back, not a summary written afterwards to look tidy. The whole of it is below; nothing is held back on another page.
Genesis AI Operating Record — Proof Example (ILLUSTRATIVE)
What you are looking at
When we decide where AI belongs in a business, we write the decision down in a fixed form we call an operating record. It covers one job, from the problem it starts with to the check that tells you whether it is still earning its place, and it shows you where we stopped the tool and put a person back in. This page is an illustrative example of that method, not a live client record. It shows you the form and one worked example of it, so you can judge how a decision would be made, not evidence that the method is already running inside a business.
This page holds two things. Part One is the record itself, empty: the fourteen questions this method asks of a job, in the same order every time. It is the same for a plumbing firm and a law firm, so you can read how a decision gets made whatever you do. Part Two is one record filled in, for an invented business, so you can see the questions carrying real weight.
Two things this record is not. It is not a claim about what AI will do for you, and it carries no promised result. And it is not a summary written afterwards to look tidy. It is the working record, with the test results in it as they came back.
PART ONE — THE RECORD, EMPTY (the same for every sector)
An operating record covers one job. It has fourteen fields in five groups. We do not treat a job as live until every field is filled and the acceptance tests have a written result.
Group A — The decision
- The problem, and how the job is done today. The specific problem in one or two sentences, and how the work gets done now without AI, including the time, the cost or the error rate as far as anybody actually knows it. This is your baseline, which is the measurement we take before a change so that later we can say whether the change worked with numbers rather than impressions.
- The decision: use AI, wait, or do not use AI. A ruling for this job, with the reason. "Do not use AI here" and "wait until X is true" are proper answers, written down with the same weight as "use".
Group B — The job and its edges
- The job we place a tool inside, and the boundary. The exact task, named plainly, and the line the tool does not cross. The boundary says where the tool's part ends and where yours begins.
- Permitted and prohibited data. What may go into the tool and what must never go into it, written as two lists rather than as a principle.
- The owner and the access list. The named person accountable for this job, and the named people allowed to run it. Access is a list of names, never "the team".
Group C — The evidence behind the choice
- The sources, and the date each was checked. Every source the decision rests on, with the date we last read it. A source with no date is treated as unverified.
- The tool chosen, and the ones turned down. What we selected, and what we considered and rejected, each with the reason. What was rejected is part of the proof.
Group D — The controls
- The human decision and check points. Every point where one of your people must decide, review or sign off before the work goes on or leaves the business.
- The acceptance tests, and what they returned. The tests the job had to pass before you relied on it, with the result of each, including any test it failed and what changed because of it.
- The runbook and the fallback. The numbered steps to run the job, and what you do when the tool is unavailable or its output is wrong. The old way of working has to stay possible.
Group E — The outcome and the upkeep
- The outcome measure. The specific thing you measure to judge whether this earns its place, and how often somebody reads it. This records what is measured. It is not a promise about what the measure will say.
- Last checked, and stale after. The date the record was last tested against reality, and the date or the trigger after which it must be checked again before anybody relies on it.
- The change log. A dated line for every material change to the job or to this record, so its history is visible.
- The next decision. The one question queued for the next review, and the condition that would open it. This is what keeps the record a live account rather than a certificate.
PART TWO — ONE RECORD FILLED IN (ILLUSTRATIVE)
The business below is invented, and so is everybody in it. There is no real client here, no measured result and no promised one. Every figure belongs to this example and to the assumptions it states about itself.
Four of the five sources in field 6 are real. We read them first-party and re-checked all four on 24 July 2026. They are named with their dates so that you can go and read them yourself. The fifth is illustrative, and it says so.
The business (invented). Kestrel Building Services, an eleven-person heating, plumbing and mechanical maintenance firm. Seven engineers, an office manager, a part-time administrator, a contracts supervisor and a working director. Domestic call-outs and repairs, and planned maintenance contracts on about forty small commercial sites. One trade account with a builders' merchant, four regular suppliers, and two subcontractors on the books.
The job recorded here: checking supplier invoices against what was ordered and what was delivered, before the weekly payment run.
1. The problem, and how the job is done today
Materials and parts arrive at Kestrel by three routes, and all three produce paper. An engineer walks into the merchant and leaves with a counter ticket. A delivery reaches the site and somebody signs a delivery note, sometimes an apprentice, sometimes the customer. An order placed by phone arrives with a note in the van. At the end of the month the merchant statement lists every one of those lines across every job.
The office manager checks supplier invoices against the order and the delivery note before the Friday payment run. In practice, and this is the firm's own description of its own week, that check happens properly on the large invoices and not on the rest. There is frequently no written order to check against, because the order was a phone call. So the classic three-way match of invoice to order to delivery is, at Kestrel, a one-way match: the invoice is checked against what somebody remembers agreeing.
The baseline, as an illustrative assumption. The firm handles roughly 260 supplier invoices a month, and takes the view that fewer than a third get checked against a document rather than against memory. That is Kestrel's own working figure for this example. It is not a measured study and it is not a statement about anybody else's business.
What that costs, in the firm's own words. A duplicated invoice gets paid twice. A price that does not match the rate agreed with the merchant gets paid at the higher figure. A quantity that was billed and never delivered gets paid in full. None of those announce themselves, and the firm's honest position is that it does not know how often they happen, which is the reason it wanted this looked at.
2. The decision: use AI, wait, or do not use AI
USE AI to point at disagreements between a supplier invoice, the order behind it and the delivery note, across every invoice rather than the large ones only.
DO NOT USE AI on subcontractor invoices under the Construction Industry Scheme. This is a separate ruling and it is not a caution. The split between labour and materials on a subcontractor's invoice sets the tax the firm deducts, and HMRC puts that check on a person: "The contractor must always check and keep records to make sure that the part of the payment for materials supplied is not overstated" (CIS 340, paragraph 3.13). If the materials part is overstated, too little tax is deducted, and paragraph 3.13 places that check, and so the consequence of getting it wrong, on the contractor as a person, which here is Kestrel. That is the exact duty paragraph 3.13 establishes, and it is why these invoices are outside the tool entirely, with the reason written here so that nobody quietly widens the scope in six months.
DO NOT USE AI to decide whether to withhold payment or to open a dispute with a supplier. That is a judgement about a commercial relationship, and about a credit line the firm needs on site next week. Nothing in the invoice tells you any of that.
WAIT on the merchant statement reconciliation. See field 14.
The reason behind the first ruling, and it matters that it is the real one. The reason is not that a general AI tool sometimes gets a number wrong. It is structural, and the UK's audit regulator states it plainly: "LLMs are often not natively strong at computational tasks, as they lack semantic awareness" (Financial Reporting Council, March 2026). So we do not place a tool where the job is arithmetic. Reading a lot of documents and pointing at where two of them disagree, with its working shown, is a different job, and it is the one we placed a tool inside.
3. The job we place a tool inside, and the boundary
The job. For each supplier invoice, the tool reads the invoice alongside the order and the delivery note for the same supply, and reports where they do not agree: a quantity billed above the quantity delivered, a unit price above the rate on the order, a line that appears on the invoice and on no delivery note, or the same invoice number seen twice. Each report cites the document, the page and the line it came from. The office manager then opens each one.
The boundary, in five parts, and every one of them is testable.
- The tool never does the sum. Every total, every extension of quantity by price, every variance figure is calculated by the accounting system or by a spreadsheet with the arithmetic in it. The model is not the calculator. The regulator's own mitigation is exactly this: giving a language model a computational tool to call "significantly mitigates risk compared to the LLM attempting the computation itself" (FRC, March 2026).
- Every flag cites its source. The document, the page and the line, so that the person reviewing it opens the page and looks. A flag that cites nothing is treated as a failure of the tool, not as a judgement call for the reader. The same regulator's reasoning: requiring a citation for everything means the model "is less likely to fabricate information", and it "supports effective human review of the output" (FRC, March 2026).
- The tool flags. It never concludes. It reports that two documents disagree and shows you both. It does not say the invoice is wrong, it does not say the supplier has overcharged, and it does not rank anybody's honesty.
- The false-alarm rate is measured, and the measurement is shown. How often the tool flags an invoice that turns out to be perfectly correct is recorded in field 9 alongside how often it catches something. A tool that flags too much is a tool people stop reading, and a control nobody reads is not a control.
- The tool never files, submits, pays or sends. It does not approve an invoice, it does not release a payment, it does not email a supplier and it does not touch the banking. Every one of those is a person's act, recorded against a person's name.
Where the tool stops and Kestrel starts. The tool hands over a list of disagreements with citations. Everything after that point is a person: opening the cited line, deciding whether it is a real discrepancy, deciding what to do about it, and paying.
4. Permitted and prohibited data
Permitted into the tool: the supplier invoice, the purchase order or written order confirmation, and the delivery note or goods received note for the same supply. The supplier's name and account number. The job reference.
Prohibited, and never entered: bank details, sort codes and account numbers for any supplier; the firm's own banking credentials; any staff personal data, including timesheets, payroll and anything about an individual's health or absence; anything relating to a supplier in an active legal dispute; and subcontractor invoices under the Construction Industry Scheme, which are out of scope under field 2 and are therefore out of scope for the data too.
5. The owner and the access list
Owner: the office manager, who runs the payment run and who is accountable for this job and for keeping this record current.
Access: the office manager and the working director, who authorises payments. Two people. The administrator and the contracts supervisor do not run it, because the person who queries a supplier and the person who releases the money have to be people the supplier can be pointed at by name.
6. The sources, and the date each was checked
Four of these five are real, public and dated, and anybody can read them.
- Financial Reporting Council, Generative and Agentic AI Guidance: risks, mitigations and illustrative examples, published 30 March 2026. Read in full. Checked 24 July 2026. This is what the boundary at field 3 is built on, and specifically the finding that language models are not natively strong at computation, and the mitigations of giving the model a calculator to call and requiring it to cite a source for everything. This is guidance for audit firms, and it does not bind Kestrel, which is not audited. We are not applying a rule that governs Kestrel; we are choosing to adopt these controls by analogy, because the limitation the guidance describes belongs to the technology rather than to auditing, and this is where we found it stated plainly by a UK regulator.
- The PCRT bodies, Topical guidance covering the application of PCRT to the ethical use of artificial intelligence tools, issued 19 January 2026. Read in full. Checked 24 July 2026. PCRT binds members and regulated firms doing UK tax work; it does not bind Kestrel's general invoice checking. Genesis has chosen to adopt two of its controls here by analogy, because they are the right way to review any AI output. First: "The output from an AI tool should also be regarded as if it were prepared by a less experienced junior colleague and reviewed with appropriate scepticism." Second: "members remain … accountable for any work produced, regardless of whether AI has been involved in producing the work or refining work already produced." The same guidance defines automation bias as "a tendency to favour output generated from automated systems, even when human reasoning or contradictory information raises questions as to whether such output is reliable or fit for purpose", which is the failure the checks at field 8 exist to resist.
- HMRC, Construction Industry Scheme: a guide for contractors and subcontractors (CIS 340), published 31 March 2014, last updated 8 July 2026. Checked 24 July 2026. Paragraph 3.13 is the basis of the second ruling at field 2, because it puts the materials check on the contractor as a person and keeps it there.
- Harber v The Commissioners for His Majesty's Revenue and Customs [2023] UKFTT 1007 (TC), decided 4 December 2023. Checked 24 July 2026. An appellant put nine First-tier Tribunal decisions before the Tribunal in support of an appeal against a penalty. None of the nine existed. They had been produced by an AI tool, and the Tribunal recorded that such tools can return results that are plausible and wrong. It is a real British case, in a real proceeding, and it is why field 3 part 2 is written as an absolute rather than as good practice. The PCRT guidance above cites the same case for the same reason.
- The tool provider's data-processing terms, and its setting for excluding business inputs from training. Checked 14 July 2026 (illustrative date, belonging to this example, and the only source in this list that is not real).
Any of these changing is a trigger to re-check the record. See field 12.
One thing we will not put in this list. No vendor's marketing claim about its own AI appears anywhere in this record, and none was used to reach any decision in it.
7. The tool chosen, and the ones turned down
Chosen: a business-tier assistant with a data-processing agreement and a setting that excludes the firm's inputs from training, running against documents that have been indexed first so that every statement it makes can cite a line.
Rejected, with the reason in each case:
- A free consumer tier. No data-processing agreement, and no clear exclusion of inputs from training. Kestrel is putting its suppliers' commercial terms into this.
- Automatic approval of any invoice the tool matched cleanly. Rejected because it crosses the boundary at field 3 part 5, and for a second reason worth stating: the failure we expect from a checking tool is that it flags too much, and automatic approval would turn that into flagging too little, silently and in the direction of paying money out.
- Connecting the tool directly to the accounting system with write access. Rejected for now. It is not needed to answer the question the firm asked, and it would put the tool inside the system of record, which the boundary keeps it out of.
- Doing nothing. Kept as the fallback and the baseline, and it remains a legitimate answer. It was not chosen because the firm's own position at field 1 is that it cannot say how much it is paying in error.
8. The human decision and check points
- Every flag is opened by a person against the cited line before anything is paid. Not a sample of them. Every one.
- Nothing is paid because the tool was quiet, and nothing is held back because the tool spoke. A flag is a reason to look. It is not a decision, and the record of the decision names the person who made it.
- The office manager decides; the director authorises the payment run. Two people, and the second one sees the flags that were dismissed as well as the ones that were acted on.
- A flag that cites nothing, or cites a line that does not carry what the flag says, is logged as a tool failure and counted in the monthly measure at field 11. It is never treated as noise for the reader to filter out.
9. The acceptance tests, and what they returned
Before the firm relied on this, we ran it over 200 supplier invoices that had already been settled. Three tests. Every flag the tool raised across all three was opened and checked by hand, which is how the false alarms below were counted. The figures are illustrative and belong to this example.
Test 1: does it catch errors that were known to be there? 6 errors in that set had come to light afterwards by other means, through a credit note or a supplier query. The tool flagged 5 of them. It did not flag the sixth, which was a price above the rate agreed with the supplier, because that rate existed only in an email between the director and the merchant's account manager, and no document in the tool's scope carried it. Result: 5 of 6, and the miss is a records problem the firm owns rather than a tool problem. The agreed rate is now written on the order.
Test 2: how often does it flag an invoice that turns out to be perfectly correct? This is the test the firm cared most about, and it is the one this tool did worst on. Across the same 200 invoices it raised 31 flags. 24 of them were false alarms. 19 of those 24 came from a single supplier whose delivery notes count in boxes of ten and whose invoices count in units, so a delivered quantity of 12 was compared against an invoiced quantity of 120 and reported as a disagreement every time. The unit conversion now sits in a fixed reference table that the arithmetic runs against, rather than being something the tool is expected to work out. On a re-run the same 200 invoices produced 12 flags, of which 5 were still false alarms. That is more than one flag in three, it is the level the firm now works with, and it is written here rather than smoothed, because the person opening those flags needs to know what to expect.
Test 3: can a person open every citation and find what the flag says is there? 31 flags, 31 citations. 3 of them pointed at the wrong line. The document was right and the line was not, and the person who opened it found it did not say what the flag claimed. Nothing was paid wrongly and nothing left the business, because the check at field 8 is to open the cited line rather than to read the flag and agree with it. That check is the only reason this was caught, and it is the reason the check is mandatory rather than recommended.
What changed because of Test 3. Opening the citation stopped being described as good practice and became the step the job cannot skip. The rule is now written into the runbook at field 10 as step 4, a wrongly cited flag is counted every month at field 11, and the record was changed on the day. It is in the change log at field 13 with that date against it.
Read those three results together, because that is the honest picture. The tool caught 5 of the 6 errors that were there to catch. It also raised 24 alarms about invoices that were perfectly correct, and it pointed at the wrong line 3 times. It is useful because a person opens every flag it raises, and it would be worse than useless if they stopped.
10. The runbook and the fallback
Runbook.
- Pull the week's supplier invoices, with the order and the delivery note for each.
- Run the tool over them.
- Read the list of flags. Do not act on any of them yet.
- Open the cited line for every flag. If the citation does not open, or the line does not carry what the flag says, log it as a tool failure and set the flag aside.
- Decide each surviving flag: real discrepancy, or not. Record the decision and who made it.
- Query the real ones with the supplier, by name, from a person.
- Release the payment run. The director authorises.
- Record the month's counts for field 11.
Fallback. If the tool is unavailable, or its output looks wrong, the payment run goes ahead on the previous manual spot-check exactly as before. Losing the tool costs coverage across the small invoices. It never stops Kestrel paying its suppliers, and it never puts the firm in the position of not being able to pay a merchant it needs on Monday morning.
11. The outcome measure
Four numbers, read monthly by the office manager and the director.
- How many flags were raised.
- What proportion of them turned out to be real. This is the honest one, and it is the one that tells you whether the tool is still worth opening.
- How many flags cited nothing, or cited the wrong line.
- Whether anything wrong was paid, from any route.
These are what Kestrel measures. They are not a claim about what those numbers will say here, and they are certainly not a claim about what they would say in your business.
12. Last checked, and stale after
- Last checked: 23 July 2026 (illustrative).
- Stale after: 23 October 2026, or immediately on any of these, whichever comes first: a change to the tool provider's data-processing terms; a change to any of the UK guidance or authority named in field 6; a change of supplier or of the document set the check runs against; or a month in which fewer than half the flags turn out to be real. On the last test that ran, 7 flags in 12 were real, so the firm is working close to its own trigger and knows it. After the stale date, nobody relies on this record until it has been checked again.
13. The change log
- 14 July 2026 (illustrative). Job first recorded. Tool selected. Boundary set at the five parts in field 3. Subcontractor invoices under the Construction Industry Scheme ruled out of scope.
- 16 July 2026 (illustrative). Test 2 returned 24 false alarms in 31 flags. Unit conversion moved into a fixed reference table. Re-run and recorded at 12 flags and 5 false alarms.
- 17 July 2026 (illustrative). Test 3 returned 3 flags citing the wrong line. Opening the cited line made a mandatory step of the runbook, and wrongly cited flags added to the monthly measure.
- 20 July 2026 (illustrative). Agreed supplier rates written onto orders, after Test 1 showed a real error the tool could not see because no document in scope carried the rate.
- 23 July 2026 (illustrative). Record reviewed. No change to the boundary. Next decision queued at field 14.
14. The next decision
The queued question: should the same checking run against the monthly merchant statement, which lists every counter ticket and delivery note across every job?
The condition that would open it: not yet, and the reason is not about the tool. Most merchant collections at Kestrel are made without a written order, so there is nothing on the firm's side to compare the statement against. Until counter collections are raised against an order, a tool pointed at the statement would be comparing the merchant's document with the merchant's other document. The question is recorded here so that it is not lost, and the answer today is wait.
Why this record looks the way it does
Three notes, because the shape of it is deliberate.
The tool is not the actor anywhere in this document. We placed a tool that points at disagreements between documents. Kestrel's office manager reviews. Kestrel's director authorises. When the tool got it wrong, a person caught it, because the job was built so that a person would.
The failure in field 9 is in the record on purpose, and it is one failure, not a catalogue. It sits where it belongs in the sequence, after the method is on the page, because that is where it means what it actually means: the control worked. We have not put it there because being open about AI earns anybody's trust, because as far as we can tell it does not. We have put it there because a record with nothing in it but passes is an advertisement, and because a control you cannot inspect is only a claim.
The honest "do not use AI here" at field 2 is a ruling, not a hedge. An entire class of document, the subcontractor invoices, is outside the tool, with the reason written down. If we could not tell you where the tool has no business being, we would not be much use to you in deciding where it does.
Illustrative, plainly
Kestrel Building Services does not exist. Its people, its suppliers, its invoice counts and every number in Part Two were invented for this example and are internally consistent with the assumptions the example states about itself. They are not market claims, they are not a promised result, and they must not be read as either.
Four of the five sources named in field 6 are real, are public, and were re-checked on 24 July 2026. The fifth is the tool provider's terms, which belongs to the example and is labelled that way.
What you see first here is the record itself, which is the same in any sector. The version filled in for a business like yours is one step in. When we work with you, we build one for your own.
The diagnosis that came before the decision.
Before any decision, there is a diagnosis. The findings-report extract shows what the audit found in the same firm's buying and paying: invoices approved against memory because no written order exists, agreed rates sitting in an email where the person checking cannot see them, and a class of subcontractor invoice that carries a tax judgement the law puts on a person. Each finding says what is happening, the root cause beneath it, what it costs, how sure we are, and the evidence it rests on. It reports what is true before it proposes anything. The whole extract is below.
Genesis Findings Report — Excerpt (ILLUSTRATIVE)
What you are looking at
This is a short excerpt from a Genesis Findings Report. A Findings Report is the document you receive after the audit and before anything is recommended or built. It records what is happening in each part of your business, why it is happening, and what it costs you. It does not yet tell you what to do. That comes later, and only once you agree the diagnosis.
This excerpt is the second half of the Specimen. The operating record you can also read shows one place where a tool was later put to work, and the boundary and the checks that hold it. This half shows the step before that: what the audit found in the first place. They are two views of the same illustrative engagement, for the same invented firm.
The firm (invented). Kestrel Building Services, an eleven-person heating, plumbing and mechanical maintenance firm. The same firm as the operating record. There is no real client here, no measured result and no promised one. Every figure belongs to this example and to the assumptions it states about itself.
The part of the business excerpted: buying materials and paying suppliers. In the full report this is one workflow section among several. One section is shown here so that you can read the shape of a finding without reading the whole audit.
The workflow: buying materials and paying suppliers
Reviewed with (illustrative): the office manager, who runs the weekly payment run, and the working director, who authorises it.
Current state
Materials reach Kestrel by three routes, and all three produce paper. An engineer walks into the builders' merchant and leaves with a counter ticket. A delivery reaches the site and somebody signs a delivery note, sometimes an apprentice, sometimes the customer. An order placed by phone arrives with a note in the van. At the end of the month the merchant statement lists every one of those lines across every job.
Before the Friday payment run, the office manager checks supplier invoices against the order and the delivery note. In practice that check is done properly on the large invoices and not on the rest, and there is often no written order to check against, because the order was a phone call. The firm's own account of its week is that the three-way match of invoice to order to delivery is, for most invoices, a one-way match: the invoice is checked against what somebody remembers agreeing.
Baseline (the firm's own working figures, illustrative).
| What we measured | The starting point | How it was captured |
|---|---|---|
| Supplier invoices a month | About 260 | The firm's own count for this example |
| Proportion checked against a document rather than memory | Fewer than one in three | The firm's own estimate, not a measured study |
| Money paid in error each month | Not known | The firm cannot say, which is why it wanted this looked at |
The last row is the important one, and it is left as "not known" on purpose. The firm does not have a number, and the report does not invent one for it.
Findings
Four findings are shown from this workflow. Each is presented as it appears in the full report: what is happening, the root cause beneath it, what it costs, how sure we are, and the evidence it rests on.
Finding 1 — Invoices are approved against memory, not against a document.
| Field | Content |
|---|---|
| What is happening | Most supplier invoices are paid after a check against what somebody remembers agreeing, because no written order exists to compare them with. |
| Root cause | Orders are placed by phone or at the merchant's counter, so the price and quantity agreed are never written down before the invoice arrives. The check has nothing to check against. |
| Impact | A duplicated invoice, a price above the rate agreed, or a quantity billed and never delivered can all pass unnoticed. The firm cannot say how often, because there is no record of what was agreed. |
| Priority | Medium-term |
| Confidence | CONFIRMED. Described consistently by the office manager and the director, and visible in the firm's own purchasing paperwork reviewed during the audit. |
| Evidence | The firm's own description of the weekly payment run, and the absence of a written order behind the majority of the month's invoices. |
Finding 2 — The rates agreed with the merchant live in email, not on the order.
| Field | Content |
|---|---|
| What is happening | The prices agreed with the builders' merchant are held in email between the director and the merchant's account manager. They are not written onto the purchase order or held anywhere the person checking the invoice can see them. |
| Root cause | Rate negotiations happen director to account manager, and the outcome is never carried across into the ordering paperwork the office uses. |
| Impact | An invoice charged above the agreed rate looks correct to the person checking it, because the agreed rate is not in front of them. The overcharge is invisible at the point it would be caught. |
| Priority | Quick win |
| Confidence | CONFIRMED. The audit saw directly where the agreed rates are held, in the director's email rather than in any document the office uses to check an invoice. |
| Evidence | The email thread holding the agreed rates, and the ordering paperwork the office works from, which does not carry them. |
Finding 3 — Subcontractor invoices under the Construction Industry Scheme carry a tax judgement that must stay with a person.
| Field | Content |
|---|---|
| What is happening | Subcontractor invoices are split between labour and materials, and that split affects the tax Kestrel deducts. The materials figure has to be checked before the payment is made, and that check is done by whoever prepares it that month. |
| Root cause | The Construction Industry Scheme places the materials check on the contractor as a named responsibility (CIS 340, paragraph 3.13). It is a legal duty, not an administrative step a tool can take over. |
| Impact | An overstated materials figure means too little tax is deducted, and paragraph 3.13 puts the duty to check it, and so the responsibility for getting it wrong, on the contractor, Kestrel. This is a compliance exposure, not a time saving to be found. |
| Priority | Strategic (a boundary to protect, not a problem to fix) |
| Confidence | CONFIRMED against HMRC guidance, for the materials-check duty paragraph 3.13 establishes. The wider monthly-return and penalty regime is real but sits in other parts of the scheme and is not claimed here from paragraph 3.13. |
| Evidence | HMRC, Construction Industry Scheme: a guide for contractors and subcontractors (CIS 340), paragraph 3.13: "The contractor must always check and keep records to make sure that the part of the payment for materials supplied is not overstated." Published 31 March 2014, last updated 8 July 2026, checked 24 July 2026. |
Finding 4 — Duplicate and overcharged payments are probably occurring, but the firm cannot yet measure them.
| Field | Content |
|---|---|
| What is happening | On the routes where invoices are not checked against a document, the firm believes some invoices are paid twice or paid at the wrong figure. It has no count. |
| Root cause | There is no measurement in place. Errors that are not looked for do not announce themselves, so they are neither caught nor counted. |
| Impact | Unknown by amount. The point of this finding is the not-knowing, which is itself the risk the firm is carrying. |
| Priority | Medium-term |
| Confidence | HYPOTHESIS. This is a reasoned inference from the current state, not a measured fact. It becomes measurable only once a check is running and its results are recorded. |
| Evidence | Findings 1 and 2 above, and the firm's own inability to state a figure for money paid in error. |
Note the labels. Findings 1, 2 and 3 are marked CONFIRMED, because each was evidenced in the audit and cross-checked against a second source or against published guidance. Finding 4 is marked HYPOTHESIS, because it is an inference and not yet a measurement. You are never handed a guess dressed as a fact.
Summary statement for this workflow
The picture in buying and paying is of a process that works on the large invoices, where attention naturally goes, and runs on memory everywhere else. The single thread beneath the findings is that the firm agrees things it never writes down, so the check at the end of the week has less to check against than it appears to. None of this is a failure of care by the people doing it. It is a records gap that leaves them checking against memory and carrying a cost they cannot size.
Where these findings led
A Findings Report stops at the diagnosis. The decision and the build are a separate step, and you agree the diagnosis first. This excerpt shows where these particular findings led only so that you can judge the quality of the thinking, and the other half of the Specimen, the operating record, is that decision written out in full.
- Finding 2 led to a records fix the firm owns, not to a tool. The agreed rates were written onto the orders. A tool would not have solved this, because the information it needed did not exist in any document. The honest answer was a process change inside the firm.
- Finding 3 led to an explicit ruling that AI has no business here. The subcontractor invoices under the Construction Industry Scheme were placed outside any tool, with the reason written down, so that nobody quietly widens the scope later. Where the law puts a judgement on a person, it stays with the person.
- Findings 1 and 4 led to the one place a tool did earn a role: reading each invoice alongside the order and the delivery note and pointing at where they disagree, so that the check reaches the smaller invoices as well as the large ones. How that tool is bounded, tested and checked, including the one test it was caught failing, is the operating record.
The reason to read the two halves together is that the finding you might expect a consultancy to seize on, the one that looks most like a case for putting a new tool in place, is the one where the honest answer was to change a habit and buy nothing. Telling you where a tool does not belong is as much of the work as telling you where it does.
Illustrative, plainly
Kestrel Building Services does not exist. Its people, its suppliers, its invoice counts and every number here were invented for this example and are internally consistent with the assumptions the example states about itself. They are not market claims, they are not a promised result, and they must not be read as either.
This excerpt carries no statistic about how common any of this is in real businesses. That is deliberate. The Specimen stands on the quality of one worked diagnosis, not on a claim about a population. The external authority named in Finding 3, HMRC's CIS 340, is real, public and dated, and you can go and read it.
What you see first in the Specimen is the method, the same shape in any sector. The version filled in for a business like yours is one step in. When we work with you, we produce one for your own.
The most useful answer was the one that sold you nothing.
Read the two halves together and one thing stands out. The finding that looks most like a case for putting a new tool in place, the messy invoice checking, is not where the money went. The agreed rates were fixed by writing them onto the order, a change the firm made itself. The subcontractor invoices were placed outside any tool, because the law puts that judgement on a person. Only one narrow job was left for a tool, and even there a person opens every flag it raises. Telling you where AI does not belong is as much of the work as telling you where it does. That is what you are judging when you read the Specimen.
Illustrative, plainly.
The firm in this Specimen is invented, and so is everybody in it. Its people, its suppliers, its invoice counts and every figure were invented for this example and are internally consistent with the assumptions it states about itself. They are not market claims, they are not a promised result, and they must not be read as either.
The Specimen carries no claim about how common any of this is in real businesses. It stands on the quality of one worked diagnosis and one worked decision. The external authorities it names, HMRC and the UK financial and accountancy regulators, are real, public and dated, and you can go and read them.
When you have read it, talk to us.
If the Specimen shows you the quality you are looking for, the next step is a free Discovery Call. Forty-five minutes, no pitch. We learn how your business runs and tell you honestly whether there is a real opportunity to help. You leave with a straight read on where you stand, whether or not you ever go further.
Book a Discovery CallFree, forty-five minutes, no pitch.
Send an enquiry instead