Poradnik
How to check a report prepared by AI? Sources, completeness and checking time
A report prepared by AI should be assessed for consistency with the sources, completeness, currency and usefulness for further work. The time to generate the document is only part of the assessment. A pilot should also include preparing the data, verification and corrections, because these decide the total human effort.
This applies to information reports, summaries and digests of documents. Working through sources and drawing recommendations from them should have a separately defined scope and acceptance criteria.
What to establish before the first test
Write down what the report will be used for and what information it must contain. Define the period, the list of sources and what to do when documents conflict or something is missing. A gap should be visible in the result rather than replaced by a plausible-sounding answer.
For figures you need the unit, the period and the source. Say whether you expect a value quoted literally, a calculation or an interpretation. That way the reviewer knows what they are actually checking.
Before you pass material into a chosen environment, agree access and processing rules with the people responsible. That is part of preparing the project. A quality test does not automatically settle whether the tool may already work on all the company’s data.
How to build quality control for a report
Prepare a list of required items and check each one by the same rules. Useful categories are: correct, missing, contradicted by the source, and needs clarification. Distinguish a factual error from a change of wording.
References to documents make verification easier, but the reviewer still has to check whether the place cited actually supports the statement. For important figures, go back to the source document.
Keep both the first version of the report and the version after correction. That shows the final quality and the effort needed to reach it. A pilot report also benefits from a record of the changes a person made.
How to compare manual work with work using AI
The comparison should cover similar material and the same expected result. If the first version measures the whole process and the second only generation, the numbers do not show a saving in work.
Worked example, all figures fictional: we compare preparing a report from 120 documents. In the manual version a person spends 480 minutes, with AI 120 minutes. The tool additionally runs for 20 minutes.
| Activity in the model comparison | Manual | With AI |
|---|---|---|
| Preparing the material | 90 min | 35 min |
| Human review and drafting | 300 min | 0 min in this scenario |
| Human checking and corrections | 90 min | 85 min |
| Total human working time | 480 min | 120 min |
| Additional tool running time | Not applicable | 20 min |
Under these assumptions human effort falls by 360 minutes, that is 75%. If in the AI version the steps run one after another without breaks, the whole run takes 140 minutes. Human working time and waiting time for a result are different quantities.
Worked example · fictional figures
Human working time [minutes / report] · hover a segment to see the stage
75% less human work in this example
Additionally: 20 minutes of tool running time.
Run one after another without breaks: 140 minutes for the whole process.
Is one correct report enough?
A single test lets you recognise the possibility and the problems. Further tests should cover material of varying quality, missing documents and cases where the data does not agree. Check too whether the result can be reproduced after the set of sources changes.
Continuing the fictional example: the report contains 38 correct items out of 40 required, and two need completing. That is 95% correct items in this sample. After corrections the document may meet every requirement, but the pilot assessment should still take account of those two interventions and the time they took.
The number of cases needed for a decision depends on how varied the documents are and on the consequences of an error. What matters is that the board knows the scope of the tests carried out and what has not yet been checked.
What should come out of the pilot
The outcome is material for a decision: quality, the time for the whole process, costs, conditions of use and the next step. At dyrektor.ai we help define those criteria and coordinate the people needed for the assessment. If the scope needs narrowing, use the criteria for choosing a pilot project.