From a sample of twenty five to the full population, with the coverage and the definition in the file

Insights
Key takeaways

The internal auditor draws a sample because the population is too large to read, while the source system has already recorded for every order, invoice or case who did what, when and in which order. DIVE reconstructs from that how the process actually ran, sets it against the drawn design with the controls per step, and tests the full population against eight controls. Every control test is a record with a version number; every execution records the population, the coverage and a frozen copy of the definition, so that the March outcome stays attached to the March definition. A hit is an indicator and not a finding: first confirm it against the source document, and only then a finding with the evidence trail attached. This explanation shows the model from event log to audit file, with what the auditor takes from it and where it chafes in practice.

A sample of twenty five items tells you something about those twenty five items. The source system, meanwhile, has recorded for every order, invoice or case who did what, when and in which order. DIVE, the process mining app from Audirium, reconstructs from that record how the process actually ran, sets it against the design and tests the full population against eight controls. The result is an audit finding with the evidence trail attached to it, and only in second instance a process diagram.

In brief

What DIVE is, in three lines.

8
control tests
Segregation of duties, duplicates, authorisation threshold, master data change before payment, time window, sequence gap, reconciliation and Benford
100%
of the population
Every test runs on the complete export, with the coverage alongside it: the share on which the test could actually form a judgement
1
frozen definition per test run
Population, coverage and the exact definition are fixed per execution, with a version number

DIVE is process mining for internal audit. You load an export from a source system (ERP, case management, service desk) and DIVE reconstructs which steps were taken, in which order, how long they took and by whom. Up to that point every process mining platform does the same.

The difference is in the last step. DIVE sets the reconstructed practice against the process as it was drawn, including the controls attached to each step, and turns the difference into a test you can explain to a third party.

The outcome then reads: this step carries the control Credit check and was skipped in 591 of 10,000 cases. With the population, the coverage and the rule used alongside it, so a reviewer can recalculate it.

What lands on the auditor's desk

Five places where the sample falls short, and one where the full population puts you on the wrong foot instead.

coverage

The sample misses precisely the exception

A control that was bypassed in 591 of 10,000 cases usually does not surface in a draw of forty files.

A bank account that changes shortly before a payment certainly does not: that pattern is close to impossible to find with a sample, and a matter of seconds on the full population.

segregation of duties

Tested on rights, not on what happened

An access rights analysis shows who is allowed to. It does not see who actually did it through an emergency account, stacked roles or a stand-in. Those routes are the interesting ones.

repeatability

Last month is no evidence for this month

If the mandate limit is adjusted along the way, the March outcome silently moves to the new definition. Without version control on the test itself, there is no saying any more what was tested against back then.

source system

The event log has to be dug up first

Three fields are mandatory: a case id, an activity and a timestamp that includes the time. A general ledger export with only a posting date is not enough, because the order within a day is then undetermined. What each Dutch source system does and does not deliver is set out below.

data quality

An administration with gaps is a finding in itself

Missing case ids, unreadable timestamps, duplicates, batch postings, cases with a single step. An export in which 8% of the rows have no case id deserves its own sentence in the report, not a silent correction.

exceptions

Not every deviation is an error

A quotation that comes back signed for amendment is rework, not a bypassed control.

If a signal touches more than 80% of the cases, it is the norm and not a deviation. If the tool does not make that distinction, a test on the full population also delivers the full population in noise.

That last point is why greater coverage does not deliver more assurance by itself.

A list of four thousand hits is not a work plan, and a hit a reviewer cannot retrace is worth less in a file than a sample that can be retraced.

So the question has two halves: how you get from twenty five to everything, and how you then get from everything back to the twenty cases that actually matter, without anything dropping out of the file along the way.

The principle: the test as a recorded item

An exploratory analysis is a button: you point at a column, you get an outcome, and once the screen closes nothing has been kept.

For an exploration that is exactly right. For the question which controls did I test on the full population this quarter, and at what threshold it is too little.

A control test in DIVE is therefore a record: a name, a type, the parameters, a reference to the control in the control framework, a weight from 1 to 5, and a version number.

Every execution leaves a test run with the population size, how much of it could be assessed, the coverage, the fingerprint of the source columns and a frozen copy of the definition as it stood at that moment.

Adjust a threshold and the version increments, and the March outcome stays attached to the March definition. Change only the name and it does not increment, because that does not move the outcome.

Population and coverage are not the same. Next to the population, everywhere, stands the share of cases on which the test could form a judgement.

In a threshold test a case without a readable amount falls outside it, in segregation of duties a case without a recorded performer.

Quietly leaving out that difference would make the coverage look better than it is, and that number is exactly what "we tested everything" stands or falls on.

The same principle runs through the rest of the app.

A hit is an indicator and not a finding; DIVE enforces the order: first confirm against the source document or dismiss it as a false positive, and only then can it be turned into a finding.

An exploration (what happens to the outcome if the norm were somewhere else) gets its own card with a dashed border and the mark exploration, not an executed test, records no execution and puts no hit at all into the work list. In an audit file that is the difference between a finding and a thought experiment.

The model underneath

Six layers. Each layer builds on the previous one, and nothing is retyped afterwards.

Layer
What goes in
What comes out
Event logthe export, loaded
CSV, TSV or Excel up to 10 GB, sent in blocks of 8 MB and resumable after a failed attempt. You map the columns yourself; DIVE proposes a mapping on the names and measures every date format against the actual values, so you can see for yourself whether 03-04 is March or April. The performer is irreversibly pseudonymised on import, what you do not map is not stored, and the raw file disappears once the log is in place.
DeliversA fixed event log per dataset, with a data quality report and the SHA-256 of the source export in the file.
▼  every question starts with a scope, even an empty one
Scopethe filter stack
Six filter types, stackable in any order: time window, attribute, start and end, performance, sequence and variant. The order really is an order: each filter looks at what the previous one left and carries its own coverage (how many cases went in, how many came out). A filter that cannot run on this export is shown as not applied, instead of being skipped in silence.
DeliversOne population that every figure afterwards refers to, visible in a bar above every screen, and as a line in the file as soon as you record it.
▼  discovered from the events, not from a drawing
Analysishow the process ran
Process discovery with nodes by frequency and arrows by waiting time, variants with a Pareto curve, throughput time (median, 90th percentile, rework, in working days if you prefer), segregation of duties on what actually happened, handovers between pseudonymised performers, and the four statistical tests that appear in every audit programme: Benford, duplicates, gaps and stratification.
DeliversEvery figure is an entry point: click it and the cases underneath open up, all the way to the full course of a single case.
▼  practice alongside the design
Testingdeviation from the norm
Conformance against a process diagram from Flowmap: control bypassed, undrawn path, dead path. The eight control tests as recorded items with test runs. And a rule based risk score per case that stacks the signals DIVE already measures, with the reason it scores stated in words for every case.
DeliversHits, ranked on a score from 0 to 100 you can retrace, in a single work list.
▼  confirm first, only then a finding
Follow-upfinding and action
A confirmed hit becomes a finding with the evidence trail attached: dataset, period, number of events, the measured count and the rule used. Every finding can be given actions with an owner, a deadline and a priority; the status of the finding follows from those actions. The layer underneath is the same shared core GRIP uses, not a status model of its own.
DeliversA list of findings per organisation, detached from a single export: a finding from March stays in place when you work with different data in April.
▼  the question that comes three years later
Filewhat happened
A frozen line for every operation: what went in and what came out, with the version of DIVE and of the calculation engine. Export as an Excel workbook (everything, a sheet per component), PDF (summary first, long lists truncated with a note of how much was not printed) or JSON. Alongside that the audit committee summary, which recalculates nothing but reads the stored outcomes.
DeliversThe answer to where did that 5.9% come from, including what was not measured.

The calculation engine is DuckDB with the event log as a Parquet file; the calculation core is built in house, without an external process mining library. A process graph over five million events costs roughly a third of a second.

Data is stored separately per organisation, as with every app in the suite.

Process discovery in DIVE: the process drawn from the events, with the thickness of an arrow as the number of times a path was taken
Process discovery on a demo dataset: 132,179 events from 22,000 cases, traced back to seven activities and 22 paths. The thickness of an arrow is the number of times that path occurs. The slider filters out the rare paths, and those rare paths are usually exactly where an audit has to go.

What the auditor takes from it

What you put in the report, and what you have to say alongside it.

eight tests

The test that shows the difference

Master data change before payment finds a master data item that changes within N days before a payment, in the same case.

Alongside it: segregation of duties, duplicates on one or more attributes, an amount above the mandate limit without an approval step, postings outside office hours or after the period close, missing numbers in a sequence, the reconciliation against a control total, and Benford, here as exceptions above the 95% threshold and explicitly as an attention signal, not as evidence.

ranking

A score you can retrace

Every exception gets 0 to 100 points from four components: financial significance (up to 40), degree of deviation within its own type (up to 25), repetition by the same person involved (up to 15) and the weight of the control (up to 20).

Every component is stated in words with the hit, because in the right of reply you have to be able to explain why this exception ranks above that one.

conformance

Control bypassed, in numbers and in a picture

Against the diagram DIVE reports three kinds of deviation: a skipped step that carries a control (with the name, type and owner of that control), a transition that occurs in the data but was not drawn, and a drawn transition that does not occur in the data.

The spatial view lays the drawing flat and hangs the actual paths above it: an arc that goes over a gateway is literally a bypassed control.

Conformance test in DIVE: bypassed controls with counts, next to the drawn path against the path actually taken
The conformance test on the same demo dataset: two bypassed controls, sixteen paths that were taken but never drawn, and one drawn path that does not occur in the data. The reconciliation with the general ledger was skipped in 3,161 of 23,797 cases (13.3 percent), with the type and owner of the control alongside it. Top right the drawn path lies against the path actually taken: the arc over the crossed out step is the bypassed control.
period against period
The cases behind a single process step in DIVE, with steps, throughput time, location and amount band per case
From a step in the diagram to the cases underneath: the 21,933 cases in which the activity "Uitbetaling samengesteld" (payment prepared) occurs, with the number of steps, the throughput time, the location and the amount band per case. From here you draw a sample or go on to the source document in the source system.

Not measured is not zero

Two exports side by side, with a comparability check first: overlapping periods, a markedly different length, a gap between them, different source columns.

A segregation of duties rule that was not reinstated in the second round stays empty, with the pointer to apply the recipe.

At the bottom the trend across all rounds, because two points are a difference and only a line shows whether an improvement holds.

explorations

What if the norm were different

The sensitivity analysis recalculates the same test across a range of thresholds around the established norm, so you can see whether a finding hangs on the norm or stands apart from it.

The dry run first shows a new definition what it would yield, including how many of those hits the team already assessed as not a finding.

Sampling precision puts attribute sampling and monetary unit sampling side by side, with the skewness of the population as the reason for the advice. All four carry the mark exploration and do not count as an executed test.

sampling

A sample a reviewer can repeat

Even on a full population you want to check a number of cases in the source system.

DIVE draws at random or monetary, from all cases or from the hits, and records the seed in the file. With the same seed, population and method a reviewer gets exactly the same cases back.

The audit committee summary delivers the work in the form it can take on a meeting agenda: what was examined, what was found, what was tested and what was not, what is being done about it, and the difference with the previous round.

Every sentence follows from a figure and a fixed threshold; no data goes to a language model, not even to turn it into nicer sentences.

The caveat is a section of its own and not a footnote: analyses that were not run, a coverage below one hundred percent (with the sentence that those cases remained untested), an active scope, records that dropped out on import.

What the source systems deliver

Every CSV or Excel works, but not every system records enough for process mining. The state of play for Dutch systems in practice, from the documentation:

SystemUsableNotes
Municipal case management systems (ZGW)yesThe VNG standard includes an audit trail with user, action and timestamp per case. Some installations still run on the older ZDS standard.
Dynamics 365 Business CentralyesAudit system fields on every table plus the Change Log with user, date and time.
TOPdeskyesProgress trail per incident or change.
AFAS ProfitprobablyGetConnectors are flexible enough; whether workflow steps carry both a time and an employee has to be determined per environment.
Exact Online, TwinfieldlimitedOnly creation and modification moments per record, no status history. One event per posting.
SAPvia extractCDHDR/CDPOS is the richest audit source there is, but it delivers an extract, not a connection.

How it runs in practice

From export to finding, and from finding to the next export.

  1. Load the export

    Drag the export in. DIVE proposes the columns on their names and shows, for every possible date format, the parse percentage and the date range on the actual values; you confirm.

    This is where process mining analyses most often go quietly wrong, so here it is made explicit.

  2. Assess the data quality

    Before anything is calculated you see what is off: missing case ids, unreadable timestamps, duplicates, batch postings, cases with a single step.

    That report goes into the file. Decide here whether the export is usable and what you will say about it in the report.

  3. Analyse and set the scope

    Process discovery, variants, throughput time, segregation of duties and handovers on the full population.

    Set the scope with the filter stack; what you type is a draft until you record it, so you can try things out without creating a line in the file you will never base a conclusion on.

  4. Test against the design

    Link a diagram from Flowmap. DIVE proposes, on name similarity, which log activity belongs to which process step; you confirm that mapping yourself, because a wrong mapping produces a tidy looking but incorrect finding.

    After that: bypassed controls, undrawn paths, dead paths.

  5. Record and run the control tests

    Record a test per control with parameters, weight and a reference to the control framework, after a dry run if you want.

    Every execution is a test run with population, coverage and frozen definition. Run it again and only the hits not yet assessed disappear; what you already ticked off stays.

  6. Validate, finding, action

    Confirm a hit against the source document, through a deep link to the record in your source system (DIVE then keeps no document at all) or an attached item.

    Only after that does it become a finding, with the evidence trail, and does it get actions with an owner and a deadline.

  7. Record it as a recipe

    Once the analysis stands the way you want it next time, record the parameters with the source system.

    Next month's export is recognised by the fingerprint of the column names.

    Applying it is never blind: first a dry run with fits, fits in part or cannot per component, and whatever cannot is named in the execution report. A period filter deliberately does not carry over.

Practical

The questions that usually come second.

  • Hosted in the European Union. On our own servers at Hetzner in Nuremberg, Germany. Data is stored separately per organisation.
  • No client data leaves the server. No cell value, case, amount or name goes from this app to an external service or to a language model. No figure, judgement or conclusion comes from a model: the analyses, scores and coverage percentages run on fixed rules and give the same outcome tomorrow on the same data.
  • Privacy by design. The performer field is pseudonymised on import with a salt per dataset, the link back is not kept, and what you do not map is not stored. If an identity is needed for the right of reply, that request goes to the administrator of the source system; DIVE records the request. An event log with a performer field touches employee representation; the app is built so that conversation becomes easier, not harder.
  • Traceable to GIAS Standard 14.1. Information is only sufficient when a knowledgeable third party can repeat the work and reach the same conclusion. The audit file exports precisely that: the column mapping, the data quality report, the assessment of every hit, the seed of every sample and the hash of the source export.
  • Large exports. Up to 10 GB per file, resumable, with a disk space check beforehand. A log of tens of millions of events is calculated once and read from the stored outcomes after that.
  • Dutch and English. The interface is bilingual and follows your preference across the whole suite. Your own content (activity names, finding texts) stays as you enter it; a few sentences composed by the server, such as filter descriptions, are at this moment still only in Dutch.

Where it stands now. DIVE has been in beta since August 2026: in use with real data, by a limited group of organisations with a role on the app, and not yet generally available.

The core (loading, analysis, conformance, the eight control tests, file and recipes) is in place; the explorations date from the end of August 2026 and are therefore younger. I would rather say that up front than have you find it out along the way.

What is deliberately not in it: task mining at workstation level, prediction of the remaining duration of cases in progress, and an AI assistant that answers questions or drafts findings. For that last one a separate consideration applies that has not been made yet, and as long as it has not, no figure at all comes out of a model.

Back to Insights