In auditing, a test basis means the auditor examines a representative sample of transactions or balances rather than reviewing every item, and then uses the results of that sample to draw conclusions about the entire account or population. Virtually every modern audit works this way, whether it’s performed by an outside accounting firm, an internal compliance team, or the IRS. Checking every invoice, receipt, and journal entry in a real business would take years and cost more than the audit is worth, while a properly designed sample can surface the same problems at a fraction of the effort.
When an audit report says the work was done “on a test basis,” it means the auditor applied procedures to less than 100 percent of the items and relied on those results to reach an opinion. The American Institute of Certified Public Accountants sets the rules in AU-C Section 530 for audits of nonpublic companies. For public companies, the Public Company Accounting Oversight Board’s AS 2315 governs the same activity. Both standards permit two approaches: statistical sampling, which uses probability theory to quantify the precision of the results, and non-statistical sampling, which relies on the auditor’s professional judgment. Either can produce enough evidence for a professional opinion when applied correctly.
What the Auditor Is Trying to Learn
Sampling serves two very different purposes, and the purpose drives everything else about the test.
The first is a test of controls. The auditor is checking whether an internal control actually works. A sample of transactions gets pulled, and each one is checked for a yes-or-no outcome: did the control operate correctly or not? For example, an auditor might pull 40 purchase orders and check whether each one carried a manager’s approval before payment. The result is a rate, such as 2 out of 40 lacking approval. Because the auditor is counting how many items have or lack a specific characteristic, this is called attribute sampling.
The second is a substantive test of details. Here the auditor is verifying dollar amounts. Instead of a pass/fail question, the auditor measures the actual value of each sampled item and compares it to what the books show. Confirming 60 accounts receivable balances against what customers say they owe is a classic example. Because the test deals with continuous dollar amounts rather than discrete outcomes, this is called variable sampling.
The distinction matters when you read audit findings. A high error rate in a test of controls means the auditor cannot rely on that control and will do more substantive work to compensate. A projected misstatement from a substantive test means the recorded balance itself may be wrong.
How Auditors Select the Sample
The method used to pick items is not arbitrary. The goal is that every item in the population has a known chance of being selected, so the results can fairly be generalized to the whole. Picking items that “look interesting” defeats the purpose.
Random Selection
Software assigns a random number to every item in the population and picks the required number. Each item has an equal probability of being chosen. Modern audit software can do this instantly across millions of transactions.
Systematic Selection
The auditor picks a random starting point and then selects every nth item. If the population contains 10,000 invoices and the sample size is 200, the interval is 50: pick a random start between 1 and 50, then take every 50th invoice. This only works when the population isn’t already sorted in a way that correlates with what’s being tested. If invoices are sorted by dollar amount, systematic selection can skew the results.
Monetary Unit Sampling
This is the most common method for substantive testing of account balances, and it treats each dollar, rather than each transaction, as the sampling unit. A $50,000 receivable has 50,000 chances of being selected; a $500 receivable has 500. Larger-dollar items are therefore far more likely to land in the sample, which is what auditors want, because bigger items carry more risk of a material error. It’s especially useful for accounts receivable, inventory, and loan portfolios where a handful of large balances dominate.
Block Selection
Picking a cluster of contiguous items, like every transaction from a specific week or department, is block selection. It’s the least reliable way to generalize to a full population, because one block may not reflect what happened at other times during the year. Auditors use it sparingly, usually as a supplement to other methods.
How the Auditor Decides the Sample Is Big Enough
Sample size isn’t pulled out of the air. It’s driven by materiality, risk, and expected error.
Materiality is the dollar threshold above which an error would change a reasonable reader’s interpretation of the financial statements. Tolerable misstatement is narrower: the most error the auditor can accept in a particular account balance while still concluding the balance is fairly stated. It’s always set lower than overall materiality, because the auditor has to leave room for the possibility that several accounts each contain some error. The lower the tolerable misstatement, the larger the sample has to be.
Risk matters too. When the auditor judges the risk of material misstatement to be high, because controls are weak, the account involves estimates, or the industry is prone to fraud, the sample gets bigger. Where strong controls and other corroborating evidence support the balance, a smaller sample may be sufficient. If prior audits found errors in an account, the auditor also increases the sample to see whether the problem persists.
And the auditor has to know the population is complete before sampling begins. A sample drawn from an incomplete list can miss whole categories of transactions. Auditors pull populations from general ledgers, transaction logs, or year-end reports and reconcile the totals to the financial statements first.
The IRS’s Statistical Thresholds
If it’s an IRS sampling audit, the government applies specific numbers. Under Revenue Procedure 2011-42, a valid statistical sample must produce a final estimate at a 95 percent one-sided confidence level.1Internal Revenue Service. Revenue Procedure 2011-42 Statistical Sampling Procedures and Evaluation Criteria If the relative precision of the estimate is 10 percent or less, the IRS may use the point estimate (the single most likely number). If relative precision falls between 10 and 15 percent, the estimate is calculated using a formula that splits the difference between the point estimate and the confidence limit. Those numbers define how much statistical uncertainty the government is allowed to carry into a proposed adjustment against you.
What the Results Actually Mean
After testing the sampled items, the auditor evaluates every error found and works out what the sample says about the population as a whole.
For substantive tests, the auditor projects the misstatements found in the sample onto the full population. If 3 percent of the sampled dollar amount was misstated, the auditor estimates that roughly 3 percent of the entire account balance contains errors, and compares that projected misstatement to the tolerable misstatement set during planning. When projected misstatement exceeds tolerable misstatement, the auditor concludes the account may be materially misstated and either expands testing, requests corrections, or modifies the audit opinion.
For tests of controls, the auditor compares the observed deviation rate (the percentage of sampled items where the control didn’t operate) to the tolerable deviation rate. If the observed rate is too high, the auditor cannot rely on the control and has to perform more substantive work to compensate.
Inconclusive results don’t just get set aside. When findings sit in an ambiguous zone, auditing standards require the auditor to expand the sample, apply alternative procedures, or both, until reaching a definitive conclusion.
Sampling Risk and Non-Sampling Risk
Any time an auditor examines less than 100 percent of the data, there’s a chance the sample won’t perfectly reflect the population. Auditing standards draw a sharp line between two kinds of risk.
Sampling risk is the possibility that the auditor’s conclusion from a sample differs from what a full examination would have shown. In substantive testing, this appears as the risk of incorrect acceptance (concluding a balance is fine when it isn’t) or incorrect rejection (concluding there’s a problem when there isn’t). Larger samples reduce sampling risk. Statistical sampling lets the auditor measure it; non-statistical sampling requires the auditor to use judgment to keep it acceptable.
Non-sampling risk covers everything else that can go wrong. Using a procedure that doesn’t fit the objective (like confirming recorded receivables when the real concern is unrecorded ones), failing to spot an error in a document that was examined, or misinterpreting the results all count. Non-sampling risk exists even in a 100 percent examination. It’s reduced through planning, supervision, and professional skepticism, not through larger samples.
The distinction has legal teeth. When an audit misses a material misstatement, whether the failure was sampling risk or non-sampling risk can decide whether the auditor faces liability. A well-designed sample that happens to miss a problem is an inherent limit of the method. An auditor who looked at the right document and missed an obvious forgery has a non-sampling problem, and a much harder position to defend.
Whether Sampling-Based Findings Hold Up
Audit conclusions reached through sampling carry real legal weight. The IRS Office of Chief Counsel and the Department of Justice have jointly concluded that “substantial authority exists for the determination of tax deficiencies based on statistical samples.”2Internal Revenue Service. IRM 4.47.3 Statistical Sampling Auditing Techniques Courts have upheld IRS sampling-based assessments in cases like Norfolk Southern Corp. v. Commissioner and Catalano v. Commissioner.3Internal Revenue Service. Field Directive: Use of Sampling Methodologies in Research Credit Cases
A sampling-based determination is only as strong as its methodology, though. The IRS Internal Revenue Manual requires that estimates of tax adjustments be “statistically sound and legally defensible.”2Internal Revenue Service. IRM 4.47.3 Statistical Sampling Auditing Techniques If the population was incomplete, the selection wasn’t truly random, or the confidence level fell short of the 95 percent threshold, a taxpayer can challenge the results. The IRS itself acknowledges that a disallowance based on sampling won’t be upheld unless the sample follows sound statistical principles, or unless the taxpayer agreed in writing to accept the results of a limited audit.3Internal Revenue Service. Field Directive: Use of Sampling Methodologies in Research Credit Cases
On the financial statement side, PCAOB standards require auditors to design responses to identified risks of material misstatement, and those responses include properly designed sampling procedures.4Public Company Accounting Oversight Board. AS 2301 The Auditors Responses to the Risks of Material Misstatement An auditor who follows the standards and still misses a misstatement isn’t automatically liable. Auditing works under a “reasonable assurance” standard, not a guarantee, and discovering a misstatement after the fact doesn’t by itself prove negligence.
What Happens After a Test Basis Audit
The consequences depend on who ran the audit and what turned up.
In an IRS examination, the proposed population adjustment is calculated so that 95 percent of the time it won’t exceed what a full examination would have found.1Internal Revenue Service. Revenue Procedure 2011-42 Statistical Sampling Procedures and Evaluation Criteria The examiner issues a proposed examination report showing the adjustments. If you disagree, you generally have 30 days to request a conference with IRS Appeals, or 60 days to file a formal written protest. The overall timeline varies with complexity: mail audits often wrap up in a few months, while field audits of business returns can run a year or longer.
For public companies, the stakes reach beyond the audit itself. When an audit reveals that previously issued financial statements can’t be relied on, the company must file a Form 8-K with the SEC within four business days disclosing that fact under Item 4.02.5U.S. Securities and Exchange Commission. Current Report on Form 8-K Frequently Asked Questions Material weaknesses in internal controls found through testing can trigger restatements, revised filings, and market fallout.
Record Retention
The paper trail from sampling work has its own rules. Section 802 of the Sarbanes-Oxley Act requires accounting firms to retain audit workpapers, including all documents containing conclusions, opinions, analyses, and financial data related to the audit, for seven years after the audit concludes.6U.S. Securities and Exchange Commission. Retention of Records Relevant to Audits and Reviews That applies to audits of public company issuers. For the organizations being audited, the IRS generally recommends keeping records that support items on a tax return until the statute of limitations for that return runs out, typically three years, or six years where substantial understatement of income is involved.
These rules exist for a reason. If a sampling-based conclusion is challenged years later in Tax Court, in litigation, or during a regulatory investigation, the workpapers are the primary evidence that the sample was properly designed, executed, and evaluated. Destroying them early can turn a defensible audit into an indefensible one.