pyxaf¶
Read and validate every iteration of the Dutch Auditfile Financieel in Python.
pyxaf is an open-source library for the Auditfile Financieel, the Dutch standard export of a general ledger that accounting software produces for tax inspectors and auditors. It reads all seven iterations of the format into one version-independent model, streams files of any size in constant memory, and validates them locally, without uploading anything, against the rules of the detected version.
| Iteration | Year | Identified by | Structure source in pyxaf | XSD check (pyxaf[xsd]) |
|---|---|---|---|---|
ADF (CLAIR1.00.00) |
1999 | fixed-width ASCII, first 12 bytes CLAIR1.00.00 |
official manual (field positions) | no XSD (fixed width) |
CLAIR2 (CLAIR2.00.00) |
2003 | no namespace, header/auditfileVersion |
official 2003 XSD, bundled | yes |
| XAF 3.0 | ~2008 | http://www.auditfiles.nl/XAF/3.0 (unverified) |
derived catalogue (no public XSD) | no |
| XAF 3.1 | ~2010 | http://www.auditfiles.nl/XAF/3.1 |
derived catalogue (no public XSD) | no |
| XAF 3.2 | 2014 (+ 2017 addendum) | http://www.auditfiles.nl/XAF/3.2 |
official XSD, bundled | yes |
| XAF 3.2.1 | 2024 | ODB namespace …/BCPP/1.0/…/XmlAuditfileFinancieel3.2.1 |
official XSD, bundled | yes |
| XAF 4.0 (4.0.3) | 2025 | ODB namespace …/BCPP/1.1/…/XmlauditfileXAF_4.0 |
official XSD, bundled | yes |
See Versions for how pyxaf tells them apart.
From 1 January 2027 only XAF 4.0 is accepted
The Belastingdienst announced that from 1 January 2027 it accepts only XAF 4.0 auditfiles. Files in older iterations stay relevant for years (seven-year retention), and pyxaf keeps reading them.
Features¶
- Zero runtime dependencies. The core uses only the standard library (
pyexpat,decimal,zipfile,gzip,csv,json) and is pure Python, so it fits small environments such as AWS Lambda. Heavier libraries are opt-in extras and are only imported when you use them. - Streaming in constant memory. Master data (company, accounts, relations, VAT codes, periods, opening balance) is read when the file is opened; transactions and lines are parsed lazily as you iterate. Memory does not grow with file size, and real files reach several GB.
- Lenient reading with graded findings. pyxaf never refuses a file because of data problems.
Every irregularity becomes a finding with a stable code (
XAF5009), a severity (ERROR,WARNING,INFO) and a position. Values aredecimal.Decimal, neverfloat, and the raw text is always kept next to the parsed value. - A local validator that checks encoding, XML, version, structure (an XSD-like check without dependencies), references, control totals and balance, uniqueness and data quality, including the official XAF 4.0 consistency rules [0001]–[0010]. It runs in one streaming pass on files of any size. See Validation.
- Normalized tables with fixed schemas for every iteration, exported to CSV and JSONL with the standard library, or to Arrow, polars, pandas and Parquet through extras. See Interoperability.
- RGS support: load the official RGS Excel release you downloaded and check every account's RGS code against it. See RGS.
- Hardened against hostile files: DOCTYPE/ENTITY declarations are refused, no "recover" parser mode, limits on nesting depth, text size and decompression. See Security.
Installation¶
pyxaf needs Python 3.11 or newer. Optional features are installed as extras:
| Extra | Installs | Enables |
|---|---|---|
xsd |
lxml ≥ 6.1.3 | validation against the bundled official XSDs (validate(..., xsd=True)) |
arrow |
nanoarrow | Arrow PyCapsule export (pyarrow.table(...), DuckDB, polars) and record batches |
polars |
polars | to_polars() (no pyarrow needed) |
pandas |
pandas, pyarrow | to_pandas() with Arrow-backed dtypes |
parquet |
pyarrow | export(..., format="parquet") |
cli |
Typer | the pyxaf command line |
all |
everything above |
pip install "pyxaf[cli]" # command line
pip install "pyxaf[polars]" # polars DataFrames (no pyarrow)
pip install "pyxaf[all]" # everything
Arrow data is built with nanoarrow
to_polars(), to_pandas() and Parquet export build their Arrow data with nanoarrow, which
the polars, pandas and parquet extras install. A missing package raises
MissingExtraError naming the extra to install.
Quick start¶
No auditfile at hand? Download a synthetic sample (2024-broken.xaf is the same ledger with one
mistyped amount):
curl -L -O https://raw.githubusercontent.com/SpireflyHQ/pyxaf/main/examples/2024.xaf \
-O https://raw.githubusercontent.com/SpireflyHQ/pyxaf/main/examples/2024-broken.xaf
import pyxaf
with pyxaf.open("2024.xaf") as af: # any iteration; .gz/.zip too
print(af.version, af.company.name, af.header.fiscal_year)
for tx in af.transactions(): # streamed, constant memory
if not tx.balanced:
print("unbalanced:", tx.journal_id, tx.number)
print(af.opening_balance().by_account()) # signed, debit-positive
report = pyxaf.validate("2024.xaf") # one streaming pass
print(report.ok, len(report.errors), len(report.warnings))
Or from the shell (with pyxaf[cli]):
Continue with Reading auditfiles.
Disclaimer¶
pyxaf is an independent open-source project. It is not affiliated with, endorsed by or supported by the Belastingdienst (Dutch Tax and Customs Administration), its Ontwikkelaars Documentatie Belastingdienst (ODB) or validation service (VTS), or the Taakgroep RGS. A clean pyxaf report does not mean that the Belastingdienst will accept a file; use the official validation service for that. pyxaf does not bundle RGS data; see RGS.
pyxaf is MIT-licensed. Please never attach real client auditfiles to issues: they contain confidential financial and personal data (see Security).