Skip to content

API reference

Generated from the docstrings with mkdocstrings. Everything listed here under pyxaf.<module> is also importable from the top-level package where noted: pyxaf.open, pyxaf.detect, pyxaf.validate, pyxaf.AuditFile, pyxaf.ValidationReport, the model classes, pyxaf.Finding, pyxaf.Severity, pyxaf.CODES, pyxaf.FormatInfo, pyxaf.Version, pyxaf.Family, pyxaf.NamespaceStatus, pyxaf.RawRecord and the exceptions. Import pyxaf.rgs and pyxaf.tables explicitly (import pyxaf.rgs).

Page section Contents
Reading open(), AuditFile, RawView (af.raw)
Detection detect()
Validation validate(), ValidationReport
Findings Finding, Severity, CODES, CodeInfo, FindingCollector
Formats FormatInfo, Version, Family, NamespaceStatus
Data model Header, Company, LedgerAccount, Transaction, Line, …
Raw records RawRecord
Tables Tables, Table, Column, TABLES
RGS load_excel(), RgsSchema, RgsCode, validate_refs()
Exceptions PyxafError and subclasses

pyxaf.open

open(source: SourceLike | Sequence[SourceLike], *, encoding: str | None = None, repair: Iterable[str] = (), negative_amounts: NegativePolicy = 'flip', max_findings_per_code: int | None = 100, max_findings: int | None = 10000, severity_overrides: Mapping[str, Severity | None] | None = None, max_depth: int = DEFAULT_MAX_DEPTH, max_text_size: int = DEFAULT_MAX_TEXT, max_decompressed_size: int | None = None) -> AuditFile

Open an auditfile of any iteration (ADF, CLAIR2, XAF 3.0–4.0).

Parameters:

Name Type Description Default
source SourceLike | Sequence[SourceLike]

Path, bytes or binary file object — or a sequence of them for a split auditfile (continuation files "Vervolgbestand x van y", or complete files per period). gzip and single-member zip archives are decompressed transparently.

required
encoding str | None

Override the declared encoding (e.g. "cp1252" for mis-declared files).

None
repair Iterable[str]

Opt-in fix-ups, each reported as a finding: "control-chars" (remove XML-illegal control characters), "latin1-as-cp1252" (decode declared ISO-8859-1 as windows-1252), "bare-ampersand" (escape & that starts no entity).

()
negative_amounts NegativePolicy

"flip" (default): a negative amount counts on the opposite side (−100 C = 100 D); "abs": the sign is ignored.

'flip'
max_findings_per_code int | None

Keep at most this many findings per code.

100
max_findings int | None

Keep at most this many findings in total.

10000
severity_overrides Mapping[str, Severity | None] | None

Change the severity of codes, or silence them with None.

None
max_depth int

Maximum XML nesting depth.

DEFAULT_MAX_DEPTH
max_text_size int

Maximum text length of one element.

DEFAULT_MAX_TEXT
max_decompressed_size int | None

Maximum decompressed size of gzip/zip input.

None

Returns:

Type Description
AuditFile

An AuditFile; use it as a context manager.

Raises:

Type Description
NotAnAuditfileError

The input is not an auditfile.

XmlSyntaxError

The XML is not well-formed (raised when the damage is reached).

ForbiddenConstructError

The XML contains a DOCTYPE/ENTITY declaration.

pyxaf.AuditFile

AuditFile(source: SourceLike | Source | Sequence[SourceLike | Source], *, encoding: str | None = None, repair: Iterable[str] = (), negative_amounts: NegativePolicy = 'flip', max_findings_per_code: int | None = 100, max_findings: int | None = 10000, severity_overrides: Mapping[str, Severity | None] | None = None, max_depth: int = DEFAULT_MAX_DEPTH, max_text_size: int = DEFAULT_MAX_TEXT, max_decompressed_size: int | None = None, _observer: Observer | None = None, _value_findings: bool = True, _findings: FindingCollector | None = None, _on_parts: Callable[[Sequence[FormatInfo]], None] | None = None)

An opened auditfile (or multi-file set): master data eagerly, transactions streamed.

Use pyxaf.open() to create one. Master data (header, company, accounts, relations, VAT codes, periods, opening balance) is read when the file is opened; transactions are parsed lazily each time transactions() or lines() is iterated (re-reading the source), so memory use does not depend on file size.

header instance-attribute

header: Header

File header.

company instance-attribute

company: Company

The administration (company).

format instance-attribute

format: FormatInfo = self._parts[0].info

Detection result of the (first) file.

findings property

findings: tuple[Finding, ...]

Findings collected while reading so far (grows as transactions are iterated).

files property

files: tuple[str, ...]

Names of the underlying files.

journals property

journals: Mapping[str, Journal]

Journals by ID.

In XML files journals enclose their transactions, so the first access may scan the whole file (cheaply, skipping transaction content).

transaction_totals property

transaction_totals: TransactionTotals | None

Control totals declared for the transactions (summed over a multi-file set).

raw property

raw: RawView

The raw records: exact text, no normalization (the fastest way to stream).

tables property

tables: Tables

Normalized tables for export (see pyxaf.tables).

transactions

transactions() -> Iterator[Transaction]

Stream all transactions in file order (each call re-reads the source).

lines

lines() -> Iterator[Line]

Stream all transaction lines (each carries its journal/transaction context).

opening_balance

opening_balance(strategy: Literal['auto', 'element', 'transactions'] = 'auto') -> OpeningBalance

Return the opening balance, wherever the file stores it.

Parameters:

Name Type Description Default
strategy Literal['auto', 'element', 'transactions']

"element" uses openingBalance; "transactions" collects the lines of transactions in period 0 or in journals of type O (scans the file); "auto" (default) uses the element when present, else the transactions.

'auto'

to_polars

to_polars(tables: Iterable[str] | None = None, **kwargs: Any) -> dict[str, Any]

Return polars DataFrames by table name (needs pyxaf[polars]).

to_pandas

to_pandas(tables: Iterable[str] | None = None, **kwargs: Any) -> dict[str, Any]

Return pandas DataFrames by table name (needs pyxaf[pandas]).

export

export(directory: str | PathLike[str], **kwargs: Any) -> list[str]

Write the normalized tables to directory (see pyxaf.tables.Tables.export()).

close

close() -> None

Release paused parsers (sources are opened per pass and closed after each).

pyxaf.reader.RawView

RawView(af: AuditFile)

Raw records of an AuditFile, exactly as written.

header property

header: RawRecord | None

The header element (ADF: the header line).

company property

company: RawRecord | None

The company element without its sections (CLAIR2/ADF: the header).

transactions

transactions() -> Iterator[tuple[str | None, RawRecord]]

Stream (journal_id, transaction_record) pairs (ADF: one record per line).

pyxaf.detect

detect(source: SourceLike, *, encoding: str | None = None) -> FormatInfo

Detect the iteration, namespace status and encoding of an auditfile cheaply.

Only the first 64 KiB (after decompression) are inspected.

Parameters:

Name Type Description Default
source SourceLike

Path, bytes or binary file object (gzip/zip are decompressed transparently).

required
encoding str | None

Override the declared encoding.

None

Raises:

Type Description
NotAnAuditfileError

If the input is not an auditfile at all.

EncryptedAuditfileError

For encrypted/vendor-compressed auditfiles.

ForbiddenConstructError

If the prolog contains a DOCTYPE.

pyxaf.validate

validate(source: SourceLike | Sequence[SourceLike], *, xsd: bool = False, rules: RuleSet = 'spec', rgs: RgsSchema | None = None, encoding: str | None = None, repair: Iterable[str] = (), negative_amounts: NegativePolicy = 'flip', severity_overrides: Mapping[str, Severity | None] | None = None, max_findings_per_code: int | None = 100, max_findings: int | None = 10000, max_depth: int = DEFAULT_MAX_DEPTH, max_text_size: int = DEFAULT_MAX_TEXT, max_decompressed_size: int | None = None) -> ValidationReport

Validate an auditfile (or multi-file set) in one streaming pass.

Parameters:

Name Type Description Default
source SourceLike | Sequence[SourceLike]

Path, bytes, binary stream, or a sequence of them for a split auditfile.

required
xsd bool

Also validate against the official XSD (needs pyxaf[xsd]).

False
rules RuleSet

"spec" (everything, per the detected version) or "vts" (only what the Belastingdienst's validation service checks: encoding, XML, schema and the numbered 4.0 consistency rules).

'spec'
rgs RgsSchema | None

An RgsSchema to check RGS codes against.

None
encoding str | None

Override the declared encoding.

None
repair Iterable[str]

Opt-in repairs (see pyxaf.open()).

()
negative_amounts NegativePolicy

Sign policy (see pyxaf.open()).

'flip'
severity_overrides Mapping[str, Severity | None] | None

Change severities per code, or silence codes with None.

None
max_findings_per_code int | None

Keep at most this many findings per code.

100
max_findings int | None

Keep at most this many findings in total.

10000
max_depth int

Maximum XML nesting depth.

DEFAULT_MAX_DEPTH
max_text_size int

Maximum text length of one element.

DEFAULT_MAX_TEXT
max_decompressed_size int | None

Maximum decompressed size of gzip/zip input.

None

Returns:

Type Description
ValidationReport

A ValidationReport. Problems in the data never raise; only unreadable input (not an auditfile, unsupported source) does.

pyxaf.ValidationReport dataclass

ValidationReport(format: FormatInfo, findings: tuple[Finding, ...], counts: Mapping[str, int], suppressed: int, checked: tuple[str, ...], stats: Mapping[str, int], rules: RuleSet = 'spec', severity_counts: Mapping[Severity, int] = dict())

Result of validate().

Attributes:

Name Type Description
format FormatInfo

Detection result.

findings tuple[Finding, ...]

Findings, most severe first, then in file order.

counts Mapping[str, int]

Occurrences per code, including findings suppressed by limits.

suppressed int

Number of findings dropped because of limits.

checked tuple[str, ...]

Validation layers that ran (see module documentation).

stats Mapping[str, int]

Counts gathered during the pass (lines, transactions, accounts…).

rules RuleSet

Rule set used.

severity_counts Mapping[Severity, int]

Occurrences per severity, including findings suppressed by limits. The verdict (ok, max_severity) is based on these, so limiting the number of findings never turns an invalid file into a valid one.

errors property

errors: tuple[Finding, ...]

Findings with severity ERROR.

warnings property

warnings: tuple[Finding, ...]

Findings with severity WARNING.

ok property

ok: bool

True when there are no ERROR findings (including suppressed ones).

max_severity property

max_severity: Severity | None

Highest severity among all findings (including suppressed ones).

to_dict

to_dict() -> dict[str, Any]

JSON-serialisable representation (schema_version 1).

to_json

to_json(**kwargs: Any) -> str

Serialise to_dict() as JSON.

pyxaf.findings

Graded, coded findings.

Every problem pyxaf notices in a file is reported as a Finding with a stable code (XAF<nnnn>), a Severity and the location where it occurred. Codes are grouped:

1xxx input container and character encoding
2xxx XML well-formedness and security
3xxx version, namespace and structure
4xxx references between records
5xxx control totals and balance
6xxx uniqueness
7xxx data quality and conventions
8xxx RGS (Referentie Grootboekschema)

The official XAF 4.0 consistency rules map to codes ending in their rule number (rule [0009] → XAF5009). CODES documents every code.

Severity

Bases: IntEnum

How serious a finding is. Ordered: INFO < WARNING < ERROR.

INFO class-attribute instance-attribute

INFO = 10

Notable, but neither wrong nor likely to affect analysis.

WARNING class-attribute instance-attribute

WARNING = 20

Likely a data problem, or something that affects analysis.

ERROR class-attribute instance-attribute

ERROR = 30

Violates the specification of the detected version.

CodeInfo dataclass

CodeInfo(code: str, severity: Severity, title: str, rule_ref: str | None = None)

Documentation of a finding code.

Finding dataclass

Finding(*, code: str, severity: Severity, message: str, file: int = 0, line: int | None = None, column: int | None = None, path: str | None = None, value: str | None = None, rule_ref: str | None = None)

One observation about an auditfile.

Attributes:

Name Type Description
code str

Stable code, e.g. "XAF5009"; see CODES.

severity Severity

Severity after rule-set adjustments and overrides.

message str

Human-readable English message.

file int

Index of the file within a multi-file set (0 for single files).

line int | None

1-based line number of the element (None if not applicable).

column int | None

0-based column number.

path str | None

Element path (/auditfile/company/...) or ADF field name.

value str | None

The offending raw value, if any.

rule_ref str | None

Reference to the official rule, e.g. "[0009]".

to_dict

to_dict() -> dict[str, object]

Return a JSON-serialisable dictionary.

FindingCollector dataclass

FindingCollector(max_per_code: int | None = 100, max_total: int | None = 10000, overrides: Mapping[str, Severity | None] = dict(), file: int = 0, _items: list[Finding] = list(), _counts: Counter[str] = Counter(), _severities: Counter[Severity] = Counter(), _keys: set[tuple[object, ...]] = set())

Collects findings with per-code and overall limits; counts what it suppresses.

Parameters:

Name Type Description Default
max_per_code int | None

Keep at most this many findings per code (None = unlimited).

100
max_total int | None

Keep at most this many findings overall (None = unlimited).

10000
overrides Mapping[str, Severity | None]

Severity per code, replacing the default severity. Mapping a code to None silences it.

dict()

counts property

counts: Mapping[str, int]

Number of occurrences per code, including suppressed ones.

severity_counts property

severity_counts: Mapping[Severity, int]

Number of findings per (effective) severity, including suppressed ones.

max_severity property

max_severity: Severity | None

Highest severity of all findings, including suppressed ones.

suppressed property

suppressed: int

Number of findings dropped because of limits.

add

add(code: str, message: str, *, line: int | None = None, column: int | None = None, path: str | None = None, value: str | None = None, severity: Severity | None = None, file: int | None = None) -> None

Record a finding (counted even when suppressed by a limit).

extend

extend(findings: Iterable[Finding]) -> None

Add already-built findings (respecting limits and overrides).

snapshot

snapshot() -> tuple[Finding, ...]

Return the findings collected so far.

pyxaf.formats

Iterations of the Auditfile Financieel and how they are identified.

KNOWN_NAMESPACES module-attribute

KNOWN_NAMESPACES: Mapping[str, tuple[Version, NamespaceStatus]] = MappingProxyType({'http://www.auditfiles.nl/xaf/3.0': (Version.XAF30, NamespaceStatus.UNVERIFIED), 'http://www.auditfiles.nl/xaf/3.1': (Version.XAF31, NamespaceStatus.OFFICIAL), 'http://www.auditfiles.nl/xaf/3.2': (Version.XAF32, NamespaceStatus.OFFICIAL), 'http://www.odb.belastingdienst.nl/belastingdienst/bcpp/1.0/structures/xmlauditfilefinancieel3.2.1': (Version.XAF321, NamespaceStatus.OFFICIAL), 'http://www.odb.belastingdienst.nl/belastingdienst/bcpp/1.1/structures/xmlauditfilexaf_4.0': (Version.XAF40, NamespaceStatus.OFFICIAL), 'http://www.odb.belastingdienst.nl/xaf/4.0': (Version.XAF40, NamespaceStatus.DOCUMENTED_VARIANT), 'http://www.auditfiles.nl/xaf/4.0': (Version.XAF40, NamespaceStatus.KNOWN_BOGUS), 'http://www.auditfiles.nl/xaf/2.0': (Version.CLAIR2, NamespaceStatus.KNOWN_BOGUS)})

Family

Bases: StrEnum

Structural family of an auditfile.

ADF class-attribute instance-attribute

ADF = 'adf'

Fixed-width ASCII auditfile (CLAIR1.00.00, 1999).

CLAIR2 class-attribute instance-attribute

CLAIR2 = 'clair2'

First XML auditfile (CLAIR2.00.00, 2003), no namespace.

XAF class-attribute instance-attribute

XAF = 'xaf'

XML Auditfile Financieel 3.0 – 4.0.

Version

Bases: StrEnum

Iteration of the Auditfile Financieel. Compares equal to its string value ("4.0").

family property

family: Family

Structural family of this version.

NamespaceStatus

Bases: StrEnum

How trustworthy the root element's namespace is.

OFFICIAL class-attribute instance-attribute

OFFICIAL = 'official'

The target namespace of the official XSD.

DOCUMENTED_VARIANT class-attribute instance-attribute

DOCUMENTED_VARIANT = 'documented-variant'

Stated in official documentation but not the XSD target namespace.

UNVERIFIED class-attribute instance-attribute

UNVERIFIED = 'unverified'

Plausible, but no official source could be found (XAF 3.0).

KNOWN_BOGUS class-attribute instance-attribute

KNOWN_BOGUS = 'known-bogus'

Does not officially exist, but occurs in files from some exporters.

UNKNOWN class-attribute instance-attribute

UNKNOWN = 'unknown'

Not in pyxaf's table.

NONE class-attribute instance-attribute

NONE = 'none'

No namespace (CLAIR2, ADF, some exporters).

FormatInfo dataclass

FormatInfo(*, family: Family | None, version: Version | None, namespace: str | None, namespace_status: NamespaceStatus, encoding: EncodingInfo, confidence: float, reasons: tuple[str, ...], continuation: tuple[int, int] | None = None, compression: str | None = None, root: str | None = None)

Result of format detection.

Attributes:

Name Type Description
family Family | None

Structural family (None only if undeterminable).

version Version | None

Detected iteration (best guess when confidence < 1).

namespace str | None

Namespace URI of the root element, as written.

namespace_status NamespaceStatus

Trust level of the namespace.

encoding EncodingInfo

Encoding information (BOM, declaration, effective codec).

confidence float

0–1; 1 means the official markers and the vocabulary agree.

reasons tuple[str, ...]

Human-readable explanation of how the version was decided.

continuation tuple[int, int] | None

(x, y) for a continuation file ("Vervolgbestand x van y"), else None.

compression str | None

"gzip", "zip" or None.

root str | None

Local name of the root element (None for ADF).

bom property

bom: bool

Whether the file starts with a byte-order mark.

pyxaf.models

Normalized, version-independent data model.

All classes are slotted, keyword-only dataclasses with English names. Master-data classes are frozen; the per-line classes (Line, Transaction, VatLine, ForeignAmount) are not, because frozen construction costs ~3 µs per object on the hot path — treat them as read-only all the same. Every object keeps the RawRecord it was built from in raw (the exact text as written), and Line.extra/Transaction.extra expose fields of that record without a normalized attribute. The "Normalized field mapping" section of the version guide documents which element of each version maps to which attribute.

Code lists (journal type, account type, relation type) are data, not closed enums: the written code is always kept next to the interpreted kind.

Side

Bases: StrEnum

Debit or credit.

JournalKind

Bases: StrEnum

Interpretation of a journal type code (jrnTp).

JournalType dataclass

JournalType(code: str | None, kind: JournalKind)

A journal type: the written code and its interpretation.

from_code classmethod

from_code(code: str | None) -> JournalType

Interpret an XAF jrnTp (B C G M O P S T Y Z) or a CLAIR2 free-text type.

AccountKind

Bases: StrEnum

Interpretation of a ledger account type.

AccountType dataclass

AccountType(code: str | None, kind: AccountKind)

A ledger account type: the written code and its interpretation.

from_code classmethod

from_code(code: str | None) -> AccountType

Interpret accTp (B/M/P) or a CLAIR2/ADF free-text type (heuristic).

RelationKind

Bases: StrEnum

Interpretation of a customer/supplier type (custSupTp).

RelationType dataclass

RelationType(code: str | None, kind: RelationKind)

A relation type: the written code and its interpretation.

from_code classmethod

from_code(code: str | None) -> RelationType

Interpret custSupTp (B/C/S; also O, and D/F as written by some exporters/ADF).

RgsRef dataclass

RgsRef(*, raw: str, code: str | None, extension: str | None, source: str, placeholder: bool = False)

A reference from a ledger account to an RGS code.

Attributes:

Name Type Description
raw str

The value as written.

code str | None

The RGS reference code (BIvaKouVvp) or None if not recognisable.

extension str | None

Anything appended after the code (sub-codes, .01…), stripped.

source str

Where the value came from: RGScode (4.0), taxonomy (3.2 conceptRef), taxonomies (3.1), leadReference/leadCode/leadCrossRef (vendor heuristics).

placeholder bool

Whether the value is a known placeholder (0000, RGS-1…).

parse classmethod

parse(raw: str, source: str) -> RgsRef

Split raw into code and extension and recognise placeholders.

Address dataclass

Address(*, kind: Literal['street', 'postal'], street: str | None = None, number: str | None = None, number_extension: str | None = None, property: str | None = None, city: str | None = None, postal_code: str | None = None, region: str | None = None, country: str | None = None, raw: RawRecord | None = None)

A street or postal address.

Header dataclass

Header(*, fiscal_year: str | None, start_date: date | None, end_date: date | None, currency: str | None, created: date | None, software_name: str | None, software_version: str | None, rgs_version: str | None = None, declared_version: str | None = None, raw: RawRecord | None = None)

File header.

fiscal_year is the raw string ("2024" or broken years "2023-2024"). declared_version is CLAIR2's auditfileVersion or ADF's version field.

Company dataclass

Company(*, name: str | None, identifier: str | None = None, commerce_number: str | None = None, tax_registration_country: str | None = None, tax_registration_id: str | None = None, addresses: tuple[Address, ...] = (), raw: RawRecord | None = None)

The administration (company) the auditfile is about.

LedgerAccount dataclass

LedgerAccount(*, id: str, description: str | None, account_type: AccountType, lead_code: str | None = None, lead_description: str | None = None, rgs: RgsRef | None = None, seq: int = 0, file: int = 0, raw: RawRecord | None = None)

A general-ledger account.

Relation dataclass

Relation(*, id: str, name: str | None, relation_type: RelationType, contact: str | None = None, tax_registration_country: str | None = None, tax_registration_id: str | None = None, commerce_number: str | None = None, email: str | None = None, telephone: str | None = None, website: str | None = None, addresses: tuple[Address, ...] = (), opening_balance: Decimal | None = None, closing_balance: Decimal | None = None, seq: int = 0, file: int = 0, raw: RawRecord | None = None)

A customer and/or supplier.

opening_balance/closing_balance (4.0 opBalDesc/clBalDesc, which despite their names are amounts) are signed, debit-positive.

VatCode dataclass

VatCode(*, id: str, description: str | None, payable_account_id: str | None = None, receivable_account_id: str | None = None, seq: int = 0, raw: RawRecord | None = None)

A VAT code definition.

Period dataclass

Period(*, key: str, number: int | None, start_date: date | None = None, end_date: date | None = None, description: str | None = None, seq: int = 0, raw: RawRecord | None = None)

An accounting period.

key is the period number as written (opaque, e.g. "01", "501"); number is its integer value when numeric.

Journal dataclass

Journal(*, id: str, description: str | None, journal_type: JournalType, offset_account_id: str | None = None, bank_account: str | None = None, seq: int = 0, file: int = 0, raw: RawRecord | None = None)

A journal (dagboek).

VatLine dataclass

VatLine(*, code: str | None, percentage: Decimal | None, amount: Decimal | None, side: Side | None, signed_amount: Decimal | None, raw: RawRecord | None = None)

VAT information on a transaction line. signed_amount is debit-positive.

ForeignAmount dataclass

ForeignAmount(*, currency: str | None, amount: Decimal | None, signed_amount: Decimal | None = None, exchange_rate: Decimal | None = None)

Amount in a foreign currency. exchange_rate is only given by ADF.

Line dataclass

Line(*, seq: int, number: str | None = None, account_id: str | None = None, amount: Decimal | None, side: Side | None, signed_amount: Decimal | None, debit: Decimal | None, credit: Decimal | None, description: str | None = None, document_ref: str | None = None, effective_date: date | None = None, settlement_date: date | None = None, relation_id: str | None = None, invoice_ref: str | None = None, order_ref: str | None = None, receiving_doc_ref: str | None = None, shipping_doc_ref: str | None = None, cost_center: str | None = None, cost_unit: str | None = None, product: str | None = None, project: str | None = None, work_cost_arrangement: str | None = None, bank_account: str | None = None, offset_bank_account: str | None = None, quantity: Decimal | None = None, vat: tuple[VatLine, ...] = (), foreign: ForeignAmount | None = None, journal_id: str | None = None, transaction_number: str | None = None, transaction_seq: int | None = None, period_key: str | None = None, transaction_date: date | None = None, file: int = 0, raw: RawRecord | None = None)

A transaction line (or opening-balance line), flattened with its transaction context.

Amounts

amount is the value as written (may be negative) and side the written debit/credit indicator. signed_amount is debit-positive after applying the negative-amount policy; debit/credit are its non-negative presentation (one of them is zero). All are None when the written amount is invalid (a finding is reported).

Context

journal_id, transaction_number, transaction_seq, period_key and transaction_date repeat the enclosing transaction's values.

extra property

extra: Mapping[str, str]

Fields of the raw record without a normalized attribute (version/vendor specific).

Transaction dataclass

Transaction(*, seq: int, journal_id: str | None, number: str | None, description: str | None, period_key: str | None, period_number: int | None, date: date | None, source: str | None = None, user: str | None = None, lines: tuple[Line, ...] = (), file: int = 0, raw: RawRecord | None = None)

A journal entry with its lines.

period_key is the period number as written (opaque); period_number its integer value.

total_debit property

total_debit: Decimal

Sum of the lines' debit values (invalid amounts count as zero).

total_credit property

total_credit: Decimal

Sum of the lines' credit values (invalid amounts count as zero).

balanced property

balanced: bool

Whether the lines balance (debit = credit).

extra property

extra: Mapping[str, str]

Fields of the raw record without a normalized attribute.

TransactionTotals dataclass

TransactionTotals(*, lines_count: int | None, total_debit: Decimal | None, total_credit: Decimal | None)

Control totals declared in the file.

linesCount/numberEntries, totalDebit and totalCredit as parsed values; None where absent or invalid.

OpeningBalance dataclass

OpeningBalance(*, source: Literal['element', 'transactions', 'none'], lines: tuple[Line, ...], date: date | None = None, description: str | None = None, declared: TransactionTotals | None = None)

Unified opening balance, wherever the file stores it.

Attributes:

Name Type Description
source Literal['element', 'transactions', 'none']

"element" (openingBalance), "transactions" (period-0 transactions or journal type O), or "none".

lines tuple[Line, ...]

The opening-balance lines (signed amounts follow the negative-amount policy).

date date | None

opBalDate (3.x) if given.

description str | None

opBalDesc (3.x) if given.

declared TransactionTotals | None

Control totals as written in the file (element only).

total_debit property

total_debit: Decimal

Sum of debit values.

total_credit property

total_credit: Decimal

Sum of credit values.

balanced property

balanced: bool

Whether debit equals credit.

by_account

by_account() -> dict[str, Decimal]

Signed (debit-positive) opening balance per account ID.

pyxaf.RawRecord

RawRecord(tag: str, fields: dict[str, str], children: dict[str, list[RawRecord]] | None, line: int, text: str | None = None, sequence: list[list[Any]] | None = None)

One XML element with its leaf children as fields and complex children as children.

Attributes:

Name Type Description
tag str

Local element name (namespace prefix removed).

fields Mapping[str, str]

Text of leaf child elements by local name, exactly as written (entities resolved, whitespace preserved; empty elements give ""). If a leaf element occurs more than once, the first occurrence is in fields and every occurrence is also available as a text-only child record in children.

children Mapping[str, Sequence[RawRecord]]

Complex child elements (and repeated leaves) by local name, in document order.

line int

1-based line of the start tag.

text str | None

Text content for a leaf element stored as a child, else None.

sequence list[list[Any]] | None

Only during validation: the names of all child elements in document order, as [name, count] runs of consecutive equal names (None otherwise).

get

get(name: str, default: str | None = None) -> str | None

Return the text of leaf child name.

child

child(name: str) -> RawRecord | None

Return the first complex child named name.

all

all(name: str) -> Sequence[RawRecord]

Return all complex children named name.

iter_children

iter_children() -> Iterator[RawRecord]

Iterate over all complex children.

to_dict

to_dict() -> dict[str, object]

Return a plain dict (children become lists of dicts).

pyxaf.tables

Normalized tables for analysis and export.

Every table has a fixed schema (column names and logical types) that is the same for all iterations. Tables are produced lazily: master-data tables from memory, transactions, lines and line_vat by streaming the file. All rows carry synthetic seq keys because real files contain duplicate numbers.

Table Rows
header one row: file header and detected version
company one row
addresses company and customer/supplier addresses
accounts ledger accounts (with RGS reference)
relations customers/suppliers
vat_codes VAT codes
periods periods
journals journals
transactions journal entries (with line counts and debit/credit totals)
lines transaction lines with their transaction context
line_vat VAT details of transaction lines (line_seq → lines.seq)
opening_balance opening-balance lines (from the element or period-0/opening transactions)

Logical types: string, int64, date, bool and decimal(p,s). Amounts are decimal(20,2) as in the XSDs, VAT percentages decimal(8,3).

TABLES module-attribute

TABLES: dict[str, tuple[Column, ...]] = {'header': _cols('format_version family fiscal_year start_date:date end_date:date currency created:date software_name software_version rgs_version declared_version'), 'company': _cols('name identifier commerce_number tax_registration_country tax_registration_id'), 'addresses': _cols('owner_type owner_id seq:int64 kind street number number_extension property city postal_code region country'), 'accounts': _cols('seq:int64 file:int64 id description account_type account_kind lead_code lead_description rgs_raw rgs_code rgs_extension rgs_source rgs_placeholder:bool'), 'relations': _cols(f'seq:int64 file:int64 id name relation_type relation_kind contact tax_registration_country tax_registration_id commerce_number email telephone website opening_balance:{AMOUNT} closing_balance:{AMOUNT}'), 'vat_codes': _cols('seq:int64 id description payable_account_id receivable_account_id'), 'periods': _cols('seq:int64 key number:int64 start_date:date end_date:date description'), 'journals': _cols('seq:int64 file:int64 id description journal_type journal_kind offset_account_id bank_account'), 'transactions': _cols(f'seq:int64 file:int64 journal_id number description period_key period_number:int64 date:date source user line_count:int64 total_debit:{AMOUNT} total_credit:{AMOUNT}'), 'lines': _cols(f'seq:int64 transaction_seq:int64 file:int64 journal_id transaction_number period_key transaction_date:date number account_id amount:{AMOUNT} side signed_amount:{AMOUNT} debit:{AMOUNT} credit:{AMOUNT} description document_ref effective_date:date settlement_date:date relation_id invoice_ref order_ref receiving_doc_ref shipping_doc_ref cost_center cost_unit product project work_cost_arrangement bank_account offset_bank_account quantity:decimal(24,6) foreign_currency foreign_amount:{AMOUNT} foreign_signed_amount:{AMOUNT} exchange_rate:decimal(24,6) vat_count:int64'), 'line_vat': _cols(f'line_seq:int64 transaction_seq:int64 index:int64 code percentage:decimal(8,3) amount:{AMOUNT} side signed_amount:{AMOUNT}'), 'opening_balance': _cols(f'seq:int64 source number account_id amount:{AMOUNT} side signed_amount:{AMOUNT} debit:{AMOUNT} credit:{AMOUNT} transaction_seq:int64 journal_id')}

Column dataclass

Column(name: str, type: str)

A table column: name and logical type.

decimal property

decimal: tuple[int, int] | None

(precision, scale) for decimal columns.

Table

Table(tables: Tables, name: str)

One normalized table: schema plus lazily produced rows.

With pyxaf[arrow] installed a table implements the Arrow PyCapsule stream interface (__arrow_c_stream__), so polars.DataFrame(table), pyarrow.table(table) and DuckDB can consume it directly.

column_names property

column_names: tuple[str, ...]

Names of the columns.

rows

rows() -> Iterator[Row]

Yield rows as tuples of Python values (Decimal, date, str, int).

dicts

dicts() -> Iterator[dict[str, Any]]

Yield rows as dictionaries.

batches

batches(batch_size: int = 65536, on_inexact: OnInexact = 'raise') -> Iterator[Any]

Yield Arrow record batches lazily (needs pyxaf[arrow]).

Each batch implements __arrow_c_array__ (a nanoarrow struct array).

to_polars

to_polars(on_inexact: OnInexact = 'raise') -> Any

Return a polars DataFrame (needs pyxaf[polars]).

to_pandas

to_pandas(on_inexact: OnInexact = 'raise') -> Any

Return a pandas DataFrame with Arrow-backed dtypes (needs pyxaf[pandas]).

Tables

Tables(af: AuditFile)

All normalized tables of an AuditFile (see module documentation).

names property

names: tuple[str, ...]

Table names.

export

export(directory: str | PathLike[str], *, format: ExportFormat = 'csv', tables: Iterable[str] | None = None, batch_size: int = 65536, on_inexact: OnInexact = 'raise') -> list[str]

Write tables to directory (one file per table); return the written paths.

csv and jsonl need no dependencies (decimals are written as exact strings, dates as ISO 8601); parquet needs pyxaf[parquet]. Streamed tables are produced in a single pass over the file.

to_polars

to_polars(tables: Iterable[str] | None = None, *, batch_size: int = 65536, on_inexact: OnInexact = 'raise') -> dict[str, Any]

Return polars DataFrames by table name (needs pyxaf[polars]; no pyarrow).

Streamed tables are built in a single pass over the file.

to_pandas

to_pandas(tables: Iterable[str] | None = None, *, on_inexact: OnInexact = 'raise') -> dict[str, Any]

Return pandas DataFrames with Arrow-backed dtypes (needs pyxaf[pandas]).

pyxaf.rgs

RGS (Referentie Grootboekschema) reference data and checks.

pyxaf does not bundle RGS data: the official RGS Excel carries no licence or reuse statement. Download the Excel release yourself (e.g. RGS 3.8-def.xlsx) and load it with load_excel(); the workbook is read with a small standard-library .xlsx reader, so no extra is needed.

Example

import pyxaf, pyxaf.rgs schema = pyxaf.rgs.load_excel("RGS-3.8-def.xlsx") # doctest: +SKIP schema.version, len(schema) # doctest: +SKIP ('3.8', 4979) schema.parent("BIvaKouVvp").code # doctest: +SKIP 'BIvaKou' report = pyxaf.validate("2024.xaf", rgs=schema) # doctest: +SKIP

RGS codes form a strict prefix hierarchy: every code consists of B (balance) or W (profit and loss) followed by three-character segments, and the parent of a code is its longest proper prefix that is itself a code. Levels run from 1 (B, W) to 5 (mutations such as …Beg, …Inv).

XlsxError

Bases: PyxafError

The workbook is not a readable .xlsx file or lacks a required part.

RgsCode dataclass

RgsCode(*, code: str, level: int | None, description: str | None, short_description: str | None = None, debit_credit: str | None = None, sort_key: str | None = None, reference_number: str | None = None, opposite_code: str | None = None, entities: frozenset[str] = frozenset(), elimination_filters: frozenset[str] = frozenset())

One RGS reference code.

Attributes:

Name Type Description
code str

The reference code, e.g. "BIvaKouVvp".

level int | None

Hierarchy level 1–5 (Nivo), None if absent or not numeric.

description str | None

Full description (Omschrijving).

short_description str | None

Short description (Omschrijving (verkort)).

debit_credit str | None

"D" or "C" (D/C) — the normal balance side.

sort_key str | None

Presentation order key (Sortering, e.g. "A.A.A010").

reference_number str | None

Referentienummer as stored. Formats are irregular ("0101010.01", "01") and some values were stored as numbers in Excel, losing a leading zero ("101015"); the value is kept exactly as written.

opposite_code str | None

ReferentieOmslagcode, the code on the other balance side used when the balance flips (e.g. a bank balance that becomes an overdraft).

entities frozenset[str]

Names of the entity-filter columns set for the code (Basis, Uitgebr, EZ/VOF, ZZP, WoCo, BV, …).

elimination_filters frozenset[str]

Names of the columns in the "Filters - te vervallen c.q. te elimineren" group that are set: the code can be dropped when the entity does not use that feature. Empty when the workbook has no such group.

RgsSchema

RgsSchema(codes: Iterable[RgsCode], *, version: str | None = None, rename_map: Mapping[str, str] | None = None, duplicate_codes: Iterable[str] = ())

One RGS release: codes, hierarchy and renames.

Usually created by load_excel().

Parameters:

Name Type Description Default
codes Iterable[RgsCode]

The codes of this release (iteration order is kept as the sheet order).

required
version str | None

The RGS version, e.g. "3.8".

None
rename_map Mapping[str, str] | None

Old code → code in this version, for codes renamed in earlier releases.

None
duplicate_codes Iterable[str]

Codes that occurred more than once in the source (the first wins).

()

Attributes:

Name Type Description
version str | None

The RGS version ("3.8") or None if it could not be determined.

codes Mapping[str, RgsCode]

Codes by reference code, in sheet order.

rename_map Mapping[str, str]

Old code → code in this version. Only codes that no longer exist in this version are included.

duplicate_codes frozenset[str]

Codes that occurred more than once in the source sheet.

resolve

resolve(code: str) -> RgsCode | None

Return the code, following rename_map for codes renamed since.

Parameters:

Name Type Description Default
code str

A reference code (surrounding whitespace is ignored).

required

Returns:

Type Description
RgsCode | None

The RgsCode in this version, or None if unknown.

parent

parent(code: str) -> RgsCode | None

Return the parent: the longest proper prefix of code that is a code.

Works for codes that are not in the schema too (the nearest existing ancestor).

children

children(code: str) -> tuple[RgsCode, ...]

Return the direct children of code in sheet order.

check_ref

check_ref(ref: RgsRef, findings: FindingCollector, *, line: int | None, account_id: str) -> None

Check one account's RGS reference against this release.

Placeholders and unrecognisable values are skipped (the validator reports them as XAF8001/XAF8002). Reports XAF8003 for a code that does not exist in this version, XAF8004 for a code that was renamed (naming the new code) and XAF8006 for a code at a level other than 4 or 5.

Parameters:

Name Type Description Default
ref RgsRef

The parsed reference.

required
findings FindingCollector

Where to report.

required
line int | None

Line of the ledger account in the file.

required
account_id str

The ledger account's ID (for the message).

required

check_version

check_version(version_text: str, findings: FindingCollector) -> None

Check the auditfile's declared RGS version (XAF 4.0 header/RGSVersion).

The element has no specified format ("3.7", "RGS 3.8", "RGS-1" occur), so it is parsed leniently with parse_version(). Reports XAF8005 when no version can be recognised or when it differs from this schema's version.

Parameters:

Name Type Description Default
version_text str

The declared version as written.

required
findings FindingCollector

Where to report.

required

parse_version

parse_version(text: str | None) -> str | None

Extract an RGS version number from free text, leniently.

Accepts "3.8", "RGS 3.8", "RGS3.8-def", "rgs-3,7", "Versie 3.5.0"; a trailing .0 is dropped. Returns None when no number is found.

Parameters:

Name Type Description Default
text str | None

Any text that may contain a version (a header value, sheet or file name).

required

Returns:

Type Description
str | None

The normalised version ("3.8") or None.

parse_ref

parse_ref(raw: str, source: str = 'RGScode') -> RgsRef

Split a raw RGS reference into code and extension (see pyxaf.models.RgsRef.parse()).

Parameters:

Name Type Description Default
raw str

The value as written in the auditfile.

required
source str

Where the value came from (RGScode, taxonomy, …).

'RGScode'

validate_refs

validate_refs(af: AuditFile, schema: RgsSchema) -> list[Finding]

Check every ledger account's RGS reference and the declared RGS version.

Reports placeholders (XAF8001), unrecognisable values (XAF8002) and everything RgsSchema.check_ref() and RgsSchema.check_version() report.

Parameters:

Name Type Description Default
af AuditFile

An open auditfile.

required
schema RgsSchema

The RGS release to check against.

required

Returns:

Type Description
list[Finding]

The findings, in account order.

load_excel

load_excel(path: str | PathLike[str] | IO[bytes], sheet: str | None = None, *, rename_sheet: str | None = None, max_member_size: int = DEFAULT_MAX_MEMBER_SIZE) -> RgsSchema

Load an official RGS Excel release.

The main sheet is detected automatically when sheet is not given: among the sheets with a Referentiecode column header in their first rows (and only one such column), names starting with Totaal are preferred, then the sheet with the most rows. A sheet with several Referentiecode columns (one per release, e.g. RGS3.8-versus-RGS3.7) provides RgsSchema.rename_map. Header matching ignores case, whitespace and punctuation.

The version is taken from the title cell above the header (RGS3.8), the sheet name, the file name or the Recap sheet, in that order.

Parameters:

Name Type Description Default
path str | PathLike[str] | IO[bytes]

Path or binary seekable stream of the .xlsx file.

required
sheet str | None

Name of the sheet with the codes (auto-detected when None).

None
rename_sheet str | None

Name of the version-comparison sheet (auto-detected when None).

None
max_member_size int

Maximum uncompressed size of a single workbook part, in bytes.

DEFAULT_MAX_MEMBER_SIZE

Returns:

Type Description
RgsSchema

The loaded RgsSchema.

Raises:

Type Description
XlsxError

If the file is not an .xlsx workbook or no sheet with RGS codes is found.

KeyError

If sheet or rename_sheet does not exist.

ForbiddenConstructError

If a workbook part contains a DOCTYPE/ENTITY declaration.

LimitExceededError

If a workbook part exceeds max_member_size.

pyxaf.errors

Exceptions raised by pyxaf.

pyxaf never refuses a file because of data problems — those become findings (see pyxaf.findings). Exceptions are reserved for inputs that cannot be read at all: not an auditfile, malformed XML, forbidden constructs, exceeded limits, missing extras.

PyxafError

Bases: Exception

Base class for all pyxaf errors.

NotAnAuditfileError

Bases: PyxafError

The input is not recognisable as any iteration of the Auditfile Financieel.

EncryptedAuditfileError

Bases: NotAnAuditfileError

The input is an encrypted or vendor-compressed auditfile (.xac, .xsc, .XFC, …).

Only the intended recipient (usually the Belastingdienst) can decrypt these.

XmlSyntaxError

XmlSyntaxError(message: str, *, line: int | None = None, column: int | None = None, file: int = 0)

Bases: PyxafError

The XML is not well-formed.

Attributes:

Name Type Description
line

1-based line number of the error (None if unknown).

column

0-based column of the error (None if unknown).

file

index of the file in a multi-file set (0 for single files).

ForbiddenConstructError

ForbiddenConstructError(message: str, *, line: int | None = None, column: int | None = None, file: int = 0)

Bases: XmlSyntaxError

The XML contains a construct pyxaf refuses for security reasons (DOCTYPE, ENTITY).

No version of the Auditfile Financieel uses a DTD; refusing them blocks entity-expansion ("billion laughs") and external-entity (XXE) attacks.

CorruptArchiveError

Bases: NotAnAuditfileError

A gzip or zip container is corrupt or truncated.

LimitExceededError

Bases: PyxafError

A configured safety limit (nesting depth, text size, decompressed size) was exceeded.

MissingExtraError

Bases: PyxafError, ImportError

An optional dependency is needed; the message says which extra to install.