API reference¶
Generated from the docstrings with mkdocstrings. Everything
listed here under pyxaf.<module> is also importable from the top-level package where noted:
pyxaf.open, pyxaf.detect, pyxaf.validate, pyxaf.AuditFile, pyxaf.ValidationReport, the
model classes, pyxaf.Finding, pyxaf.Severity, pyxaf.CODES, pyxaf.FormatInfo,
pyxaf.Version, pyxaf.Family, pyxaf.NamespaceStatus, pyxaf.RawRecord and the exceptions.
Import pyxaf.rgs and pyxaf.tables explicitly (import pyxaf.rgs).
| Page section | Contents |
|---|---|
| Reading | open(), AuditFile, RawView (af.raw) |
| Detection | detect() |
| Validation | validate(), ValidationReport |
| Findings | Finding, Severity, CODES, CodeInfo, FindingCollector |
| Formats | FormatInfo, Version, Family, NamespaceStatus |
| Data model | Header, Company, LedgerAccount, Transaction, Line, … |
| Raw records | RawRecord |
| Tables | Tables, Table, Column, TABLES |
| RGS | load_excel(), RgsSchema, RgsCode, validate_refs() |
| Exceptions | PyxafError and subclasses |
pyxaf.open
¶
open(source: SourceLike | Sequence[SourceLike], *, encoding: str | None = None, repair: Iterable[str] = (), negative_amounts: NegativePolicy = 'flip', max_findings_per_code: int | None = 100, max_findings: int | None = 10000, severity_overrides: Mapping[str, Severity | None] | None = None, max_depth: int = DEFAULT_MAX_DEPTH, max_text_size: int = DEFAULT_MAX_TEXT, max_decompressed_size: int | None = None) -> AuditFile
Open an auditfile of any iteration (ADF, CLAIR2, XAF 3.0–4.0).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
SourceLike | Sequence[SourceLike]
|
Path, bytes or binary file object — or a sequence of them for a split auditfile (continuation files "Vervolgbestand x van y", or complete files per period). gzip and single-member zip archives are decompressed transparently. |
required |
encoding
|
str | None
|
Override the declared encoding (e.g. |
None
|
repair
|
Iterable[str]
|
Opt-in fix-ups, each reported as a finding: |
()
|
negative_amounts
|
NegativePolicy
|
|
'flip'
|
max_findings_per_code
|
int | None
|
Keep at most this many findings per code. |
100
|
max_findings
|
int | None
|
Keep at most this many findings in total. |
10000
|
severity_overrides
|
Mapping[str, Severity | None] | None
|
Change the severity of codes, or silence them with |
None
|
max_depth
|
int
|
Maximum XML nesting depth. |
DEFAULT_MAX_DEPTH
|
max_text_size
|
int
|
Maximum text length of one element. |
DEFAULT_MAX_TEXT
|
max_decompressed_size
|
int | None
|
Maximum decompressed size of gzip/zip input. |
None
|
Returns:
| Type | Description |
|---|---|
AuditFile
|
An |
Raises:
| Type | Description |
|---|---|
NotAnAuditfileError
|
The input is not an auditfile. |
XmlSyntaxError
|
The XML is not well-formed (raised when the damage is reached). |
ForbiddenConstructError
|
The XML contains a DOCTYPE/ENTITY declaration. |
pyxaf.AuditFile
¶
AuditFile(source: SourceLike | Source | Sequence[SourceLike | Source], *, encoding: str | None = None, repair: Iterable[str] = (), negative_amounts: NegativePolicy = 'flip', max_findings_per_code: int | None = 100, max_findings: int | None = 10000, severity_overrides: Mapping[str, Severity | None] | None = None, max_depth: int = DEFAULT_MAX_DEPTH, max_text_size: int = DEFAULT_MAX_TEXT, max_decompressed_size: int | None = None, _observer: Observer | None = None, _value_findings: bool = True, _findings: FindingCollector | None = None, _on_parts: Callable[[Sequence[FormatInfo]], None] | None = None)
An opened auditfile (or multi-file set): master data eagerly, transactions streamed.
Use pyxaf.open() to create one. Master data (header, company, accounts, relations, VAT
codes, periods, opening balance) is read when the file is opened; transactions are parsed
lazily each time transactions() or lines() is iterated (re-reading the source),
so memory use does not depend on file size.
format
instance-attribute
¶
format: FormatInfo = self._parts[0].info
Detection result of the (first) file.
findings
property
¶
findings: tuple[Finding, ...]
Findings collected while reading so far (grows as transactions are iterated).
journals
property
¶
journals: Mapping[str, Journal]
Journals by ID.
In XML files journals enclose their transactions, so the first access may scan the whole file (cheaply, skipping transaction content).
transaction_totals
property
¶
transaction_totals: TransactionTotals | None
Control totals declared for the transactions (summed over a multi-file set).
raw
property
¶
raw: RawView
The raw records: exact text, no normalization (the fastest way to stream).
transactions
¶
transactions() -> Iterator[Transaction]
Stream all transactions in file order (each call re-reads the source).
lines
¶
lines() -> Iterator[Line]
Stream all transaction lines (each carries its journal/transaction context).
opening_balance
¶
opening_balance(strategy: Literal['auto', 'element', 'transactions'] = 'auto') -> OpeningBalance
Return the opening balance, wherever the file stores it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
strategy
|
Literal['auto', 'element', 'transactions']
|
|
'auto'
|
to_polars
¶
Return polars DataFrames by table name (needs pyxaf[polars]).
to_pandas
¶
Return pandas DataFrames by table name (needs pyxaf[pandas]).
export
¶
Write the normalized tables to directory (see pyxaf.tables.Tables.export()).
pyxaf.reader.RawView
¶
RawView(af: AuditFile)
pyxaf.detect
¶
detect(source: SourceLike, *, encoding: str | None = None) -> FormatInfo
Detect the iteration, namespace status and encoding of an auditfile cheaply.
Only the first 64 KiB (after decompression) are inspected.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
SourceLike
|
Path, bytes or binary file object (gzip/zip are decompressed transparently). |
required |
encoding
|
str | None
|
Override the declared encoding. |
None
|
Raises:
| Type | Description |
|---|---|
NotAnAuditfileError
|
If the input is not an auditfile at all. |
EncryptedAuditfileError
|
For encrypted/vendor-compressed auditfiles. |
ForbiddenConstructError
|
If the prolog contains a DOCTYPE. |
pyxaf.validate
¶
validate(source: SourceLike | Sequence[SourceLike], *, xsd: bool = False, rules: RuleSet = 'spec', rgs: RgsSchema | None = None, encoding: str | None = None, repair: Iterable[str] = (), negative_amounts: NegativePolicy = 'flip', severity_overrides: Mapping[str, Severity | None] | None = None, max_findings_per_code: int | None = 100, max_findings: int | None = 10000, max_depth: int = DEFAULT_MAX_DEPTH, max_text_size: int = DEFAULT_MAX_TEXT, max_decompressed_size: int | None = None) -> ValidationReport
Validate an auditfile (or multi-file set) in one streaming pass.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
SourceLike | Sequence[SourceLike]
|
Path, bytes, binary stream, or a sequence of them for a split auditfile. |
required |
xsd
|
bool
|
Also validate against the official XSD (needs |
False
|
rules
|
RuleSet
|
|
'spec'
|
rgs
|
RgsSchema | None
|
An |
None
|
encoding
|
str | None
|
Override the declared encoding. |
None
|
repair
|
Iterable[str]
|
Opt-in repairs (see |
()
|
negative_amounts
|
NegativePolicy
|
Sign policy (see |
'flip'
|
severity_overrides
|
Mapping[str, Severity | None] | None
|
Change severities per code, or silence codes with |
None
|
max_findings_per_code
|
int | None
|
Keep at most this many findings per code. |
100
|
max_findings
|
int | None
|
Keep at most this many findings in total. |
10000
|
max_depth
|
int
|
Maximum XML nesting depth. |
DEFAULT_MAX_DEPTH
|
max_text_size
|
int
|
Maximum text length of one element. |
DEFAULT_MAX_TEXT
|
max_decompressed_size
|
int | None
|
Maximum decompressed size of gzip/zip input. |
None
|
Returns:
| Type | Description |
|---|---|
ValidationReport
|
A |
pyxaf.ValidationReport
dataclass
¶
ValidationReport(format: FormatInfo, findings: tuple[Finding, ...], counts: Mapping[str, int], suppressed: int, checked: tuple[str, ...], stats: Mapping[str, int], rules: RuleSet = 'spec', severity_counts: Mapping[Severity, int] = dict())
Result of validate().
Attributes:
| Name | Type | Description |
|---|---|---|
format |
FormatInfo
|
Detection result. |
findings |
tuple[Finding, ...]
|
Findings, most severe first, then in file order. |
counts |
Mapping[str, int]
|
Occurrences per code, including findings suppressed by limits. |
suppressed |
int
|
Number of findings dropped because of limits. |
checked |
tuple[str, ...]
|
Validation layers that ran (see module documentation). |
stats |
Mapping[str, int]
|
Counts gathered during the pass (lines, transactions, accounts…). |
rules |
RuleSet
|
Rule set used. |
severity_counts |
Mapping[Severity, int]
|
Occurrences per severity, including findings suppressed by limits.
The verdict ( |
pyxaf.findings
¶
Graded, coded findings.
Every problem pyxaf notices in a file is reported as a Finding with a stable code
(XAF<nnnn>), a Severity and the location where it occurred. Codes are grouped:
1xxx |
input container and character encoding |
2xxx |
XML well-formedness and security |
3xxx |
version, namespace and structure |
4xxx |
references between records |
5xxx |
control totals and balance |
6xxx |
uniqueness |
7xxx |
data quality and conventions |
8xxx |
RGS (Referentie Grootboekschema) |
The official XAF 4.0 consistency rules map to codes ending in their rule number
(rule [0009] → XAF5009). CODES documents every code.
Severity
¶
Bases: IntEnum
How serious a finding is. Ordered: INFO < WARNING < ERROR.
CodeInfo
dataclass
¶
CodeInfo(code: str, severity: Severity, title: str, rule_ref: str | None = None)
Documentation of a finding code.
Finding
dataclass
¶
Finding(*, code: str, severity: Severity, message: str, file: int = 0, line: int | None = None, column: int | None = None, path: str | None = None, value: str | None = None, rule_ref: str | None = None)
One observation about an auditfile.
Attributes:
| Name | Type | Description |
|---|---|---|
code |
str
|
Stable code, e.g. |
severity |
Severity
|
Severity after rule-set adjustments and overrides. |
message |
str
|
Human-readable English message. |
file |
int
|
Index of the file within a multi-file set (0 for single files). |
line |
int | None
|
1-based line number of the element ( |
column |
int | None
|
0-based column number. |
path |
str | None
|
Element path ( |
value |
str | None
|
The offending raw value, if any. |
rule_ref |
str | None
|
Reference to the official rule, e.g. |
FindingCollector
dataclass
¶
FindingCollector(max_per_code: int | None = 100, max_total: int | None = 10000, overrides: Mapping[str, Severity | None] = dict(), file: int = 0, _items: list[Finding] = list(), _counts: Counter[str] = Counter(), _severities: Counter[Severity] = Counter(), _keys: set[tuple[object, ...]] = set())
Collects findings with per-code and overall limits; counts what it suppresses.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
max_per_code
|
int | None
|
Keep at most this many findings per code ( |
100
|
max_total
|
int | None
|
Keep at most this many findings overall ( |
10000
|
overrides
|
Mapping[str, Severity | None]
|
Severity per code, replacing the default severity. Mapping a code to
|
dict()
|
counts
property
¶
Number of occurrences per code, including suppressed ones.
severity_counts
property
¶
severity_counts: Mapping[Severity, int]
Number of findings per (effective) severity, including suppressed ones.
max_severity
property
¶
max_severity: Severity | None
Highest severity of all findings, including suppressed ones.
add
¶
add(code: str, message: str, *, line: int | None = None, column: int | None = None, path: str | None = None, value: str | None = None, severity: Severity | None = None, file: int | None = None) -> None
Record a finding (counted even when suppressed by a limit).
pyxaf.formats
¶
Iterations of the Auditfile Financieel and how they are identified.
KNOWN_NAMESPACES
module-attribute
¶
KNOWN_NAMESPACES: Mapping[str, tuple[Version, NamespaceStatus]] = MappingProxyType({'http://www.auditfiles.nl/xaf/3.0': (Version.XAF30, NamespaceStatus.UNVERIFIED), 'http://www.auditfiles.nl/xaf/3.1': (Version.XAF31, NamespaceStatus.OFFICIAL), 'http://www.auditfiles.nl/xaf/3.2': (Version.XAF32, NamespaceStatus.OFFICIAL), 'http://www.odb.belastingdienst.nl/belastingdienst/bcpp/1.0/structures/xmlauditfilefinancieel3.2.1': (Version.XAF321, NamespaceStatus.OFFICIAL), 'http://www.odb.belastingdienst.nl/belastingdienst/bcpp/1.1/structures/xmlauditfilexaf_4.0': (Version.XAF40, NamespaceStatus.OFFICIAL), 'http://www.odb.belastingdienst.nl/xaf/4.0': (Version.XAF40, NamespaceStatus.DOCUMENTED_VARIANT), 'http://www.auditfiles.nl/xaf/4.0': (Version.XAF40, NamespaceStatus.KNOWN_BOGUS), 'http://www.auditfiles.nl/xaf/2.0': (Version.CLAIR2, NamespaceStatus.KNOWN_BOGUS)})
Family
¶
Version
¶
NamespaceStatus
¶
Bases: StrEnum
How trustworthy the root element's namespace is.
OFFICIAL
class-attribute
instance-attribute
¶
The target namespace of the official XSD.
DOCUMENTED_VARIANT
class-attribute
instance-attribute
¶
Stated in official documentation but not the XSD target namespace.
UNVERIFIED
class-attribute
instance-attribute
¶
Plausible, but no official source could be found (XAF 3.0).
KNOWN_BOGUS
class-attribute
instance-attribute
¶
Does not officially exist, but occurs in files from some exporters.
FormatInfo
dataclass
¶
FormatInfo(*, family: Family | None, version: Version | None, namespace: str | None, namespace_status: NamespaceStatus, encoding: EncodingInfo, confidence: float, reasons: tuple[str, ...], continuation: tuple[int, int] | None = None, compression: str | None = None, root: str | None = None)
Result of format detection.
Attributes:
| Name | Type | Description |
|---|---|---|
family |
Family | None
|
Structural family ( |
version |
Version | None
|
Detected iteration (best guess when |
namespace |
str | None
|
Namespace URI of the root element, as written. |
namespace_status |
NamespaceStatus
|
Trust level of the namespace. |
encoding |
EncodingInfo
|
Encoding information (BOM, declaration, effective codec). |
confidence |
float
|
0–1; 1 means the official markers and the vocabulary agree. |
reasons |
tuple[str, ...]
|
Human-readable explanation of how the version was decided. |
continuation |
tuple[int, int] | None
|
|
compression |
str | None
|
|
root |
str | None
|
Local name of the root element ( |
pyxaf.models
¶
Normalized, version-independent data model.
All classes are slotted, keyword-only dataclasses with English names. Master-data classes are
frozen; the per-line classes (Line, Transaction, VatLine,
ForeignAmount) are not, because frozen construction costs ~3 µs per object on the hot
path — treat them as read-only all the same. Every object keeps
the RawRecord it was built from in raw (the exact text as written),
and Line.extra/Transaction.extra expose fields of that record without a normalized
attribute. The
"Normalized field mapping" section of the version guide documents which element of each version
maps to which attribute.
Code lists (journal type, account type, relation type) are data, not closed enums: the written
code is always kept next to the interpreted kind.
Side
¶
Bases: StrEnum
Debit or credit.
JournalKind
¶
Bases: StrEnum
Interpretation of a journal type code (jrnTp).
JournalType
dataclass
¶
JournalType(code: str | None, kind: JournalKind)
A journal type: the written code and its interpretation.
from_code
classmethod
¶
from_code(code: str | None) -> JournalType
Interpret an XAF jrnTp (B C G M O P S T Y Z) or a CLAIR2 free-text type.
AccountKind
¶
Bases: StrEnum
Interpretation of a ledger account type.
AccountType
dataclass
¶
AccountType(code: str | None, kind: AccountKind)
A ledger account type: the written code and its interpretation.
from_code
classmethod
¶
from_code(code: str | None) -> AccountType
Interpret accTp (B/M/P) or a CLAIR2/ADF free-text type (heuristic).
RelationKind
¶
Bases: StrEnum
Interpretation of a customer/supplier type (custSupTp).
RelationType
dataclass
¶
RelationType(code: str | None, kind: RelationKind)
A relation type: the written code and its interpretation.
from_code
classmethod
¶
from_code(code: str | None) -> RelationType
Interpret custSupTp (B/C/S; also O, and D/F as written by some exporters/ADF).
RgsRef
dataclass
¶
RgsRef(*, raw: str, code: str | None, extension: str | None, source: str, placeholder: bool = False)
A reference from a ledger account to an RGS code.
Attributes:
| Name | Type | Description |
|---|---|---|
raw |
str
|
The value as written. |
code |
str | None
|
The RGS reference code ( |
extension |
str | None
|
Anything appended after the code (sub-codes, |
source |
str
|
Where the value came from: |
placeholder |
bool
|
Whether the value is a known placeholder ( |
Address
dataclass
¶
Address(*, kind: Literal['street', 'postal'], street: str | None = None, number: str | None = None, number_extension: str | None = None, property: str | None = None, city: str | None = None, postal_code: str | None = None, region: str | None = None, country: str | None = None, raw: RawRecord | None = None)
A street or postal address.
Header
dataclass
¶
Header(*, fiscal_year: str | None, start_date: date | None, end_date: date | None, currency: str | None, created: date | None, software_name: str | None, software_version: str | None, rgs_version: str | None = None, declared_version: str | None = None, raw: RawRecord | None = None)
File header.
fiscal_year is the raw string ("2024" or broken years "2023-2024").
declared_version is CLAIR2's auditfileVersion or ADF's version field.
Company
dataclass
¶
Company(*, name: str | None, identifier: str | None = None, commerce_number: str | None = None, tax_registration_country: str | None = None, tax_registration_id: str | None = None, addresses: tuple[Address, ...] = (), raw: RawRecord | None = None)
The administration (company) the auditfile is about.
LedgerAccount
dataclass
¶
LedgerAccount(*, id: str, description: str | None, account_type: AccountType, lead_code: str | None = None, lead_description: str | None = None, rgs: RgsRef | None = None, seq: int = 0, file: int = 0, raw: RawRecord | None = None)
A general-ledger account.
Relation
dataclass
¶
Relation(*, id: str, name: str | None, relation_type: RelationType, contact: str | None = None, tax_registration_country: str | None = None, tax_registration_id: str | None = None, commerce_number: str | None = None, email: str | None = None, telephone: str | None = None, website: str | None = None, addresses: tuple[Address, ...] = (), opening_balance: Decimal | None = None, closing_balance: Decimal | None = None, seq: int = 0, file: int = 0, raw: RawRecord | None = None)
A customer and/or supplier.
opening_balance/closing_balance (4.0 opBalDesc/clBalDesc, which despite their
names are amounts) are signed, debit-positive.
VatCode
dataclass
¶
VatCode(*, id: str, description: str | None, payable_account_id: str | None = None, receivable_account_id: str | None = None, seq: int = 0, raw: RawRecord | None = None)
A VAT code definition.
Period
dataclass
¶
Period(*, key: str, number: int | None, start_date: date | None = None, end_date: date | None = None, description: str | None = None, seq: int = 0, raw: RawRecord | None = None)
An accounting period.
key is the period number as written (opaque, e.g. "01", "501"); number is
its integer value when numeric.
Journal
dataclass
¶
Journal(*, id: str, description: str | None, journal_type: JournalType, offset_account_id: str | None = None, bank_account: str | None = None, seq: int = 0, file: int = 0, raw: RawRecord | None = None)
A journal (dagboek).
VatLine
dataclass
¶
VatLine(*, code: str | None, percentage: Decimal | None, amount: Decimal | None, side: Side | None, signed_amount: Decimal | None, raw: RawRecord | None = None)
VAT information on a transaction line. signed_amount is debit-positive.
ForeignAmount
dataclass
¶
ForeignAmount(*, currency: str | None, amount: Decimal | None, signed_amount: Decimal | None = None, exchange_rate: Decimal | None = None)
Amount in a foreign currency. exchange_rate is only given by ADF.
Line
dataclass
¶
Line(*, seq: int, number: str | None = None, account_id: str | None = None, amount: Decimal | None, side: Side | None, signed_amount: Decimal | None, debit: Decimal | None, credit: Decimal | None, description: str | None = None, document_ref: str | None = None, effective_date: date | None = None, settlement_date: date | None = None, relation_id: str | None = None, invoice_ref: str | None = None, order_ref: str | None = None, receiving_doc_ref: str | None = None, shipping_doc_ref: str | None = None, cost_center: str | None = None, cost_unit: str | None = None, product: str | None = None, project: str | None = None, work_cost_arrangement: str | None = None, bank_account: str | None = None, offset_bank_account: str | None = None, quantity: Decimal | None = None, vat: tuple[VatLine, ...] = (), foreign: ForeignAmount | None = None, journal_id: str | None = None, transaction_number: str | None = None, transaction_seq: int | None = None, period_key: str | None = None, transaction_date: date | None = None, file: int = 0, raw: RawRecord | None = None)
A transaction line (or opening-balance line), flattened with its transaction context.
Amounts
amount is the value as written (may be negative) and side the written debit/credit
indicator. signed_amount is debit-positive after applying the negative-amount policy;
debit/credit are its non-negative presentation (one of them is zero).
All are None when the written amount is invalid (a finding is reported).
Context
journal_id, transaction_number, transaction_seq, period_key and
transaction_date repeat the enclosing transaction's values.
extra
property
¶
Fields of the raw record without a normalized attribute (version/vendor specific).
Transaction
dataclass
¶
Transaction(*, seq: int, journal_id: str | None, number: str | None, description: str | None, period_key: str | None, period_number: int | None, date: date | None, source: str | None = None, user: str | None = None, lines: tuple[Line, ...] = (), file: int = 0, raw: RawRecord | None = None)
TransactionTotals
dataclass
¶
TransactionTotals(*, lines_count: int | None, total_debit: Decimal | None, total_credit: Decimal | None)
Control totals declared in the file.
linesCount/numberEntries, totalDebit and totalCredit as parsed values;
None where absent or invalid.
OpeningBalance
dataclass
¶
OpeningBalance(*, source: Literal['element', 'transactions', 'none'], lines: tuple[Line, ...], date: date | None = None, description: str | None = None, declared: TransactionTotals | None = None)
Unified opening balance, wherever the file stores it.
Attributes:
| Name | Type | Description |
|---|---|---|
source |
Literal['element', 'transactions', 'none']
|
|
lines |
tuple[Line, ...]
|
The opening-balance lines (signed amounts follow the negative-amount policy). |
date |
date | None
|
|
description |
str | None
|
|
declared |
TransactionTotals | None
|
Control totals as written in the file (element only). |
by_account
¶
Signed (debit-positive) opening balance per account ID.
pyxaf.RawRecord
¶
RawRecord(tag: str, fields: dict[str, str], children: dict[str, list[RawRecord]] | None, line: int, text: str | None = None, sequence: list[list[Any]] | None = None)
One XML element with its leaf children as fields and complex children as children.
Attributes:
| Name | Type | Description |
|---|---|---|
tag |
str
|
Local element name (namespace prefix removed). |
fields |
Mapping[str, str]
|
Text of leaf child elements by local name, exactly as written (entities resolved,
whitespace preserved; empty elements give |
children |
Mapping[str, Sequence[RawRecord]]
|
Complex child elements (and repeated leaves) by local name, in document order. |
line |
int
|
1-based line of the start tag. |
text |
str | None
|
Text content for a leaf element stored as a child, else |
sequence |
list[list[Any]] | None
|
Only during validation: the names of all child elements in document order,
as |
pyxaf.tables
¶
Normalized tables for analysis and export.
Every table has a fixed schema (column names and logical types) that is the same for all
iterations. Tables are produced lazily: master-data tables from memory, transactions,
lines and line_vat by streaming the file. All rows carry synthetic seq keys because
real files contain duplicate numbers.
| Table | Rows |
|---|---|
header |
one row: file header and detected version |
company |
one row |
addresses |
company and customer/supplier addresses |
accounts |
ledger accounts (with RGS reference) |
relations |
customers/suppliers |
vat_codes |
VAT codes |
periods |
periods |
journals |
journals |
transactions |
journal entries (with line counts and debit/credit totals) |
lines |
transaction lines with their transaction context |
line_vat |
VAT details of transaction lines (line_seq → lines.seq) |
opening_balance |
opening-balance lines (from the element or period-0/opening transactions) |
Logical types: string, int64, date, bool and decimal(p,s). Amounts are
decimal(20,2) as in the XSDs, VAT percentages decimal(8,3).
TABLES
module-attribute
¶
TABLES: dict[str, tuple[Column, ...]] = {'header': _cols('format_version family fiscal_year start_date:date end_date:date currency created:date software_name software_version rgs_version declared_version'), 'company': _cols('name identifier commerce_number tax_registration_country tax_registration_id'), 'addresses': _cols('owner_type owner_id seq:int64 kind street number number_extension property city postal_code region country'), 'accounts': _cols('seq:int64 file:int64 id description account_type account_kind lead_code lead_description rgs_raw rgs_code rgs_extension rgs_source rgs_placeholder:bool'), 'relations': _cols(f'seq:int64 file:int64 id name relation_type relation_kind contact tax_registration_country tax_registration_id commerce_number email telephone website opening_balance:{AMOUNT} closing_balance:{AMOUNT}'), 'vat_codes': _cols('seq:int64 id description payable_account_id receivable_account_id'), 'periods': _cols('seq:int64 key number:int64 start_date:date end_date:date description'), 'journals': _cols('seq:int64 file:int64 id description journal_type journal_kind offset_account_id bank_account'), 'transactions': _cols(f'seq:int64 file:int64 journal_id number description period_key period_number:int64 date:date source user line_count:int64 total_debit:{AMOUNT} total_credit:{AMOUNT}'), 'lines': _cols(f'seq:int64 transaction_seq:int64 file:int64 journal_id transaction_number period_key transaction_date:date number account_id amount:{AMOUNT} side signed_amount:{AMOUNT} debit:{AMOUNT} credit:{AMOUNT} description document_ref effective_date:date settlement_date:date relation_id invoice_ref order_ref receiving_doc_ref shipping_doc_ref cost_center cost_unit product project work_cost_arrangement bank_account offset_bank_account quantity:decimal(24,6) foreign_currency foreign_amount:{AMOUNT} foreign_signed_amount:{AMOUNT} exchange_rate:decimal(24,6) vat_count:int64'), 'line_vat': _cols(f'line_seq:int64 transaction_seq:int64 index:int64 code percentage:decimal(8,3) amount:{AMOUNT} side signed_amount:{AMOUNT}'), 'opening_balance': _cols(f'seq:int64 source number account_id amount:{AMOUNT} side signed_amount:{AMOUNT} debit:{AMOUNT} credit:{AMOUNT} transaction_seq:int64 journal_id')}
Column
dataclass
¶
A table column: name and logical type.
Table
¶
Table(tables: Tables, name: str)
One normalized table: schema plus lazily produced rows.
With pyxaf[arrow] installed a table implements the Arrow PyCapsule stream interface
(__arrow_c_stream__), so polars.DataFrame(table), pyarrow.table(table) and
DuckDB can consume it directly.
Tables
¶
Tables(af: AuditFile)
All normalized tables of an AuditFile (see module documentation).
export
¶
export(directory: str | PathLike[str], *, format: ExportFormat = 'csv', tables: Iterable[str] | None = None, batch_size: int = 65536, on_inexact: OnInexact = 'raise') -> list[str]
Write tables to directory (one file per table); return the written paths.
csv and jsonl need no dependencies (decimals are written as exact strings, dates
as ISO 8601); parquet needs pyxaf[parquet]. Streamed tables are produced in a
single pass over the file.
to_polars
¶
to_polars(tables: Iterable[str] | None = None, *, batch_size: int = 65536, on_inexact: OnInexact = 'raise') -> dict[str, Any]
Return polars DataFrames by table name (needs pyxaf[polars]; no pyarrow).
Streamed tables are built in a single pass over the file.
to_pandas
¶
to_pandas(tables: Iterable[str] | None = None, *, on_inexact: OnInexact = 'raise') -> dict[str, Any]
Return pandas DataFrames with Arrow-backed dtypes (needs pyxaf[pandas]).
pyxaf.rgs
¶
RGS (Referentie Grootboekschema) reference data and checks.
pyxaf does not bundle RGS data: the official RGS Excel carries no licence or reuse statement.
Download the Excel release yourself (e.g. RGS 3.8-def.xlsx) and load it with
load_excel(); the workbook is read with a small standard-library .xlsx reader, so no
extra is needed.
Example
import pyxaf, pyxaf.rgs schema = pyxaf.rgs.load_excel("RGS-3.8-def.xlsx") # doctest: +SKIP schema.version, len(schema) # doctest: +SKIP ('3.8', 4979) schema.parent("BIvaKouVvp").code # doctest: +SKIP 'BIvaKou' report = pyxaf.validate("2024.xaf", rgs=schema) # doctest: +SKIP
RGS codes form a strict prefix hierarchy: every code consists of B (balance) or W
(profit and loss) followed by three-character segments, and the parent of a code is its longest
proper prefix that is itself a code. Levels run from 1 (B, W) to 5 (mutations such as
…Beg, …Inv).
XlsxError
¶
Bases: PyxafError
The workbook is not a readable .xlsx file or lacks a required part.
RgsCode
dataclass
¶
RgsCode(*, code: str, level: int | None, description: str | None, short_description: str | None = None, debit_credit: str | None = None, sort_key: str | None = None, reference_number: str | None = None, opposite_code: str | None = None, entities: frozenset[str] = frozenset(), elimination_filters: frozenset[str] = frozenset())
One RGS reference code.
Attributes:
| Name | Type | Description |
|---|---|---|
code |
str
|
The reference code, e.g. |
level |
int | None
|
Hierarchy level 1–5 ( |
description |
str | None
|
Full description ( |
short_description |
str | None
|
Short description ( |
debit_credit |
str | None
|
|
sort_key |
str | None
|
Presentation order key ( |
reference_number |
str | None
|
|
opposite_code |
str | None
|
|
entities |
frozenset[str]
|
Names of the entity-filter columns set for the code ( |
elimination_filters |
frozenset[str]
|
Names of the columns in the "Filters - te vervallen c.q. te elimineren" group that are set: the code can be dropped when the entity does not use that feature. Empty when the workbook has no such group. |
RgsSchema
¶
RgsSchema(codes: Iterable[RgsCode], *, version: str | None = None, rename_map: Mapping[str, str] | None = None, duplicate_codes: Iterable[str] = ())
One RGS release: codes, hierarchy and renames.
Usually created by load_excel().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
codes
|
Iterable[RgsCode]
|
The codes of this release (iteration order is kept as the sheet order). |
required |
version
|
str | None
|
The RGS version, e.g. |
None
|
rename_map
|
Mapping[str, str] | None
|
Old code → code in this version, for codes renamed in earlier releases. |
None
|
duplicate_codes
|
Iterable[str]
|
Codes that occurred more than once in the source (the first wins). |
()
|
Attributes:
| Name | Type | Description |
|---|---|---|
version |
str | None
|
The RGS version ( |
codes |
Mapping[str, RgsCode]
|
Codes by reference code, in sheet order. |
rename_map |
Mapping[str, str]
|
Old code → code in this version. Only codes that no longer exist in this version are included. |
duplicate_codes |
frozenset[str]
|
Codes that occurred more than once in the source sheet. |
resolve
¶
resolve(code: str) -> RgsCode | None
Return the code, following rename_map for codes renamed since.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
code
|
str
|
A reference code (surrounding whitespace is ignored). |
required |
Returns:
| Type | Description |
|---|---|
RgsCode | None
|
The |
parent
¶
parent(code: str) -> RgsCode | None
Return the parent: the longest proper prefix of code that is a code.
Works for codes that are not in the schema too (the nearest existing ancestor).
children
¶
children(code: str) -> tuple[RgsCode, ...]
Return the direct children of code in sheet order.
check_ref
¶
check_ref(ref: RgsRef, findings: FindingCollector, *, line: int | None, account_id: str) -> None
Check one account's RGS reference against this release.
Placeholders and unrecognisable values are skipped (the validator reports them as
XAF8001/XAF8002). Reports XAF8003 for a code that does not exist in this
version, XAF8004 for a code that was renamed (naming the new code) and XAF8006
for a code at a level other than 4 or 5.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ref
|
RgsRef
|
The parsed reference. |
required |
findings
|
FindingCollector
|
Where to report. |
required |
line
|
int | None
|
Line of the ledger account in the file. |
required |
account_id
|
str
|
The ledger account's ID (for the message). |
required |
check_version
¶
check_version(version_text: str, findings: FindingCollector) -> None
Check the auditfile's declared RGS version (XAF 4.0 header/RGSVersion).
The element has no specified format ("3.7", "RGS 3.8", "RGS-1" occur), so it
is parsed leniently with parse_version(). Reports XAF8005 when no version can be
recognised or when it differs from this schema's version.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
version_text
|
str
|
The declared version as written. |
required |
findings
|
FindingCollector
|
Where to report. |
required |
parse_version
¶
Extract an RGS version number from free text, leniently.
Accepts "3.8", "RGS 3.8", "RGS3.8-def", "rgs-3,7", "Versie 3.5.0"; a
trailing .0 is dropped. Returns None when no number is found.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str | None
|
Any text that may contain a version (a header value, sheet or file name). |
required |
Returns:
| Type | Description |
|---|---|
str | None
|
The normalised version ( |
parse_ref
¶
parse_ref(raw: str, source: str = 'RGScode') -> RgsRef
Split a raw RGS reference into code and extension (see pyxaf.models.RgsRef.parse()).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
raw
|
str
|
The value as written in the auditfile. |
required |
source
|
str
|
Where the value came from ( |
'RGScode'
|
validate_refs
¶
Check every ledger account's RGS reference and the declared RGS version.
Reports placeholders (XAF8001), unrecognisable values (XAF8002) and everything
RgsSchema.check_ref() and RgsSchema.check_version() report.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
af
|
AuditFile
|
An open auditfile. |
required |
schema
|
RgsSchema
|
The RGS release to check against. |
required |
Returns:
| Type | Description |
|---|---|
list[Finding]
|
The findings, in account order. |
load_excel
¶
load_excel(path: str | PathLike[str] | IO[bytes], sheet: str | None = None, *, rename_sheet: str | None = None, max_member_size: int = DEFAULT_MAX_MEMBER_SIZE) -> RgsSchema
Load an official RGS Excel release.
The main sheet is detected automatically when sheet is not given: among the sheets with
a Referentiecode column header in their first rows (and only one such column), names
starting with Totaal are preferred, then the sheet with the most rows. A sheet with
several Referentiecode columns (one per release, e.g. RGS3.8-versus-RGS3.7) provides
RgsSchema.rename_map. Header matching ignores case, whitespace and punctuation.
The version is taken from the title cell above the header (RGS3.8), the sheet name, the
file name or the Recap sheet, in that order.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str | PathLike[str] | IO[bytes]
|
Path or binary seekable stream of the |
required |
sheet
|
str | None
|
Name of the sheet with the codes (auto-detected when |
None
|
rename_sheet
|
str | None
|
Name of the version-comparison sheet (auto-detected when |
None
|
max_member_size
|
int
|
Maximum uncompressed size of a single workbook part, in bytes. |
DEFAULT_MAX_MEMBER_SIZE
|
Returns:
| Type | Description |
|---|---|
RgsSchema
|
The loaded |
Raises:
| Type | Description |
|---|---|
XlsxError
|
If the file is not an |
KeyError
|
If |
ForbiddenConstructError
|
If a workbook part contains a DOCTYPE/ENTITY declaration. |
LimitExceededError
|
If a workbook part exceeds |
pyxaf.errors
¶
Exceptions raised by pyxaf.
pyxaf never refuses a file because of data problems — those become findings
(see pyxaf.findings). Exceptions are reserved for inputs that cannot be read at all:
not an auditfile, malformed XML, forbidden constructs, exceeded limits, missing extras.
PyxafError
¶
Bases: Exception
Base class for all pyxaf errors.
NotAnAuditfileError
¶
Bases: PyxafError
The input is not recognisable as any iteration of the Auditfile Financieel.
EncryptedAuditfileError
¶
Bases: NotAnAuditfileError
The input is an encrypted or vendor-compressed auditfile (.xac, .xsc, .XFC, …).
Only the intended recipient (usually the Belastingdienst) can decrypt these.
XmlSyntaxError
¶
Bases: PyxafError
The XML is not well-formed.
Attributes:
| Name | Type | Description |
|---|---|---|
line |
1-based line number of the error ( |
|
column |
0-based column of the error ( |
|
file |
index of the file in a multi-file set (0 for single files). |
ForbiddenConstructError
¶
ForbiddenConstructError(message: str, *, line: int | None = None, column: int | None = None, file: int = 0)
Bases: XmlSyntaxError
The XML contains a construct pyxaf refuses for security reasons (DOCTYPE, ENTITY).
No version of the Auditfile Financieel uses a DTD; refusing them blocks entity-expansion ("billion laughs") and external-entity (XXE) attacks.
CorruptArchiveError
¶
Bases: NotAnAuditfileError
A gzip or zip container is corrupt or truncated.
LimitExceededError
¶
Bases: PyxafError
A configured safety limit (nesting depth, text size, decompressed size) was exceeded.
MissingExtraError
¶
Bases: PyxafError, ImportError
An optional dependency is needed; the message says which extra to install.