Reading CUF files¶
read() reads the whole file (CUF files are small), checks it on the way and returns a
CufFile. It only raises when the input cannot be
read at all: not CUF-XML (NotCufError), malformed XML (XmlSyntaxError), a DTD
(ForbiddenConstructError) or an exceeded limit (LimitExceededError). Everything else becomes a
finding.
The data model¶
CufFile
├── created, encoding, namespace, raw
├── project Project (PROJECTGEGEVENS)
├── sort_code_schemes SortCodeScheme (SORTEERCODES) → entries: SortCodeEntry
├── estimate Estimate (BEGROTING) → stated totals, children
│ └── children Bundle | Line (BUNDELING | BEGROTINGSREGEL), in document order
│ Bundle → children, quantity_lines, sort_codes, stated, multiplier
│ Line → resources: ResourceLine (MAMO_REGEL), quantity_lines, sort_codes
└── tail Tail | None (STAARTGEGEVENS) → contract_sum, items: TailItem
All names are English; the glossary maps every Dutch element and
attribute to its class and field. The classes are frozen dataclasses. Numbers are
decimal.Decimal, dates datetime.date / datetime.datetime, and every value is None when
the attribute is absent, empty or invalid. The specification's defaults ("empty counts as 0",
"an empty factor counts as 1") are applied when calculating, never in the
model, so the model shows exactly what the file says.
for line in cuf.lines: # every estimate line, in document order
print(line.code, line.description, line.quantity, line.unit, line.material_price)
bundle = cuf.bundles[0]
bundle.children # bundles and lines inside it
bundle.stated.labour # the LOONKOSTEN as written (a control total)
bundle.multiplier # DOORREKEN_HOEVEELHEID, if any
Flat views (cuf.bundles, cuf.lines, cuf.resource_lines, cuf.quantity_lines) list every
node in document order; node.seq is its position. Line.bundle_seq and Bundle.parent_seq
point to the enclosing bundle.
Numbers¶
CUF-XML numbers are XDR numbers: digits with an optional sign, decimal point and exponent
(-123.456E+10), in the range of a double. pycuf reads zero and magnitudes from
2.2250738585072014E-308 to 1.7976931348623157E+308, with at most 1,074 decimals (enough to
write any double exactly). Anything else, such as 1e999 or 0e-1000000000000, is an invalid
value: CUF3018 ("outside the supported range"), None in the model, the text in raw. A
short attribute can therefore never make a calculation or an export enormous.
Navigating the tree¶
cuf.parent(line) # the enclosing bundle (None at the top level)
cuf.ancestors(line) # all enclosing bundles, outermost first
for node in cuf.walk(): # every bundle and line, depth first
...
cuf.sort_codes(line) # {'PLANCODE': '200-10', 'ADMICODE': 'A10.002'}
sort_codes() returns the effective sort codes: CUF-XML says a sort code on a bundle applies to
everything beneath it unless a lower node has its own code for the same scheme. Pass
inherited=False for the node's own codes only.
Text lines (headings and remarks without a quantity or prices) have line.is_text == True.
The raw layer and vendor extensions¶
Every model object keeps its RawElement in .raw: the element with all its attributes
(including vendor extensions) as parsed, its line number and its path. Values are the parsed
text, not the bytes: references such as & are resolved, XML turns line breaks and tabs
inside attribute values into spaces, and comments and processing instructions are not kept.
line.raw.get("BTW") # '' – the raw text
line.raw.line, line.raw.path # 812, '/CUF/BEGROTING[1]/BUNDELING[3]/BEGROTINGSREGEL[7]'
line.extra # {'WERKBESCHRIJVING': '…'} – attributes CUF-XML lacks
cuf.raw.iter("MAMO_REGEL") # walk the raw tree yourself
Attributes that CUF-XML 4.003 does not define are reported once per element and attribute
(CUF3012) and are always available in extra and raw. Elements it does not define are
reported (CUF3011) and kept in the raw tree.
Encodings and namespaces¶
Files declare their encoding in the XML declaration; Ibis writes Windows-1252, others UTF-8.
pycuf honours the declaration, including bytes that Windows-1252 does not define. A file without
a declaration is read as UTF-8; if it is really Windows-1252, pass encoding="cp1252".
The default namespace of CUF files varies by producer (x-schema:CufSchema.xml,
x-schema:Forum CUF-XML-schema4003.xsd, or none) and some are not even valid URIs. pycuf
ignores namespaces and matches elements by name; cuf.namespace and cuf.namespaces record what
was declared.
Lenient parsing¶
By default read() accepts three common deviations and reports each acceptance:
| Option | Accepts | Finding |
|---|---|---|
"decimal-comma" |
12,5 as 12.5 |
CUF7003 |
"dmy-date" |
3-5-2024 and 3-5-2024 13:10:29 |
CUF7002 |
"textual-boolean" |
true and false (exactly so) as 1 and 0 |
CUF7004 |
Pass lenient=() to read strictly; such values are then reported as invalid (CUF3018,
CUF3019, CUF3020) and left None. Negative quantities and prices on estimate and resource
lines are valid (omitted work, proceeds from salvaged material) but reported as CUF7001 (INFO).
Repairs¶
Some exporters write XML that is not well-formed. Two opt-in repairs fix the common cases, and each repair is reported:
bare-ampersand escapes an & that starts no entity (as in SYSTEEMHUIS="Bakker & Spees");
control-chars removes characters XML forbids.
Options of read()¶
| Option | Default | Meaning |
|---|---|---|
policy |
"usage-rules" |
default calculation policy of the file |
lenient |
both | accepted deviations (see above) |
encoding |
as declared | override the declared encoding |
repair |
none | opt-in repairs (see above) |
severity_overrides |
none | change or silence finding codes |
max_findings_per_code, max_findings |
100, 10 000 | finding limits |
max_size, max_depth, max_attribute_size |
256 MiB, 64, 1 Mi | safety limits |