API reference¶
Generated from the docstrings with mkdocstrings. Everything
listed under pycuf.<module> is also importable from the top-level package where noted:
pycuf.read, pycuf.validate, pycuf.CufFile, pycuf.ValidationReport, pycuf.Totals,
pycuf.Policy, the model classes, pycuf.Finding, pycuf.Severity, pycuf.CODES,
pycuf.RawElement and the exceptions. Import pycuf.policy, pycuf.tables and
pycuf.spec for the presets, table schemas and the format catalogue.
| Page section | Contents |
|---|---|
| Reading | read(), CufFile |
| Policies | Policy, presets, resolve_policy() |
| Calculating | Totals |
| Validation | validate(), ValidationReport |
| Findings | Finding, Severity, CODES, CodeInfo |
| Data model | Estimate, Bundle, Line, ResourceLine, Costs, … |
| Raw elements | RawElement |
| Tables | Tables, Table, Column, TABLES |
| Format catalogue | ELEMENTS, ElementSpec, AttributeSpec |
| Exceptions | PycufError and subclasses |
pycuf.read
¶
read(source: SourceLike, *, policy: Policy | PresetName = DEFAULT, lenient: Iterable[Leniency] = ('decimal-comma', 'dmy-date', 'textual-boolean'), encoding: str | None = None, repair: Iterable[str] = (), severity_overrides: Mapping[str, Severity | None] | None = None, max_findings_per_code: int | None = 100, max_findings: int | None = 10000, max_size: int | None = DEFAULT_MAX_SIZE, max_depth: int = DEFAULT_MAX_DEPTH, max_attribute_size: int = DEFAULT_MAX_ATTRIBUTE_SIZE) -> CufFile
Read a CUF-XML file.
The file is read completely and checked on the way; data problems never raise but become
CufFile.findings. Nothing is computed yet: call CufFile.totals() or
CufFile.validate().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
SourceLike
|
Path, bytes or binary file object. |
required |
policy
|
Policy | PresetName
|
The file's default calculation policy, a |
DEFAULT
|
lenient
|
Iterable[Leniency]
|
Deviations to accept, each reported as a finding: |
('decimal-comma', 'dmy-date', 'textual-boolean')
|
encoding
|
str | None
|
Override the declared encoding (e.g. |
None
|
repair
|
Iterable[str]
|
Opt-in fix-ups of malformed XML, each reported as a finding:
|
()
|
severity_overrides
|
Mapping[str, Severity | None] | None
|
Change the severity of finding codes, or silence them with |
None
|
max_findings_per_code
|
int | None
|
Keep at most this many findings per code. |
100
|
max_findings
|
int | None
|
Keep at most this many findings in total. |
10000
|
max_size
|
int | None
|
Refuse inputs larger than this many bytes ( |
DEFAULT_MAX_SIZE
|
max_depth
|
int
|
Maximum XML nesting depth. |
DEFAULT_MAX_DEPTH
|
max_attribute_size
|
int
|
Maximum length of one attribute value. |
DEFAULT_MAX_ATTRIBUTE_SIZE
|
Returns:
| Type | Description |
|---|---|
CufFile
|
A |
Raises:
| Type | Description |
|---|---|
NotCufError
|
The input is not CUF-XML (another XML format, or not XML at all). |
XmlSyntaxError
|
The XML is not well-formed. |
ForbiddenConstructError
|
The XML contains a DOCTYPE or ENTITY declaration. |
LimitExceededError
|
A safety limit was exceeded. |
ValueError
|
An option has an invalid value. |
pycuf.CufFile
¶
CufFile(*, name: str, root: RawElement, encoding: EncodingInfo, namespaces: Mapping[str, str], findings: FindingCollector, policy: Policy, lenient: frozenset[str])
A CUF file that has been read: the typed model plus the findings of reading it.
Use pycuf.read() to create one. The whole file is held in memory (CUF files are
small). Nothing is computed when reading: calculate() and check() apply a
Policy afterwards, so the same file can be evaluated under several policies.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Path of the file, or |
created |
datetime | None
|
|
project |
Project
|
Project data ( |
sort_code_schemes |
tuple[SortCodeScheme, ...]
|
Declared sort-code schemes ( |
estimate |
Estimate
|
The estimate tree ( |
tail |
Tail | None
|
The tail ( |
encoding |
EncodingInfo
|
How the bytes were decoded. |
namespace |
str | None
|
The default namespace as written ( |
namespaces |
Mapping[str, str]
|
All namespace declarations of the root element (prefix → URI). |
raw |
RawElement
|
The root |
bundles |
tuple[Bundle, ...]
|
All bundles in document order ( |
lines |
tuple[Line, ...]
|
All estimate lines in document order ( |
resource_lines |
tuple[ResourceLine, ...]
|
All resource (MAMO) lines in document order. |
quantity_lines |
tuple[QuantityLine, ...]
|
All quantity take-off lines in document order. |
policy
instance-attribute
¶
policy: Policy = policy
The default policy for totals() and validate() (from read()).
findings
property
¶
findings: tuple[Finding, ...]
Findings of reading the file: encoding, XML, structure, values and sort codes.
Totals are not checked when reading; see check() and pycuf.validate().
parent
¶
parent(node: Bundle | Line | ResourceLine) -> Bundle | Line | None
The bundle (or, for a resource line, the estimate line) that contains node.
None for nodes directly under the estimate.
ancestors
¶
ancestors(node: Bundle | Line | ResourceLine) -> tuple[Bundle, ...]
The enclosing bundles of node, outermost first.
walk
¶
Yield every bundle and estimate line in document order (depth first).
scheme
¶
scheme(name: str) -> SortCodeScheme | None
The declared sort-code scheme called name (SORTERING), if any.
sort_codes
¶
sort_codes(node: Bundle | Line | ResourceLine, *, inherited: bool = True) -> dict[str, str | None]
Sort codes of node by scheme name.
CUF-XML says that a sort code on a bundle or line applies to everything beneath it,
unless a lower node has its own code for the same scheme. With inherited=True (the
default) the result contains those effective codes; otherwise only the node's own.
totals
¶
totals(*, policy: Policy | PresetName | None = None) -> Totals
Compute the costs of every node under policy (default: policy).
Results are cached per policy. See Totals and
pycuf.policy.
validate
¶
validate(*, policy: Policy | PresetName | None = None, severity_overrides: Mapping[str, Severity | None] | None = None) -> ValidationReport
Check the file: the findings of reading it plus stated totals against computed ones.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
policy
|
Policy | PresetName | None
|
The policy for computing (default: |
None
|
severity_overrides
|
Mapping[str, Severity | None] | None
|
Change the severity of finding codes, or silence them with
|
None
|
export
¶
export(directory: str | PathLike[str], *, policy: Policy | PresetName | None = None, **kwargs: Any) -> list[str]
Write the normalized tables to directory; see Tables.export().
policy selects the policy of the computed columns (default: policy).
to_polars
¶
to_polars(tables: Iterable[str] | None = None, *, policy: Policy | PresetName | None = None, **kwargs: Any) -> dict[str, Any]
Polars DataFrames by table name; see Tables.to_polars().
to_pandas
¶
to_pandas(tables: Iterable[str] | None = None, *, policy: Policy | PresetName | None = None, **kwargs: Any) -> dict[str, Any]
Pandas DataFrames by table name; see Tables.to_pandas().
pycuf.policy
¶
Calculation policies: how to interpret the CUF-XML calculation rules.
CUF-XML's calculation rules are ambiguous in places, and real exporters and importers read them
differently. A Policy makes every interpretation explicit. It is an immutable value:
pass it to CufFile.totals <pycuf.CufFile.totals>() or pycuf.validate(), give it
to pycuf.read() as the file's default, or derive a variant with Policy.replace()::
import pycuf
from pycuf.policy import ERP
cuf = pycuf.read("begroting.xml")
cuf.totals() # the default policy (the 2006 usage rules)
cuf.totals(policy="schema") # a preset, by name
cuf.totals(policy=ERP.replace(abs_tol=Decimal("0.05")))
| Preset | Meaning |
|---|---|
usage-rules |
The "Gebruikersregels CUF 4003" (2006); the default (USAGE_RULES) |
schema |
The literal reading of the 4.003 schema comments (SCHEMA) |
erp |
Labour resource lines as ERP importers read them (ERP) |
ZeroFactorRule
module-attribute
¶
What a factor written as 0 means: "one" (ignore it) or "zero" (literally 0).
EstimateStyle
module-attribute
¶
Whether DOORREKEN_HOEVEELHEID multiplies: "element" yes, "traditional" no,
"auto" yes when the file contains such multipliers (reported as CUF7005).
LabourRule
module-attribute
¶
How labour resource lines are priced: the specification's formula, or ERP practice.
ResourceBasis
module-attribute
¶
Whether a resource line's HOEVEELHEID is for the whole estimate line ("total") or per
unit of the estimate line ("per-unit").
ResourceRule
module-attribute
¶
When resource (MAMO) lines price their estimate line: only when the line states no prices
itself ("fallback"), never ("ignore") or whenever there are resource lines
("prefer").
Rounding
module-attribute
¶
Rounding mode for comparing computed with stated values (decimal constants work too).
PresetName
module-attribute
¶
Names of the presets in PRESETS.
USAGE_RULES
module-attribute
¶
USAGE_RULES: Final = Policy()
The 2006 usage rules ("Gebruikersregels CUF 4003"): the default.
SCHEMA
module-attribute
¶
SCHEMA: Final = Policy(zero_factor='zero', resources='ignore')
The literal reading of the 4.003 schema comments: a factor of 0 is 0, resource lines are only informative.
ERP
module-attribute
¶
ERP: Final = Policy(labour='erp')
Labour resource lines as ERP importers read them (HOEVEELHEID = hours, PRIJS = rate).
PRESETS
module-attribute
¶
PRESETS: Final[Mapping[PresetName, Policy]] = MappingProxyType({'usage-rules': USAGE_RULES, 'schema': SCHEMA, 'erp': ERP})
The presets by name; anywhere a policy is accepted, its name is accepted too.
Policy
dataclass
¶
Policy(*, zero_factor: ZeroFactorRule = 'one', estimate_style: EstimateStyle = 'auto', labour: LabourRule = 'spec', resources: ResourceRule = 'fallback', resource_basis: ResourceBasis = 'total', abs_tol: Decimal = Decimal('0.01'), rel_tol: Decimal = Decimal(0), rounding: Rounding = 'ROUND_HALF_UP')
How to interpret the CUF-XML calculation rules; immutable, see replace().
Attributes:
| Name | Type | Description |
|---|---|---|
zero_factor |
ZeroFactorRule
|
What a factor written as |
estimate_style |
EstimateStyle
|
Whether a bundle's |
labour |
LabourRule
|
How labour ( |
resources |
ResourceRule
|
When resource (MAMO) lines price their estimate line. The specification
makes the estimate line leading and resource lines informative, but some exporters
only price the resource lines. |
resource_basis |
ResourceBasis
|
What a resource line's quantity refers to. |
abs_tol |
Decimal
|
Absolute tolerance for stated-against-computed checks. |
rel_tol |
Decimal
|
Relative tolerance ( |
rounding |
Rounding
|
Rounding mode used to round a computed value to the precision of the stated value before comparing. |
replace
¶
Return a copy with some fields changed (validated like the constructor).
agrees
¶
Whether a stated value agrees with a computed one under this policy's tolerances.
When stated is written with decimals, computed is first rounded to as many
decimals with rounding. A stated value without decimals (100, 1E+2) is
compared as written: exporters drop trailing zeros, so 100 usually means 100.00.
The result does not depend on your decimal context; non-finite values never agree.
pycuf.Totals
dataclass
¶
Totals(*, policy: Policy, estimate_style: Literal['traditional', 'element'], style_inferred: bool, estimate: Costs, findings: tuple[Finding, ...] = (), _costs: Mapping[Bundle | Line, Costs] = dict(), _multipliers: Mapping[Bundle | Line, Decimal] = dict(), _resources: Mapping[ResourceLine, Costs] = dict(), _from_resources: frozenset[Line] = frozenset())
Computed costs of a CUF file under one policy.
Look up any node with totals[node]: for an estimate line its costs before multipliers,
for a bundle the sum of its direct children (the figure its stated totals should show).
extended() gives a node's contribution to the estimate, with every multiplier applied.
Attributes:
| Name | Type | Description |
|---|---|---|
policy |
Policy
|
The policy used. |
estimate_style |
Literal['traditional', 'element']
|
The resolved estimate style (never |
style_inferred |
bool
|
Whether the style was inferred ( |
estimate |
Costs
|
The computed estimate totals (excluding markups and VAT). |
findings |
tuple[Finding, ...]
|
What applying the policy involved ( |
multiplier
¶
The product of the multipliers applied to node's costs in extended().
For a bundle this includes its own DOORREKEN_HOEVEELHEID; for a line, those of its
enclosing bundles. Always 1 with estimate_style="traditional".
extended
¶
node's contribution to the estimate: its costs times multiplier().
resource
¶
resource(resource: ResourceLine) -> Costs
Costs of one resource line (informative; see the policy's resources rule).
pycuf.validate
¶
validate(source: SourceLike, *, policy: Policy | PresetName = DEFAULT, lenient: Iterable[Leniency] = ('decimal-comma', 'dmy-date', 'textual-boolean'), encoding: str | None = None, repair: Iterable[str] = (), severity_overrides: Mapping[str, Severity | None] | None = None, max_findings_per_code: int | None = 100, max_findings: int | None = 10000, max_size: int | None = DEFAULT_MAX_SIZE, max_depth: int = DEFAULT_MAX_DEPTH, max_attribute_size: int = DEFAULT_MAX_ATTRIBUTE_SIZE) -> ValidationReport
Read and validate a CUF file: read(source, ...).validate(), plus fatal problems.
Malformed XML, forbidden constructs and exceeded limits do not raise here: they become an
ERROR finding (CUF2001–CUF2003) in an otherwise empty report. Input that is not
CUF-XML at all still raises NotCufError.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
SourceLike
|
Path, bytes or binary file object. |
required |
policy
|
Policy | PresetName
|
Calculation policy or preset name (see |
DEFAULT
|
lenient
|
Iterable[Leniency]
|
Deviations to accept (see |
('decimal-comma', 'dmy-date', 'textual-boolean')
|
encoding
|
str | None
|
Override the declared encoding. |
None
|
repair
|
Iterable[str]
|
Opt-in repairs (see |
()
|
severity_overrides
|
Mapping[str, Severity | None] | None
|
Change severities per code, or silence codes with |
None
|
max_findings_per_code
|
int | None
|
Keep at most this many findings per code. |
100
|
max_findings
|
int | None
|
Keep at most this many findings in total. |
10000
|
max_size
|
int | None
|
Refuse inputs larger than this many bytes. |
DEFAULT_MAX_SIZE
|
max_depth
|
int
|
Maximum XML nesting depth. |
DEFAULT_MAX_DEPTH
|
max_attribute_size
|
int
|
Maximum length of one attribute value. |
DEFAULT_MAX_ATTRIBUTE_SIZE
|
pycuf.ValidationReport
dataclass
¶
ValidationReport(*, name: str, findings: tuple[Finding, ...], counts: Mapping[str, int], suppressed: int, policy: Policy | None, estimate_style: Literal['traditional', 'element'] | None, stats: Mapping[str, int], severity_counts: Mapping[str, int] = dict())
Result of validate() or pycuf.CufFile.validate().
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The file that was validated. |
findings |
tuple[Finding, ...]
|
Findings, most severe first, then in file order. |
counts |
Mapping[str, int]
|
Occurrences per code, including findings suppressed by limits. |
suppressed |
int
|
Number of findings dropped because of limits. |
severity_counts |
Mapping[str, int]
|
Occurrences per severity name ( |
policy |
Policy | None
|
The calculation policy used ( |
estimate_style |
Literal['traditional', 'element'] | None
|
The resolved estimate style ( |
stats |
Mapping[str, int]
|
Counts of what the file contains. |
pycuf.findings
¶
Graded, coded findings.
Every problem pycuf notices in a file is reported as a Finding with a stable code
(CUF<nnnn>), a Severity and the location where it occurred. Codes are grouped:
1xxx |
input and character encoding |
2xxx |
XML well-formedness and security |
3xxx |
structure and values (against CUF-XML 4.003) |
4xxx |
sort codes (SORTEERCODES / SORTEERCODE) |
5xxx |
stated totals against computed totals |
7xxx |
data quality, conventions and lenient parsing |
CODES documents every code. Codes are never renumbered or reused.
Severity
¶
CodeInfo
dataclass
¶
CodeInfo(code: str, severity: Severity, title: str)
Documentation of a finding code.
Finding
dataclass
¶
Finding(*, code: str, severity: Severity, message: str, line: int | None = None, column: int | None = None, path: str | None = None, value: str | None = None)
One observation about a CUF file.
Attributes:
| Name | Type | Description |
|---|---|---|
code |
str
|
Stable code, e.g. |
severity |
Severity
|
Severity after overrides. |
message |
str
|
Human-readable English message. |
line |
int | None
|
1-based line number of the element ( |
column |
int | None
|
0-based column number. |
path |
str | None
|
Element path, e.g. |
value |
str | None
|
The offending raw value, if any. |
FindingCollector
dataclass
¶
FindingCollector(max_per_code: int | None = 100, max_total: int | None = 10000, overrides: Mapping[str, Severity | None] = dict(), _items: list[Finding] = list(), _counts: Counter[str] = Counter(), _keys: set[tuple[object, ...]] = set())
Collects findings with per-code and overall limits; counts what it suppresses.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
max_per_code
|
int | None
|
Keep at most this many findings per code ( |
100
|
max_total
|
int | None
|
Keep at most this many findings overall ( |
10000
|
overrides
|
Mapping[str, Severity | None]
|
Severity per code, replacing the default severity. Mapping a code to
|
dict()
|
counts
property
¶
Number of occurrences per code, including suppressed ones.
add
¶
add(code: str, message: str, *, line: int | None = None, column: int | None = None, path: str | None = None, value: str | None = None, severity: Severity | None = None) -> None
Record a finding (counted even when suppressed by a limit).
pycuf.models
¶
The typed, English-named data model of a CUF file.
All classes are frozen, slotted, keyword-only dataclasses. Every object keeps the
RawElement it was built from in raw (the parsed text of every attribute),
and extra exposes attributes that CUF-XML 4.003 does not define (vendor extensions). The Dutch
name of every attribute is listed in the glossary of the documentation and in
pycuf.spec.ELEMENTS.
Numbers are decimal.Decimal and None when the attribute is absent, empty or invalid
(an invalid value is reported as a finding; the raw text stays in raw). The spec's defaults
("empty counts as 0", "an empty factor counts as 1") are applied when computing, never here, so
the model shows what the file says.
The tree nodes (Bundle, Line, ResourceLine, QuantityLine)
compare by identity, so they can be used as dictionary keys (for example to look up computed
costs).
COST_FIELDS
module-attribute
¶
COST_FIELDS: tuple[str, ...] = ('hours', 'labour', 'material', 'equipment', 'subcontracting', 'other')
The fields of Costs, in CUF order (UREN, LOONKOSTEN, …, OVERIGE_KOSTEN).
CostType
¶
Bases: StrEnum
The five cost types of CUF-XML (KOSTENSOORT); the value is the Dutch code.
Costs
dataclass
¶
Costs(*, hours: Decimal = _ZERO, labour: Decimal = _ZERO, material: Decimal = _ZERO, equipment: Decimal = _ZERO, subcontracting: Decimal = _ZERO, other: Decimal = _ZERO)
A cost breakdown: labour hours plus the amounts of the five cost types.
Supports +, multiplication by a number and get(); total is the sum of the
five amounts (hours are not money). The arithmetic runs in pycuf's own decimal context, so
your context does not change the results.
StatedTotals
dataclass
¶
StatedTotals(*, hours: Decimal | None = None, labour: Decimal | None = None, material: Decimal | None = None, equipment: Decimal | None = None, subcontracting: Decimal | None = None, other: Decimal | None = None)
Totals as written on a bundle or the estimate (UREN … OVERIGE_KOSTEN).
A field is None when the attribute is absent, empty or invalid. These are control totals:
CUF-XML says the estimate lines are leading, and real files often carry wrong or zero totals,
so compare them with computed totals instead of trusting them.
items
¶
Yield (name, value) for hours and the five cost types.
SortCode
dataclass
¶
SortCode(*, scheme: str | None, value: str | None, raw: RawElement | None = None)
A sort code assigned to a node (SORTEERCODE).
Attributes:
| Name | Type | Description |
|---|---|---|
scheme |
str | None
|
Name of the sort-code scheme ( |
value |
str | None
|
The code ( |
SortCodeEntry
dataclass
¶
SortCodeEntry(*, code: str | None, description: str | None = None, unit: str | None = None, reference_quantity: Decimal | None = None, raw: RawElement | None = None)
One code in a sort-code scheme's table (SORTEERCODE_REGEL).
SortCodeScheme
dataclass
¶
SortCodeScheme(*, name: str | None, purpose: str | None = None, entries: tuple[SortCodeEntry, ...] = (), raw: RawElement | None = None)
A sort-code scheme declared in the file (SORTEERCODES).
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str | None
|
Name of the scheme ( |
purpose |
str | None
|
What the scheme is for ( |
entries |
tuple[SortCodeEntry, ...]
|
The scheme's code table (often empty). |
Project
dataclass
¶
Project(*, cuf_version: str | None = None, software_house: str | None = None, estimate_date: date | None = None, number: str | None = None, name: str | None = None, estimator: str | None = None, client: str | None = None, address: str | None = None, start_date: date | None = None, currency: str | None = None, euro_rate: Decimal | None = None, notes: str | None = None, raw: RawElement | None = None)
General project data (PROJECTGEGEVENS).
extra
property
¶
Attributes not defined in CUF-XML 4.003 (vendor extensions).
QuantityLine
dataclass
¶
QuantityLine(*, seq: int, room: str | None = None, description: str | None = None, count: Decimal | None = None, length: Decimal | None = None, width: Decimal | None = None, height: Decimal | None = None, factor_1: Decimal | None = None, factor_2: Decimal | None = None, raw: RawElement | None = None)
A quantity take-off line (HOEVEELHEDENSTAAT_REGEL); informative only.
product
property
¶
Product of the numbers that are filled in (None if none is).
Take-off software usually multiplies count, dimensions and factors; the result is informative and does not change any quantity.
ResourceLine
dataclass
¶
ResourceLine(*, seq: int, line_seq: int, cost_type: CostType | None, cost_type_code: str | None = None, code: str | None = None, description: str | None = None, unit: str | None = None, quantity: Decimal | None = None, count_unit: str | None = None, count: Decimal | None = None, duration_unit: str | None = None, duration: Decimal | None = None, production_unit: str | None = None, production: Decimal | None = None, price: Decimal | None = None, price_factor: Decimal | None = None, factor_code: str | None = None, hours_per_unit: Decimal | None = None, hourly_rate: Decimal | None = None, hourly_rate_code: str | None = None, ean_code: str | None = None, article_group: str | None = None, order_unit: str | None = None, quantity_per_order_unit: Decimal | None = None, supplier_code: str | None = None, comment: str | None = None, quantity_lines: tuple[QuantityLine, ...] = (), sort_codes: tuple[SortCode, ...] = (), raw: RawElement | None = None)
A resource line (MAMO_REGEL): one cost type of an estimate line, broken down.
Resource lines are informative: when they disagree with their estimate line, the estimate line is leading.
Attributes:
| Name | Type | Description |
|---|---|---|
cost_type |
CostType | None
|
The interpreted cost type, |
cost_type_code |
str | None
|
|
extra
property
¶
Attributes not defined in CUF-XML 4.003 (vendor extensions).
Line
dataclass
¶
Line(*, seq: int, bundle_seq: int | None = None, depth: int = 0, code: str | None = None, description: str | None = None, unit: str | None = None, quantity: Decimal | None = None, count_unit: str | None = None, count: Decimal | None = None, duration_unit: str | None = None, duration: Decimal | None = None, production_unit: str | None = None, production: Decimal | None = None, quantity_factor: Decimal | None = None, factor_code: str | None = None, hours_per_unit: Decimal | None = None, hourly_rate: Decimal | None = None, hourly_rate_code: str | None = None, material_price: Decimal | None = None, equipment_price: Decimal | None = None, subcontract_price: Decimal | None = None, other_price: Decimal | None = None, provisional_sum: bool | None = None, vat_rate: Decimal | None = None, ean_code: str | None = None, article_group: str | None = None, order_unit: str | None = None, quantity_per_order_unit: Decimal | None = None, supplier_code: str | None = None, comment: str | None = None, resources: tuple[ResourceLine, ...] = (), quantity_lines: tuple[QuantityLine, ...] = (), sort_codes: tuple[SortCode, ...] = (), raw: RawElement | None = None)
An estimate line (BEGROTINGSREGEL): a quantity priced per cost type.
Attributes:
| Name | Type | Description |
|---|---|---|
seq |
int
|
Position among all estimate lines of the file (0-based, document order). |
bundle_seq |
int | None
|
|
depth |
int
|
Number of enclosing bundles. |
provisional_sum |
bool | None
|
|
vat_rate |
Decimal | None
|
|
priced
property
¶
Whether the line itself states hours or a unit price for any cost type.
is_text
property
¶
Whether this is a text line: no quantity, no prices and no resource lines.
Exporters use such lines for headings and remarks (IBIS marks them with an stk
sort code).
extra
property
¶
Attributes not defined in CUF-XML 4.003 (vendor extensions).
Bundle
dataclass
¶
Bundle(*, seq: int, parent_seq: int | None = None, depth: int = 1, code: str | None = None, coding_method: str | None = None, description: str | None = None, unit: str | None = None, reference_quantity: Decimal | None = None, multiplier: Decimal | None = None, stated: StatedTotals = StatedTotals(), comment: str | None = None, children: tuple[Bundle | Line, ...] = (), quantity_lines: tuple[QuantityLine, ...] = (), sort_codes: tuple[SortCode, ...] = (), raw: RawElement | None = None)
A bundle (BUNDELING): a chapter, paragraph or element grouping lines and bundles.
Attributes:
| Name | Type | Description |
|---|---|---|
seq |
int
|
Position among all bundles of the file (0-based, document order). |
parent_seq |
int | None
|
|
depth |
int
|
Nesting level (1 = directly under the estimate). |
multiplier |
Decimal | None
|
|
reference_quantity |
Decimal | None
|
|
stated |
StatedTotals
|
The totals as written (excluding the bundle's own multiplier). |
children |
tuple[Bundle | Line, ...]
|
Bundles and estimate lines, in document order. |
Estimate
dataclass
¶
Estimate(*, stated: StatedTotals = StatedTotals(), children: tuple[Bundle | Line, ...] = (), raw: RawElement | None = None)
The estimate (BEGROTING): direct costs, excluding markups and VAT.
Attributes:
| Name | Type | Description |
|---|---|---|
stated |
StatedTotals
|
The totals as written. |
children |
tuple[Bundle | Line, ...]
|
Top-level bundles and estimate lines, in document order. |
TailItem
dataclass
¶
TailItem(*, description: str | None = None, amount: Decimal | None = None, raw: RawElement | None = None)
One tail item (VRIJE_GROOTHEID), e.g. overheads, profit and risk or the VAT amount.
Tail
dataclass
¶
Tail(*, contract_sum: Decimal | None = None, items: tuple[TailItem, ...] = (), raw: RawElement | None = None)
The tail (STAARTGEGEVENS): markups and the contract sum.
Attributes:
| Name | Type | Description |
|---|---|---|
contract_sum |
Decimal | None
|
|
items |
tuple[TailItem, ...]
|
The tail items as written. CUF does not say which of them add up to the contract sum (a VAT amount is often listed too). |
RawElement
¶
RawElement(name: str, attributes: Mapping[str, str], children: Sequence[RawElement], line: int, column: int, path: str, text: str | None = None)
One XML element with its attributes and child elements.
Attributes:
| Name | Type | Description |
|---|---|---|
tag |
str
|
Element name without namespace prefix ( |
name |
str
|
Element name as written, including a prefix if there was one ( |
attributes |
Mapping[str, str]
|
Attribute values by name as written (prefixes kept, |
children |
Sequence[RawElement]
|
Child elements in document order. |
line |
int
|
1-based line of the start tag. |
column |
int
|
0-based column of the start tag. |
path |
str
|
Location such as |
text |
str | None
|
Non-whitespace text content, if any (CUF-XML elements have none). |
get
¶
Return the raw value of attribute name.
iter
¶
iter(tag: str | None = None) -> Iterator[RawElement]
Iterate over this element and all descendants (optionally only those named tag).
to_dict
¶
Return a plain dict: the attributes plus "children" (a list of dicts).
pycuf.tables
¶
Normalized tables for analysis and export.
Every table has a fixed schema (column names and logical types). Computed columns (hours,
labour … total) depend on the calculation policy; they use the file's policy unless another
one is given (Tables(cuf, policy="erp"), cuf.export(dir, policy="erp")).
| Table | Rows |
|---|---|
project |
one row: file and project data, stated and computed estimate totals |
bundles |
bundles (BUNDELING), with stated and computed totals |
lines |
estimate lines (BEGROTINGSREGEL), with computed costs |
resource_lines |
resource lines (MAMO_REGEL), with computed costs |
quantity_lines |
quantity take-off lines (HOEVEELHEDENSTAAT_REGEL) |
sort_codes |
effective sort codes of bundles, lines and resource lines |
sort_code_schemes |
declared sort-code schemes (SORTEERCODES) |
sort_code_entries |
the code tables of the schemes (SORTEERCODE_REGEL) |
tail_items |
tail items (VRIJE_GROOTHEID) |
Rows are joined by synthetic keys, because codes in real files are neither unique nor
mandatory: lines.bundle_seq → bundles.seq, resource_lines.line_seq → lines.seq,
and owner_type/owner_seq for quantity lines and sort codes.
Logical types: string, int64, bool, date, timestamp and number. A
number is a decimal.Decimal in Python and decimal128(28, 15) in Arrow: 13 integer
digits and 15 decimals, which holds every value seen in real CUF files exactly.
OnInexact
module-attribute
¶
What Arrow-based exports do with a number that does not fit decimal128(28, 15) exactly
(more than 15 decimals, or 13 integer digits): raise ValueError (default), write null, or
round half-even to 15 decimals.
Table
¶
Table(tables: Tables, name: str)
One normalized table: its schema plus rows built from the file.
With pycuf[arrow] installed a table implements the Arrow PyCapsule interface
(__arrow_c_stream__, __arrow_c_array__, __arrow_c_schema__), so
polars.DataFrame(table), pyarrow.table(table) and DuckDB consume it directly.
Tables
¶
Tables(cuf: CufFile, policy: Policy | PresetName | None = None)
All normalized tables of a CufFile (see the module documentation).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
cuf
|
CufFile
|
The file. |
required |
policy
|
Policy | PresetName | None
|
Policy for the computed columns (default: the file's policy). |
None
|
totals
instance-attribute
¶
totals: Totals = cuf.totals(policy=policy)
The computed totals behind the computed columns.
export
¶
export(directory: str | PathLike[str], *, format: ExportFormat = 'csv', tables: Iterable[str] | None = None, on_inexact: OnInexact = 'raise') -> list[str]
Write tables to directory (one file per table); return the written paths.
csv and jsonl need no dependencies (numbers are written as exact strings in plain
notation, dates in ISO 8601); parquet needs pycuf[parquet]. A computed number
whose plain notation would take more than 2,000 characters (only absurd inputs produce
one) raises ValueError naming the table, column and row.
pycuf.spec
¶
The CUF-XML 4.003 catalogue: every element and attribute, with types and English names.
This is pycuf's machine-readable description of the format. It drives the structure checks in
pycuf.validate, the mapping from Dutch attribute names to the English attributes of
pycuf.models, and the generated glossary in the documentation. It is written from the
CUF-XML 4.003 schema (Forum Systeemhuizen Bouw, now maintained by Ketenstandaard Bouw en
Techniek) and its usage rules ("Gebruikersregels CUF 4003"), in pycuf's own words.
Data types follow Microsoft XDR, which the 4.003 schema uses: number, date,
datetime, boolean, string and enum (a closed list of values).
ELEMENTS
module-attribute
¶
ELEMENTS: Mapping[str, ElementSpec] = MappingProxyType({e.name: e for e in (_e('CUF', 'CufFile', 'The document: project data, sort-code schemes, the estimate and the tail.', [_a('AANMAAKDATUMTIJD', 'datetime', 'created', 'Date and time the CUF file was written.', required=True)], {'PROJECTGEGEVENS': (1, 1), 'SORTEERCODES': (0, None), 'BEGROTING': (1, 1), 'STAARTGEGEVENS': (1, 1)}, ordered=True), _e('PROJECTGEGEVENS', 'Project', 'General data about the project and the estimate.', [_a('CUF_VERSIE', 'enum', 'cuf_version', 'CUF-XML version of the file; always 4.003 for this version.', required=True, values=frozenset({VERSION})), _a('SYSTEEMHUIS', 'enum', 'software_house', 'Vendor of the software that wrote the file.', values=SOFTWARE_HOUSES), _a('AANMAAKDATUM', 'date', 'estimate_date', 'Date the estimate was created in the estimating software.', required=True), _a('PROJECTNUMMER', 'string', 'number', 'Project number.'), _a('PROJECTNAAM', 'string', 'name', 'Project name.'), _a('CALCULATOR', 'string', 'estimator', 'Person who made the estimate.'), _a('OPDRACHTGEVER', 'string', 'client', 'Client who commissions the work.'), _a('ADRESGEGEVENS', 'string', 'address', 'Location of the building site.'), _a('PROJECTSTARTDATUM', 'date', 'start_date', 'Start date of construction.'), _a('VALUTA', 'enum', 'currency', 'Currency of all amounts (ISO 4217 code).', required=True, values=CURRENCIES), _a('EURO_KOERS', 'number', 'euro_rate', 'Units of the currency that equal one euro.'), _a('VRIJE_TEKST', 'string', 'notes', 'Additional free text.')]), _e('SORTEERCODES', 'SortCodeScheme', 'Declares one sort-code scheme used in the estimate, optionally with its codes.', [_a('SORTERING', 'string', 'name', 'Name of the scheme (PLANCODE, ADMICODE, …).', required=True), _a('FUNCTIE', 'string', 'purpose', 'What the scheme is used for.')], {'SORTEERCODE_REGEL': (0, None)}), _e('SORTEERCODE_REGEL', 'SortCodeEntry', 'One code of a sort-code scheme.', [_a('CODE', 'string', 'code', 'The sort code.'), _a('OMSCHRIJVING', 'string', 'description', 'Description of the code.'), _a('EENHEID', 'string', 'unit', 'Unit of the reference quantity.'), _a('TERUGDEEL_HOEVEELHEID', 'number', 'reference_quantity', 'Quantity of this code in the estimate, for unit rates (informative).')]), _e('SORTEERCODE', 'SortCode', 'Assigns a sort code of one scheme to a bundle, line or resource line.', [_a('SORTERING', 'string', 'scheme', 'Name of the scheme the code belongs to.', required=True), _a('WAARDE', 'string', 'value', 'The sort code.')]), _e('HOEVEELHEDENSTAAT_REGEL', 'QuantityLine', 'A quantity take-off line explaining where a quantity comes from (informative).', [_a('RUIMTE_NAAM', 'string', 'room', 'Room or location of the measured part.'), _a('OMSCHRIJVING', 'string', 'description', 'Description of the measurement.'), _a('AANTAL', 'number', 'count', 'Number of times the part occurs.'), _a('LENGTE', 'number', 'length', 'Length.'), _a('BREEDTE', 'number', 'width', 'Width.'), _a('HOOGTE', 'number', 'height', 'Height.'), _a('FACTOR_1', 'number', 'factor_1', 'Multiplier, e.g. a unit conversion.'), _a('FACTOR_2', 'number', 'factor_2', 'Second multiplier.')]), _e('MAMO_REGEL', 'ResourceLine', 'A resource line (MAMO): breaks an estimate line down into one cost type.', [_a('KOSTENSOORT', 'enum', 'cost_type', 'Cost type: LOON, MATERIAAL, MATERIEEL, ONDERAANNEMING or OVERIG.', required=True, values=COST_TYPES), _a('CODE', 'string', 'code', 'Resource code.'), _a('OMSCHRIJVING', 'string', 'description', 'Description of the resource.'), *_line_quantities(mamo=True), _a('PRIJS', 'number', 'price', 'Unit price (cost types other than LOON).'), _a('PRIJS_FACTOR', 'number', 'price_factor', 'Multiplier on the price or hourly rate (1.10 = 10% surcharge).'), _a('FACTORCODE', 'string', 'factor_code', 'Code of the price factor.'), _a('UUR_NORM', 'number', 'hours_per_unit', 'Labour hours per unit (LOON).'), _a('UUR_TARIEF', 'number', 'hourly_rate', 'Hourly rate (LOON).'), _a('UUR_TARIEFCODE', 'string', 'hourly_rate_code', 'Code of the hourly rate.'), *_article()], {**_QTY_LINE, **_SORT}), _e('BEGROTINGSREGEL', 'Line', 'An estimate line: a quantity with unit prices for up to five cost types.', [_a('CODE', 'string', 'code', 'Line code (does not define the hierarchy).'), _a('OMSCHRIJVING', 'string', 'description', 'Description of the work.'), *_line_quantities(), _a('HOEVEELHEID_FACTOR', 'number', 'quantity_factor', 'Multiplier on the quantity (1.10 = 10% waste); empty counts as 1.'), _a('FACTORCODE', 'string', 'factor_code', 'Code of the quantity factor.'), _a('UUR_NORM', 'number', 'hours_per_unit', 'Labour hours per unit.'), _a('UUR_TARIEF', 'number', 'hourly_rate', 'Hourly rate.'), _a('UUR_TARIEFCODE', 'string', 'hourly_rate_code', 'Code of the hourly rate.'), _a('MATERIAALPRIJS', 'number', 'material_price', 'Material price per unit.'), _a('MATERIEELPRIJS', 'number', 'equipment_price', 'Equipment price per unit.'), _a('ONDERAANNEMINGSPRIJS', 'number', 'subcontract_price', 'Subcontracting price per unit.'), _a('OVERIGE_KOSTEN', 'number', 'other_price', 'Other costs per unit.'), _a('STELPOST', 'boolean', 'provisional_sum', 'Whether the line is a provisional sum (stelpost).'), _a('BTW', 'number', 'vat_rate', 'VAT percentage for the line and everything below it.', required=True), *_article()], {**_QTY_LINE, 'MAMO_REGEL': (0, None), **_SORT}), _e('BUNDELING', 'Bundle', 'A bundle: groups lines and bundles (chapter, paragraph or element).', [_a('CODE', 'string', 'code', 'Bundle code (does not define the hierarchy).'), _a('CODERING_METHODE', 'string', 'coding_method', 'Coding system of the code (NLSFB, Bouwdelen, …).'), _a('OMSCHRIJVING', 'string', 'description', 'Description of the bundle.'), _a('EENHEID', 'string', 'unit', 'Unit of the reference quantity and multiplier.'), _a('TERUGDEEL_HOEVEELHEID', 'number', 'reference_quantity', 'Quantity of the bundle, for a unit rate (informative).'), _a('DOORREKEN_HOEVEELHEID', 'number', 'multiplier', 'Element quantity that multiplies all costs below; empty counts as 1.'), *_totals(required=False), _a('COMMENTAAR', 'string', 'comment', 'Free-text comment.')], {'BUNDELING': (0, None), 'BEGROTINGSREGEL': (0, None), **_QTY_LINE, **_SORT}), _e('BEGROTING', 'Estimate', 'The estimate: direct costs excluding markups and VAT, as a tree of bundles.', _totals(required=True), {'BUNDELING': (0, None), 'BEGROTINGSREGEL': (0, None)}), _e('STAARTGEGEVENS', 'Tail', 'The tail (staart): markups and the contract sum.', [_a('AANNEEMSOM', 'number', 'contract_sum', 'Contract sum including markups, excluding VAT.', required=True)], {'VRIJE_GROOTHEID': (0, None)}), _e('VRIJE_GROOTHEID', 'TailItem', 'One tail item, such as overheads, profit and risk, or the VAT amount.', [_a('OMSCHRIJVING', 'string', 'description', 'Name of the tail item.'), _a('BEDRAG', 'number', 'amount', 'Amount of the tail item.', required=True)]))})
VERSION
module-attribute
¶
The CUF-XML version this catalogue describes (the latest; there is no newer one).
SUPPORTED_VERSIONS
module-attribute
¶
CUF_VERSIE values pycuf reads as CUF-XML (4.000–4.002 differ only in details).
SOFTWARE_HOUSES
module-attribute
¶
SOFTWARE_HOUSES = frozenset(['ADMICOM', 'ARKEY', 'BRINK', 'CTB', 'DUNCAN', 'ENK', 'KOOIJMAN', 'KPD', 'KRAAN', 'NCCWCASA', 'PIRAMIDE', 'PLUSINTEGRATION', 'SCAB', 'SETZ', 'STABIPLAN', 'TWEESNOEKEN', 'VANENBURG', 'VANMEIJEL', 'VMV'])
ElementSpec
dataclass
¶
ElementSpec(name: str, model: str, description: str, attributes: Mapping[str, AttributeSpec], children: Mapping[str, tuple[int, int | None]] = (lambda: MappingProxyType({}))(), ordered: bool = False)
One CUF-XML element.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The Dutch element name ( |
model |
str
|
Name of the class in |
description |
str
|
Short English description. |
attributes |
Mapping[str, AttributeSpec]
|
Attribute specifications by Dutch name, in schema order. |
children |
Mapping[str, tuple[int, int | None]]
|
Allowed child elements: name → |
ordered |
bool
|
Whether children must appear in the listed order (only |
AttributeSpec
dataclass
¶
AttributeSpec(name: str, type: ValueType, field: str, description: str, required: bool = False, values: frozenset[str] | None = None)
One attribute of a CUF-XML element.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The Dutch attribute name as written in files ( |
type |
ValueType
|
XDR data type. |
field |
str
|
Name of the corresponding attribute in |
description |
str
|
Short English description. |
required |
bool
|
Whether CUF-XML 4.003 requires the attribute. |
values |
frozenset[str] | None
|
Allowed values for |
pycuf.errors
¶
Exceptions raised by pycuf.
pycuf never refuses a file because of data problems: those become findings (see
pycuf.findings). Exceptions are reserved for input that cannot be read at all: not CUF-XML,
malformed XML, forbidden constructs, exceeded limits and missing extras.
PycufError
¶
Bases: Exception
Base class for all pycuf errors.
NotCufError
¶
Bases: PycufError
The input is not a CUF-XML file (for example another XML format or the pre-XML CUF 3).
XmlSyntaxError
¶
Bases: PycufError
The XML is not well-formed.
Attributes:
| Name | Type | Description |
|---|---|---|
line |
1-based line number of the error ( |
|
column |
0-based column of the error ( |
ForbiddenConstructError
¶
Bases: XmlSyntaxError
The XML contains a construct pycuf refuses for security reasons (DOCTYPE, ENTITY).
CUF-XML never uses a DTD; refusing them blocks entity-expansion ("billion laughs") and external-entity (XXE) attacks.
LimitExceededError
¶
Bases: PycufError
A configured safety limit (input size, nesting depth, attribute size) was exceeded.
MissingExtraError
¶
Bases: PycufError, ImportError
An optional dependency is needed; the message says which extra to install.