
Supply Chain
Part of Advertising clean rooms
Preparing datasets for a clean room project
Define clean-room table grain, match keys, eligibility, duplicates and output requirements before loading data.
Start with one agreed analysis and its intended output. Define the row grain, the fields needed to match and calculate the result, and whether the selected clean-room product permits them. Clean data makes a result easier to interpret, but does not guarantee a high match rate.
Specify the result first
Suppose the intended report counts purchases among customers eligible to match with campaign exposure records. Define the purchase event, campaign and outcome dates, currency, treatment of refunds, and whether the measure is orders, distinct customers or both. Identify only the fields needed for that report.
Build a field-level schema sheet for the join key, purchase event, event date, campaign, currency and chosen outcome measure. For each field, record its meaning, type and representation approved for the selected product, source owner, permitted use, null treatment and update frequency; state the grain of each table and whether each column is a join column or otherwise.
Joining impression rows to order rows can multiply records when a customer has several of each. Plan the grouping or deduplication before reporting a total.
Check the join key and coverage
The parties need compatible keys at the intended identity level. Record how each key is created, normalised, refreshed and removed, and agree its representation before loading; test that representation on both sides.
Identical column names do not establish that values describe the same person or account. Check formats and missing-key counts on each side using records each owner is authorised to inspect.
Snowflake data offerings use a data offering specification that names each dataset's source object, the columns to include and whether each column is a join column or otherwise. The offering creates a view containing only those columns, and some column categories can be renamed. Check the specified columns and names against the planned schema before registering the offering.
Snowflake offerings are live views of source data, with source policies active. Moving or renaming underlying tables, or changing their access permissions, makes previously registered links unusable, so check source objects and permissions before linking.
AWS Clean Rooms requires prepared data tables before they can be queried; a configured table refers to an existing table and has an analysis rule, with a data access budget as an option. AWS Clean Rooms can query data in its original format from Amazon S3, Amazon Athena or Snowflake, accessing the dataset at query run time.
Keep separate counts for source records, records eligible for this use, records with a usable key and records that match. Agree project-specific coverage checks before loading, and state the denominator of any match rate. A match does not show that media was delivered.
Product requirements differ, so build for the selected route rather than assuming a universal file format. Confirm the chosen product's permitted source, column requirements and field representations before loading.
Key metrics to track during clean room data loading
- Total source records
- Count from original dataset
- Records with usable join key
- After cleaning and normalisation
- Matched records
- Records that successfully joined across datasets
Agree on the handling rules
Before loading or linking data, settle the event time zone and reporting cut-off; duplicate, late, corrected and refunded records; treatment of missing values; approved purposes and recipients; and permitted groupings and outputs. Assign an owner to each rule.
For an Australian project, have the privacy owner document why each identifier is needed, its permitted purpose, recipients and data flow under the Privacy Act 1988 (Cth) and the Australian Privacy Principles (APPs). The OAIC's Guide to Securing Personal Information covers APP 11 security, including technical and organisational measures, and says personal information should be destroyed or de-identified when no longer needed unless an exception applies.
Use OAIC guidance on data breach notification to document who handles escalation; it covers mandatory reporting of serious data breaches.
Rehearse the intended calculation
Before using test rows, have the data owner and privacy owner confirm that they are synthetic or authorised for this purpose. Check a valid match, an unmatched record, repeated events and a missing key against the agreed specification.
Compare the result with the intended calculation and check which fields, groupings and outputs the configured rule permits or suppresses. Record the applicable output rule and test its behaviour; do not assume a numeric threshold. If a required field or breakdown cannot be used, revise the proposed analysis before relying on the result.



