Category: Privacy engineering
All blog posts in this category.

- 21 Sep, 2026
Anonymous, pseudonymised or aggregated data: what the 2026 EDPB guidelines mean for analytics
In analytics products, the words “anonymous”, “anonymised”, “pseudonymised”, “aggregated” and “hashed” are often used too quickly. They create a sense of safety, but they do not always describe the same legal or technical reality. The topic is back in focus in 2026. The European Data Protection Board adopted guidelines on anonymisation, open for public consultation until 30 October 2026. The text aims to clarify the notion of anonymous data and takes recent Court of Justice of the European Union case law into account. For analytics teams, this is a good moment to review both product claims and architecture. Data is not anonymous merely because it no longer contains a name. A truncated IP address, a hashed identifier or an aggregated statistic can reduce risk, but they do not automatically move data outside the scope of personal data. Why this matters for analytics An analytics tool rarely collects a single signal. Even when it is privacy-first, it may process URLs, referrers, campaign parameters, timestamps, countries, devices, events, journeys and rare page views. Each signal may look harmless in isolation. Combined, they may sometimes make a person or behaviour recognisable, especially at low volumes: a page visited by only one person, a very specific segment, a rare conversion, a query containing personal data, or a campaign link with an identifier. The question is not only: “Do we collect direct identifiers?” The question is: could an actor with means reasonably likely to be used still relate the information to a person, and would that person then be identifiable, directly or indirectly? Anonymous, pseudonymised and aggregated are not the same Pseudonymised data replaces or masks some identifiers, but a link with a person can still exist. A hash, stable identifier or separate key can remain personal data if linkage is possible. Aggregated data groups several observations. It often reduces risk, but it does not automatically guarantee anonymity. A statistic on a very small segment can still reveal information about a person or a small group. Anonymous data no longer relates to an identified or identifiable person. That is a demanding threshold. It is not enough to remove visible names. Teams must consider reasonably available means, linkage possibilities and the context of the actor holding or receiving the data. For a blog post, product page or privacy notice, the cautious approach is to reserve “anonymous” for situations that have actually been tested. In other cases, terms such as “minimal collection”, “aggregation”, “no persistent identifier” or “reduced re-identification risk” are often more accurate. Three risks to test Anonymisation methods are usually assessed through three families of risk. 1. Isolation Can a record or unique behaviour be singled out in the dataset? Analytics example: a low-traffic internal page receives a single visit from one country during a precise time window. Even without a name, that observation may be distinctive. 2. Linkage Can several records be linked together or matched with another source? Analytics example: a hashed identifier is used in several exports, or a combination of country, device, referrer and URL makes it possible to follow the same visitor across tables. 3. Inference Can new information be inferred about a person or a small group? Analytics example: a segment such as “visitors to the enterprise pricing page from a specific prospect domain” can reveal commercial intent if volumes are too low. These risks are contextual. They depend on granularity, access to raw data, other available information and the people who can view the reports. What this changes in an analytics setup The first change is vocabulary. Teams should avoid calling data “anonymous” simply because it is cookieless or because no direct identifier is present. Cookieless does not automatically mean anonymous. The second change is architectural. Anonymisation is not assessed only at collection time. It must be considered across the pipeline: collection, preprocessing, storage, aggregation, dashboard display, exports, APIs and deletion. The third change is documentation. If an organisation claims that some data is anonymous, it should be able to explain why. Which columns remain? Which combinations were tested? What thresholds prevent small segments? Who can access raw data? How long is it kept? Example: URLs and campaign parameters URLs are a good example of underestimated risk. A visited page can reveal business context. A URL parameter can contain an email address, CRM identifier, confirmation token, campaign key or customer number. Even in a cookieless tool, storing the full URL without filtering can create accidental collection of personal data. Anonymisation will not necessarily fix the problem if sensitive data is stored in clear form before processing. The better approach is to act upstream:filter or remove risky parameters; retain useful attribution parameters in a controlled way; document filtering rules; avoid unnecessary raw exports; apply display thresholds to low-volume segments.Example: multi-site reporting Multi-site dashboards add another risk. They often consolidate data from several domains, brands, countries or entities. That is useful for governance, but it can also make some behaviours more recognisable when volumes are low. A global manager does not always need every combination of site, page, country, device and hour. A mature dashboard adapts granularity to the role: summary for leadership, operational detail for the team, restricted access to sensitive data. Anonymisation is therefore not only a technical process. It is also a question of access rights and reporting design. Audit checklist To audit analytics claims and architecture, check:is raw data still accessible? are full URLs stored? are risky URL parameters filtered before storage? are stable identifiers used, even if hashed? are small segments hidden or grouped? do exports contain more detail than dashboards? are retention periods justified? are access rights aligned with real need? does marketing use the right vocabulary? can the team explain why a dataset would be considered anonymous?How to use the consultation window The EDPB consultation is open until 30 October 2026. Not every team will submit feedback, but every team can use the deadline as an audit trigger. A useful exercise is to build a data collection summary: data collected, purposes, transformations, granularity levels, retention, access, exports and public wording. The objective is not to prove that everything is anonymous. The objective is to be accurate. A product can be privacy-first without claiming that every dataset is anonymous at every stage. Conclusion The 2026 EDPB guidelines reinforce a simple idea: anonymisation is not a magic word. It is a result to demonstrate in a specific context. For analytics, this clarification is healthy. It pushes teams to collect less, reduce granularity, filter accidental data, control exports and choose more precise wording. Privacy credibility does not come from maximalist vocabulary. It comes from coherent architecture and honest documentation. FAQ Is hashed data anonymous? Not by default. Hashing can be useful, but it does not automatically remove identifiability if the value can be linked back, guessed or combined with other data. Are aggregated analytics reports always anonymous? No. Aggregation reduces risk, but small segments, rare events or unusual combinations can still reveal information about a person or organisation. The relevant question is whether re-identification remains reasonably possible in context. Why does the consultation matter for web analytics? Because analytics vendors and customers often rely on the words anonymous, pseudonymised and aggregated. The EDPB draft gives teams a better checklist for testing those claims. SourcesEDPB - Guidelines 02/2026 on Anonymisation, public consultation EDPB - EDPB sheds light on anonymisation and web scraping for generative AI EDPB - Public consultations

- 01 Jun, 2026
Which URL parameters should privacy-first analytics filter?
A URL can look harmless while carrying far more information than the page path. https://example.com/confirmation? email=alice@example.com& order_id=84721& utm_source=newsletter& session_token=abc123If analytics collects the full URL, those values may appear in events, logs, exports, screenshots and shared reports. The problem often starts before the analytics platform: the application placed excessive information in the address. A privacy-first approach applies two controls:do not put personal or sensitive information in URLs; send analytics only the parameters that have an explicit purpose.Filtering is not a single patch. It is defence in depth. Why query parameters need their own policy The part after ? is the query string. It contains key-value pairs separated by &. Parameters can be used to:attribute a campaign; paginate or sort a list; select language; prefill a form; identify a resource; carry a token; manage an experiment; preserve a search filter.Browsers, servers, CDNs, monitoring systems and third-party scripts can all observe parts of a URL. OWASP notes that sensitive values in query strings can appear in browser history, logs, intermediary systems and sometimes referrer data, even over HTTPS. HTTPS protects transport between endpoints. It does not hide the URL from authorised systems that process it. Prefer an allowlist to an endless blocklist A blocklist names forbidden parameters: email phone token user_idIt fails when somebody introduces customer_email, invitee, auth or another unknown key. An allowlist names the few parameters justified for analytics: utm_source utm_medium utm_campaign utm_contentEverything else is removed before transmission or storage. This is usually more robust for minimal collection. It also reduces report fragmentation: /products/?sort=price, /products/?sort=name and /products/?session=xyz can map to one stable path when those variants do not answer a business question. Some applications genuinely need functional parameters. The method is not “delete everything after ?” but classify each family. A six-category decision framework 1. Approved campaign parameters Examples: utm_source utm_medium utm_campaign utm_contentThey can help read acquisition when naming is controlled and person-level identifiers are forbidden. The guide to UTM tags, referrers and direct traffic explains the taxonomy. Possible decision:collect a short list; normalise case and values; keep campaign dimensions separate from page path; remove visible parameters after capture when that does not break the journey.2. Functional parameters with no analytics value Examples: sort view page theme currencyThey may be required by the interface without belonging in the page report. Keeping them can generate hundreds of rows. Possible decision:exclude them from the analytics page URL; emit a dedicated event only when a product decision depends on the behaviour; keep a reduced category such as filter_applied, not the free-form value.3. Potentially useful content parameters Examples: lang category plan variantBefore approval, ask:Does the value change a decision? Is there a closed set of valid values? Can it contain free text or an identifier?When the answers are safe, transform it into a controlled dimension. Otherwise remove it. 4. Business identifiers Examples: order_id invoice customer ticket workspaceThey can connect a visit to a case, order or account. Even without a name, linkage can make them personal data. Recommended decision:do not send them to general-purpose analytics; measure an aggregate category or status; handle diagnostics in a separate operational system with appropriate access and retention.5. Personal data and free text Examples: email name phone address search messageFree text is especially risky. Internal search terms can include names, medical issues, addresses or confidential phrases. Recommended decision:prevent the value from entering the URL; remove it from analytics payloads; check logs and third-party tools too; measure only a category or the fact that a search occurred, if needed.6. Secrets and tokens Examples: token code jwt signature password_reset inviteThese values must not be captured. They may grant access to an action or resource. Recommended decision:revisit the journey design; use short-lived, limited-use tokens when a URL is technically necessary; prevent logging; remove the parameter from the address promptly; exclude the page from analytics when controls are unreliable.Filter at several layers Layer 1: the application The best protection is not creating an excessive URL. Do not prefill forms with clear-text email addresses in the query string. Do not place customer IDs in marketing links. Do not copy free-form searches into the page title. This reduces exposure across every system, not only analytics. Layer 2: before the analytics request Build a cleaned representation: const current = new URL(window.location.href); const allowed = new Set([ "utm_source", "utm_medium", "utm_campaign", "utm_content", ]);const clean = new URL(current.origin + current.pathname);for (const [key, value] of current.searchParams) { if (allowed.has(key)) { clean.searchParams.set(key, value.toLowerCase().slice(0, 100)); } }const analyticsPage = clean.pathname; const campaign = Object.fromEntries(clean.searchParams);This illustrates the principle, not a universal implementation. You must also handle repeated keys, validate values, cap length, reject free text, test encoding, account for routing and ensure errors never fall back to the raw URL. Sending page path and campaign dimensions as separate fields is often safer. Layer 3: the collector or proxy Server-side validation protects against browser bugs and old scripts. Reject unknown fields, cap values and log only an error code without copying rejected data. This is important when many sites share an endpoint. Layer 4: the analytics platform Some platforms provide redaction or exclusion. GA4 can redact email patterns and administrator-defined query parameters. Matomo can exclude query parameters from page reports. These settings help, but they do not replace earlier controls. A value may cross a tag manager, log or proxy before being hidden in a report. Layer 5: exports Historical exports can retain values filtered later in the platform. Include warehouses, backups, CSV files and BI connectors in the deletion process. Your data collection summary should distinguish received, transformed, stored and exposed data. Normalise pages without losing useful context A content report should usually group variants of the same resource: /products?sort=price&page=1 /products?sort=name&page=1 /products?utm_source=newsletter /products?session=abcThe primary page dimension can remain: /productsUseful context can be separate: campaign_source=newsletter sort_used=trueThis creates readable reports and avoids high-cardinality dimensions. When the query defines the content Some applications use ?article=42 or ?category=security as the resource identifier. Removing it without replacement would merge distinct pages. Options include:migrating to stable paths such as /articles/42; deriving a controlled, non-personal content dimension.Do not preserve the raw identifier automatically. First assess whether it links to a person or case. SEO cleanup and analytics cleanup differ SEO teams may use canonical URLs, redirects, indexing rules and consistent internal links. Analytics teams choose the representation stored in reports. A canonical tag does not stop a script from collecting the full URL. Removing a parameter from analytics does not change search-engine crawling. Document the two decisions separately. A practical test protocol Test 1: parameter corpus Create test URLs with:approved UTM tags; an unknown key; an encoded email; a numeric identifier; a very long value; repeated keys; special characters; a dummy token; free-form search text.Test 2: network observation Inspect the exact browser payload. Search for the dummy sensitive value across every request, not just the primary analytics request. Test 3: logs and storage Check the collector, CDN, application errors and raw data. Absence from the dashboard does not prove the value was never stored. Test 4: reports and exports Inspect page reports, custom dimensions, API output and a representative export. Test 5: failure behaviour Disable a rule, send an unknown key and simulate an invalid payload. The system should fail safely without logging the full URL. Govern the allowlist Maintain a small register:Parameter Status Purpose Allowed values Owner Reviewutm_source Allowed Acquisition Marketing taxonomy Growth Quarterlyutm_medium Allowed Channel Closed list Growth Quarterlylang Derived Content fr, en Product Twice yearlyemail Forbidden None None Engineering Permanenttoken Forbidden Security None Security PermanentEvery new key must answer the same question: which decision justifies collection? For multi-site environments, use one common default allowlist and document every property-specific exception. Conclusion The right filter is not a long list of forbidden words. It is a simple policy:no personal data or secrets in URLs; a short allowlist for useful campaign signals; controlled dimensions for product needs; cleaning before transmission plus server-side validation; tests across network, logs, storage and exports.This improves privacy, security and report clarity at the same time. Measuring fewer URL variants often produces a better view of the pages that matter. FAQ Should analytics remove the entire query string? It is a sound default for the page dimension, but some applications use parameters to define content. Derive a controlled dimension instead of storing the raw URL. Can UTM tags contain an email address? No. Email addresses in URLs can spread across many systems. Use campaign categories, never person-level identifiers. Is GA4 redaction enough? It reduces specific risks inside GA4 but may not cover logs, other tags, proxies or exports. Filter as early as possible and verify every layer. Can a hashed identifier stay in the URL? Hashing does not automatically make data anonymous. If it can distinguish, link or recover a person, it may remain personal data and should not be transmitted without a justified design. How should internal search terms be handled? Avoid sending free text. Measure search usage, a controlled category or aggregate statistics after assessing the need. SourcesOWASP, Information exposure through query strings in URL MDN, URLSearchParams Google Analytics, Data redaction Matomo, Excluding URL query parameters from tracked URLs Regulation (EU) 2016/679, Article 25 and Article 5 principles EDPB, Guidelines on data protection by design and by default