Analytics consent: what to verify before promising “no cookie banner”

Analytics consent: what to verify before promising “no cookie banner”

“Cookieless analytics” is often shortened to “consent-free analytics” and then to “no cookie banner”. Those statements are not equivalent. A tool can avoid HTTP cookies while reading or writing information on a device through another mechanism. A product may offer a limited audience-measurement configuration while other modules require a different assessment. And even when analytics fits a strict framework, videos, support widgets, advertising pixels or embedded forms elsewhere on the site may still require consent. The right question is not, “Is the tool cookieless?” It is:Which trackers and processing operations are actually deployed on this site, in this configuration, for which purposes and under which conditions?This is an assessment framework, not legal advice. It must be adapted to the countries, uses and setup involved. Do not confuse three layers 1. Storage or access technology A cookie is one technique. The ePrivacy framework more broadly addresses storing information on a user's terminal or accessing information already stored there, as transposed in national law. Local storage, SDKs, pixels, fingerprinting mechanisms and other terminal access can therefore raise consent questions without a traditional HTTP cookie. Cookieless is a technical characteristic, not a complete legal classification. 2. The ePrivacy tracker regime In France, Article 82 of the Data Protection Act implements the tracker rules. The general principle is prior information and consent for covered operations, with exceptions including operations strictly necessary for a service expressly requested. The CNIL also describes conditions under which certain audience-measurement trackers may fall within an exemption. This is a narrow framework, not a general exemption for all analytics. 3. Personal-data processing under the GDPR Even when a terminal operation does not require ePrivacy consent in a particular configuration, GDPR duties can still apply if personal data are processed. Purposes, legal basis, transparency, minimisation, retention, recipients, transfers, security and rights may still need to be documented. No banner does not mean no processing or no information. The French limited audience-measurement conditions The CNIL states that, to remain strictly necessary for the service and potentially fall within the described exemption, trackers must in particular:be strictly limited to measuring the audience of the site or app; operate exclusively on behalf of the publisher; produce anonymous statistics only; avoid combining the data with other processing; avoid transmitting non-anonymous data to third parties; avoid global tracking across websites or apps.The CNIL also recommends informing users, limiting tracker lifetime, for example to thirteen months without automatic extension, retaining collected information for no more than twenty-five months, and reviewing those periods. Each condition matters. Strictly limited purpose Technical performance, viewed content and navigation problems may fit the described logic. Advertising audiences, CRM enrichment, ad personalisation and cross-service tracking do not share the same purpose. One product interface may offer both. Audit the enabled feature, not only the vendor name. Exclusively for the publisher The provider should not turn the collection into data for its own targeting, profiling or incompatible cross-client measurement. Review the contract, product documentation and subprocessors. A marketing statement is not enough. Anonymous statistics “Anonymous” is a demanding word. Removing a name, truncating an IP address or hashing an identifier does not automatically create anonymity. If a signal still distinguishes or connects a person, use cautious terminology. Ask the vendor to explain transformations and re-identification risk. No global cross-site tracking A shared identifier used to deduplicate people across properties changes the scope. This matters for groups and agencies consolidating audiences. A multi-site dashboard can aggregate indicators without requiring a cross-site person identifier. The checklist before any no-banner promise 1. Inventory the whole site Do not begin and end with analytics. Include:analytics; tag managers; embedded video and maps; support chat; forms; fraud prevention; experimentation; session replay; advertising; social widgets; security and CDN tooling; partner scripts; mobile SDKs where relevant.Run a tracker audit before and after each consent choice, across several pages and journeys. Strict analytics does not neutralise an advertising pixel elsewhere. 2. State real purposes For every component, state what it enables:aggregate audience statistics; campaign analysis; personalisation; advertising; security; interaction recording; support; product experimentation.“Improve the service” is too broad to govern a configuration. 3. Identify terminal operations Document cookies, local storage, session storage, cache identifiers, SDKs, pixels, device characteristics, consent signals and withdrawal. A scanner showing no cookies does not close the assessment. 4. Inspect collected data and transformations The data collection summary should answer:Is the IP address received, used and stored? Is the full URL transmitted? Is the user-agent raw or reduced? Is a visitor identifier created? Is it stable across days or sites? Are UTM parameters retained? Can free-form events contain text? Which data are aggregated? At what point can a record no longer single someone out?An “anonymous mode” that nobody can explain is not evidence. 5. Review vendor use Ask whether the vendor:acts only as a processor for this collection; reuses data for its own purposes; combines data between customers; produces benchmarks from individual-level data; trains another product; sends data to subprocessors; makes international transfers.Benchmarking can sometimes be designed on separated aggregate data. It still needs to be understood. 6. Verify the exact configuration Documentation may say “can be configured to meet the criteria”. That does not mean your default account does. Keep evidence of:configuration export or screenshots; script version; collection parameters; disabled modules; allowed domains; retention; sharing options; verification date; owner.The CNIL tells publishers to request documentation and operating instructions from providers. 7. Review retention Separate tracker or identifier lifetime, raw events, statistics, technical logs, backups and exports. Test automated deletion. A dashboard retention setting may not cover files exported by your team. 8. Inform visitors Even when consent is not required for a strictly framed measurement setup, the CNIL recommends informing users, for example in the privacy notice. Depending on context, explain purpose, relevant data, general operation, duration, provider, recipients, rights, contact and relevant transfers. “We use privacy-friendly analytics” is not enough. 9. Test refusal and withdrawal Where part of the stack relies on consent:covered trackers must not start before the choice; refusal must follow applicable interface requirements; withdrawal must have an effect; the signal must reach all relevant tags; new pages and components must respect the choice.Test behaviour, not only the CMP appearance. 10. Validate and retain the assessment The controller makes the final decision, with DPO or legal support where appropriate. Record:countries; purposes; inventory; criteria reviewed; vendor evidence; configuration; tests; residual risks; date and owners; review triggers.The answer may differ for a French corporate site, an authenticated app and an international property group. Cookieless, consent mode and no banner Cookieless The term can mean no persistent cookie, no cookie in one mode, alternative storage, identifier-free events, server-derived identifiers or simply no advertising cookie. Ask for the technical definition. Consent mode A consent mode communicates user choices to tags and can change their behaviour. Depending on the product and setup, signals may still be sent without advertising cookies. It helps implement a decision. It does not decide whether no-consent collection is legally permitted, and it does not turn advertising into strictly necessary measurement. No banner This statement can only be assessed across the complete site. It may be reasonable when no non-essential component runs before consent and the audience measurement genuinely meets the applicable framework. It is misleading when based only on the absence of an analytics cookie. Claims to avoidabsolute GDPR or legal-compliance claims; blanket consent-exemption claims; claims that cookie-free analytics automatically remove every banner; claims of official CNIL certification; claims of official CNIL approval; “No personal data” “No legal assessment required”The CNIL explicitly states that a solution cannot present itself as certified or approved by the authority merely because of the audience-measurement self-assessment. More accurate wording includes:“cookieless by default”; “designed for minimal collection”; “can be configured for limited audience measurement”; “exemption depends on purposes, configuration and context”; “users remain informed”; “the complete site stack must be audited”.Precision protects credibility as well as compliance. When a banner remains necessary Depending on applicable law and configuration, consent is generally still relevant for:personalised advertising; retargeting; ad-network sharing; cross-site tracking; profile enrichment; some session-replay uses; non-essential personalisation; third-party embeds with non-essential trackers; analytics beyond a limited measurement purpose.The existing guide to session replay and the CNIL consultation explains why detailed behavioural recording should not be treated like aggregate audience statistics. A simple decision process Case A: strictly limited measurement Minimal collection, no cross-site tracking, no vendor reuse, anonymous statistics, controlled retention, information and documentation. Action: assess and document the local framework, then inspect the rest of the site. Case B: enriched analytics after consent The team wants detailed events, advanced attribution or more persistent identifiers. Action: block the relevant capabilities until consent, transmit the choice correctly and document the processing. Case C: mixed stack Minimal measurement runs by default, with extended modules enabled after consent. Action: separate the modes technically, prevent reporting changes from silently expanding collection, and test every transition. Clear separation is more credible than one setting claimed to fit every use. Conclusion A no-banner promise cannot be inferred from “cookieless”. It follows from an assessment of the complete site, purposes, terminal operations and configuration. Before communicating, verify:every component; purposes; terminal access; data and identifiers; vendor use; configuration; retention; transparency; consent behaviour where applicable; the documented decision.The result may be a no-banner strict stack, a consent-based extended stack, or a clearly separated combination. Quality comes from the distinction, not the slogan. FAQ Is cookieless analytics automatically exempt from consent? No. Assess other terminal operations, purposes, data, identifiers and applicable national law. Cookieless is a technical feature, not a legal conclusion. Does the CNIL certify exempt analytics tools? No. The CNIL provides criteria and a self-assessment tool but says providers cannot present that self-assessment as official certification or approval. Can visitors be informed without a banner? Yes, when consent is not required for the relevant collection, information can be provided in a privacy notice or another appropriate location. It must remain clear and accurate. Do UTM tags prevent an exemption? Not automatically, but their use and combination must remain compatible with the limited purpose, minimisation and absence of cross-site tracking. They must never contain personal data. Who decides whether the site can operate without a banner? The controller makes and documents the decision, supported by a DPO or legal adviser where needed. A vendor alone cannot guarantee the answer for every site. SourcesCNIL, Audience-measurement cookies and consent conditions CNIL, What does the law say about cookies and trackers? Directive 2002/58/EC on privacy and electronic communications EDPB, Guidelines 05/2020 on consent EDPB, Guidelines 2/2023 on the technical scope of Article 5(3) ePrivacy

Read article →
Multi-site analytics dashboard: manage 5, 10 or 30 websites without losing clarity

Multi-site analytics dashboard: manage 5, 10 or 30 websites without losing clarity

Tracking one website is a measurement problem. Tracking ten becomes a governance problem. Each team initially creates its own analytics property, event names and dashboard. Months later, the group has ten definitions of “conversion”, three time zones, incompatible campaign taxonomies and accounts with unclear ownership. The missing piece is not another chart. It is a shared structure. A useful multi-site dashboard must support two movements:compare properties on a consistent baseline; drill into each site without erasing its business context.Combining everything creates an abstract average. Separating everything hides the portfolio. The right architecture preserves both levels. Start with a property map List every site and its role before choosing metrics.Property Role Main audience Meaningful conversion OwnerCorporate site Trust Prospects, partners Qualified contact CommunicationsProduct A Acquisition SMBs Demo request Growth AProduct B Acquisition Mid-market Meeting booked Growth BHelp centre Support Customers Self-service resolution SupportBlog Discovery B2B audience Signup or product visit ContentSites with different purposes should not be ranked only by traffic. A help centre can perform well by reducing support demand even when it generates no demos. The map should also record:domains and subdomains; production environment; analytics tool and property ID; time zone; currency where relevant; creation date; business owner; technical owner; access list; collection mode; retention; active, migrating or archived status.This becomes the reference inventory. Define a common measurement contract The multi-site baseline is not a dashboard. It is a compact measurement contract applied to every property. Common dimensions Use shared definitions for:page or path; referrer domain; source, medium and campaign; country or region; device class; date and time zone; primary events; conversion status.Apply one URL parameter policy and one UTM taxonomy. Common events A small library is enough: form_submitted demo_requested signup_completed download_completed outbound_clicked search_usedEvery event needs a definition, trigger, allowed properties, owner, test and version. One name must not represent different actions. Conversely, three names for the same contact request prevent comparison. Common quality rules Document:test-environment filtering; bot handling; internal-domain handling; consent and collection modes; path normalisation; time zone; deployment process; alert thresholds.The data collection summary can hold the shared baseline and property-specific exceptions. Separate three reading levels A sound multi-site system does not put every chart on one page. Level 1: portfolio view This answers management questions:Which sites gain or lose useful traffic? Where are conversions moving? Which property has an anomaly? Which team needs investigation? Which site stopped sending data?Keep it short. A table with one row per property is often more useful than twenty small charts.Site Visits Change Useful conversions Rate Main channel Data statusCorporate 24,500 +6% 132 0.54% Organic OKProduct A 18,100 -4% 284 1.57% Paid search ReviewProduct B 9,600 +12% 96 1.00% Partners OKHelp 41,000 +2% n/a n/a Direct OKFigures are illustrative. Data status matters: a fall means something different when collection broke. Level 2: property view Each site retains its business dashboard:acquisition; landing pages; content; conversions; events; trends; data quality.A SaaS property may track trials, while a help centre tracks unsuccessful searches or support escalation. Level 3: diagnostics Analysts and engineers need:events by version; collection errors; unknown parameters; client/server discrepancies; time-series breaks; unexpected domains; test traffic; ingestion delay.Keep diagnostics out of executive reporting, but do not omit them. Otherwise every anomaly becomes a manual investigation. KPIs that can be compared Visits and page views They show scale but naturally favour larger sites. Always include trend and context. Common meaningful conversions A shared conversion group can include demo requests, qualified contacts, verified signups or confirmed purchases. Preserve the conversion mix too. A total can hide a shift toward lower-value actions. Conversion rate Rates compare different property sizes only when the denominator is identical. Document whether it uses visits, visitors, sessions or landing-page entries. Channel share Organic, paid, email, partner, referral and direct shares reveal dependence. This requires one campaign taxonomy. Collection health Add technical KPIs:time of last received event; event-volume change; rejected-event share; unknown parameters; pages missing path or title; abrupt direct-traffic movement.Data reliability is a governance KPI. What not to add naively Unique visitors The same person can visit several domains. Adding each site's unique visitors counts them more than once. A global identifier for deduplication materially changes collection. The CNIL notes that using the same identifier across several sites for global tracking falls outside the French consent-exemption conditions it describes for certain audience-measurement trackers. A lightweight report can instead use:visits by property; a clearly labelled non-deduplicated reach sum; or an aggregate method that does not require a person-level cross-site identifier.Heterogeneous conversions A brochure download is not automatically equal to a sale. Show a common total and its composition. Simple averages An average of ten conversion rates gives equal weight to a site with 100 visits and one with 100,000. Use a weighted overall rate or show the distribution. Unaligned periods Time zones and campaign calendars can move events between days or weeks. Normalise time before comparison. Compare without punishing small sites Multi-site views easily become rankings. That is rarely helpful. Use four axes:current level; change over time; local target; measurement confidence.A niche site can have low volume, healthy growth and high-value outcomes. A large site can hide paid-channel dependence or broken tracking. Trends and comparison bands are more useful than a podium. Structure access Multi-site operations increase excess-access risk. Define roles:portfolio owner: sees all properties and manages standards; site owner: administers one property; analyst: views and exports as needed; contributor: sees reports without changing collection; agency or partner: access limited to contracted properties; technical support: temporary, logged access when required.Avoid shared accounts. Review access quarterly, remove access at contract end and apply least privilege. Establish a governance cycle Weekly: monitor health Automate simple alerts for no data, abnormal shifts, unknown domains, rejected events and sudden direct-traffic changes. Monthly: discuss decisions Ask property owners:What changed? What action follows? Which hypothesis will be tested?Reporting should not become a reading of numbers. Quarterly: review the contract Check common events, UTM naming, inactive properties, access, retention, vendors, configuration differences and business goals. At launch: use a checklist Before adding a site:assign owners; set time zone; apply the collection baseline; configure filters; test events; verify consent behaviour; add it to the portfolio; document exceptions; create alerts; schedule the first review.Choose an architecture One property per site This is usually clearest for access, retention and configuration. It requires a portfolio layer for comparison. One shared property with a site dimension It can simplify some reports but mixes permissions, configurations and collection risks. One mistake affects the full dataset. One property per site plus a consolidated view This is often the best compromise: operational separation with portfolio aggregation. Some vendors provide roll-up or consolidated views. Verify plan requirements, deduplication method, permissions and exactly which data are combined. The key factor is not only the number of sites. It is their independence across teams, brands, purposes, regions, access and privacy settings. A one-page dashboard model Top stripportfolio visits; useful conversions; weighted overall rate; healthy property count; open anomaly count.Central table One row per property with trend, conversion, channel and status. Acquisition block Channel shares by site. Content block Top landing pages and rising pages, filterable by property. Quality block Missing data, rejected events, access reviews and recent deployments. Each block links to a detailed view. The portfolio dashboard signals; it does not explain everything. Conclusion Multi-site measurement works when governance comes before visualisation. You need:a clear property map; a common measurement contract; documented exceptions; three reading levels; comparable indicators; restricted access; a review cycle; consolidation that does not force person-level cross-site tracking.The best dashboard does not make every site identical. It gives them a shared language while preserving their role. FAQ Should every website have its own analytics property? It is often the clearest way to separate access and configuration. A consolidated view can compare them. A shared property can work when purposes and permissions are genuinely shared. Can unique visitors be added across sites? Not as deduplicated reach. One person can appear in several properties. Label the number as non-deduplicated or use an appropriate aggregate approach without introducing a global identifier by default. How many KPIs belong in the portfolio view? Five to eight well-defined columns are usually enough: volume, trend, conversion, rate, main channel and collection health. Details belong in property views. How should different site goals be handled? Keep a small common baseline and add local indicators. Compare each site with its own target and trend, not only with other sites. How often should access be reviewed? Quarterly review is a reasonable practice, with immediate removal when employees or vendors leave. SourcesCNIL, Audience-measurement cookies and consent conditions Google Analytics, Analytics account structure Google Analytics, Roll-up properties Matomo, Roll-Up Reporting Plausible, Consolidated view Regulation (EU) 2016/679, purpose limitation and data minimisation principles

Which URL parameters should privacy-first analytics filter?

Which URL parameters should privacy-first analytics filter?

A URL can look harmless while carrying far more information than the page path. https://example.com/confirmation? email=alice@example.com& order_id=84721& utm_source=newsletter& session_token=abc123If analytics collects the full URL, those values may appear in events, logs, exports, screenshots and shared reports. The problem often starts before the analytics platform: the application placed excessive information in the address. A privacy-first approach applies two controls:do not put personal or sensitive information in URLs; send analytics only the parameters that have an explicit purpose.Filtering is not a single patch. It is defence in depth. Why query parameters need their own policy The part after ? is the query string. It contains key-value pairs separated by &. Parameters can be used to:attribute a campaign; paginate or sort a list; select language; prefill a form; identify a resource; carry a token; manage an experiment; preserve a search filter.Browsers, servers, CDNs, monitoring systems and third-party scripts can all observe parts of a URL. OWASP notes that sensitive values in query strings can appear in browser history, logs, intermediary systems and sometimes referrer data, even over HTTPS. HTTPS protects transport between endpoints. It does not hide the URL from authorised systems that process it. Prefer an allowlist to an endless blocklist A blocklist names forbidden parameters: email phone token user_idIt fails when somebody introduces customer_email, invitee, auth or another unknown key. An allowlist names the few parameters justified for analytics: utm_source utm_medium utm_campaign utm_contentEverything else is removed before transmission or storage. This is usually more robust for minimal collection. It also reduces report fragmentation: /products/?sort=price, /products/?sort=name and /products/?session=xyz can map to one stable path when those variants do not answer a business question. Some applications genuinely need functional parameters. The method is not “delete everything after ?” but classify each family. A six-category decision framework 1. Approved campaign parameters Examples: utm_source utm_medium utm_campaign utm_contentThey can help read acquisition when naming is controlled and person-level identifiers are forbidden. The guide to UTM tags, referrers and direct traffic explains the taxonomy. Possible decision:collect a short list; normalise case and values; keep campaign dimensions separate from page path; remove visible parameters after capture when that does not break the journey.2. Functional parameters with no analytics value Examples: sort view page theme currencyThey may be required by the interface without belonging in the page report. Keeping them can generate hundreds of rows. Possible decision:exclude them from the analytics page URL; emit a dedicated event only when a product decision depends on the behaviour; keep a reduced category such as filter_applied, not the free-form value.3. Potentially useful content parameters Examples: lang category plan variantBefore approval, ask:Does the value change a decision? Is there a closed set of valid values? Can it contain free text or an identifier?When the answers are safe, transform it into a controlled dimension. Otherwise remove it. 4. Business identifiers Examples: order_id invoice customer ticket workspaceThey can connect a visit to a case, order or account. Even without a name, linkage can make them personal data. Recommended decision:do not send them to general-purpose analytics; measure an aggregate category or status; handle diagnostics in a separate operational system with appropriate access and retention.5. Personal data and free text Examples: email name phone address search messageFree text is especially risky. Internal search terms can include names, medical issues, addresses or confidential phrases. Recommended decision:prevent the value from entering the URL; remove it from analytics payloads; check logs and third-party tools too; measure only a category or the fact that a search occurred, if needed.6. Secrets and tokens Examples: token code jwt signature password_reset inviteThese values must not be captured. They may grant access to an action or resource. Recommended decision:revisit the journey design; use short-lived, limited-use tokens when a URL is technically necessary; prevent logging; remove the parameter from the address promptly; exclude the page from analytics when controls are unreliable.Filter at several layers Layer 1: the application The best protection is not creating an excessive URL. Do not prefill forms with clear-text email addresses in the query string. Do not place customer IDs in marketing links. Do not copy free-form searches into the page title. This reduces exposure across every system, not only analytics. Layer 2: before the analytics request Build a cleaned representation: const current = new URL(window.location.href); const allowed = new Set([ "utm_source", "utm_medium", "utm_campaign", "utm_content", ]);const clean = new URL(current.origin + current.pathname);for (const [key, value] of current.searchParams) { if (allowed.has(key)) { clean.searchParams.set(key, value.toLowerCase().slice(0, 100)); } }const analyticsPage = clean.pathname; const campaign = Object.fromEntries(clean.searchParams);This illustrates the principle, not a universal implementation. You must also handle repeated keys, validate values, cap length, reject free text, test encoding, account for routing and ensure errors never fall back to the raw URL. Sending page path and campaign dimensions as separate fields is often safer. Layer 3: the collector or proxy Server-side validation protects against browser bugs and old scripts. Reject unknown fields, cap values and log only an error code without copying rejected data. This is important when many sites share an endpoint. Layer 4: the analytics platform Some platforms provide redaction or exclusion. GA4 can redact email patterns and administrator-defined query parameters. Matomo can exclude query parameters from page reports. These settings help, but they do not replace earlier controls. A value may cross a tag manager, log or proxy before being hidden in a report. Layer 5: exports Historical exports can retain values filtered later in the platform. Include warehouses, backups, CSV files and BI connectors in the deletion process. Your data collection summary should distinguish received, transformed, stored and exposed data. Normalise pages without losing useful context A content report should usually group variants of the same resource: /products?sort=price&page=1 /products?sort=name&page=1 /products?utm_source=newsletter /products?session=abcThe primary page dimension can remain: /productsUseful context can be separate: campaign_source=newsletter sort_used=trueThis creates readable reports and avoids high-cardinality dimensions. When the query defines the content Some applications use ?article=42 or ?category=security as the resource identifier. Removing it without replacement would merge distinct pages. Options include:migrating to stable paths such as /articles/42; deriving a controlled, non-personal content dimension.Do not preserve the raw identifier automatically. First assess whether it links to a person or case. SEO cleanup and analytics cleanup differ SEO teams may use canonical URLs, redirects, indexing rules and consistent internal links. Analytics teams choose the representation stored in reports. A canonical tag does not stop a script from collecting the full URL. Removing a parameter from analytics does not change search-engine crawling. Document the two decisions separately. A practical test protocol Test 1: parameter corpus Create test URLs with:approved UTM tags; an unknown key; an encoded email; a numeric identifier; a very long value; repeated keys; special characters; a dummy token; free-form search text.Test 2: network observation Inspect the exact browser payload. Search for the dummy sensitive value across every request, not just the primary analytics request. Test 3: logs and storage Check the collector, CDN, application errors and raw data. Absence from the dashboard does not prove the value was never stored. Test 4: reports and exports Inspect page reports, custom dimensions, API output and a representative export. Test 5: failure behaviour Disable a rule, send an unknown key and simulate an invalid payload. The system should fail safely without logging the full URL. Govern the allowlist Maintain a small register:Parameter Status Purpose Allowed values Owner Reviewutm_source Allowed Acquisition Marketing taxonomy Growth Quarterlyutm_medium Allowed Channel Closed list Growth Quarterlylang Derived Content fr, en Product Twice yearlyemail Forbidden None None Engineering Permanenttoken Forbidden Security None Security PermanentEvery new key must answer the same question: which decision justifies collection? For multi-site environments, use one common default allowlist and document every property-specific exception. Conclusion The right filter is not a long list of forbidden words. It is a simple policy:no personal data or secrets in URLs; a short allowlist for useful campaign signals; controlled dimensions for product needs; cleaning before transmission plus server-side validation; tests across network, logs, storage and exports.This improves privacy, security and report clarity at the same time. Measuring fewer URL variants often produces a better view of the pages that matter. FAQ Should analytics remove the entire query string? It is a sound default for the page dimension, but some applications use parameters to define content. Derive a controlled dimension instead of storing the raw URL. Can UTM tags contain an email address? No. Email addresses in URLs can spread across many systems. Use campaign categories, never person-level identifiers. Is GA4 redaction enough? It reduces specific risks inside GA4 but may not cover logs, other tags, proxies or exports. Filter as early as possible and verify every layer. Can a hashed identifier stay in the URL? Hashing does not automatically make data anonymous. If it can distinguish, link or recover a person, it may remain personal data and should not be transmitted without a justified design. How should internal search terms be handled? Avoid sending free text. Measure search usage, a controlled category or aggregate statistics after assessing the need. SourcesOWASP, Information exposure through query strings in URL MDN, URLSearchParams Google Analytics, Data redaction Matomo, Excluding URL query parameters from tracked URLs Regulation (EU) 2016/679, Article 25 and Article 5 principles EDPB, Guidelines on data protection by design and by default

UTM tags, referrers and direct traffic: read acquisition sources correctly

UTM tags, referrers and direct traffic: read acquisition sources correctly

A rise in direct traffic does not necessarily mean more people typed your domain into a browser. A UTM-tagged visit does not prove that one campaign created the demand. A missing referrer does not prove that the visit had no source. These concepts appear in the same acquisition reports but describe different signals:UTM parameters are labels deliberately added to a URL; the referrer is information a browser may transmit; direct traffic is a classification used when the analytics system has no more specific usable source under its rules.Reliable reporting starts with that distinction. It also accepts that web attribution is a reconstruction from incomplete signals, not a complete history of a person's journey. UTM tags are declarations attached to a link A campaign URL might look like this: https://www.example.com/guide/?utm_source=newsletter&utm_medium=email&utm_campaign=launch_juneCommon parameters are:utm_source: the declared origin, such as linkedin, newsletter or a partner; utm_medium: the channel family, such as paid_social, email or referral; utm_campaign: the initiative name; utm_content: a creative, placement or link variant; utm_term: historically used for keywords, and best used only when there is a clear need.Google documents additional manual campaign parameters, but most small teams gain little from more dimensions. Three required fields and one optional variant are usually enough. UTM values are not detected by the browser. A person or system writes them into the link. Treat them as declared campaign metadata, with the strengths and weaknesses of any declared data. What UTM tags do well They help when the referrer is missing, generic or insufficient:newsletters; QR codes; PDF documents; email signatures; organic or paid social posts; partner campaigns; in-app links.They also distinguish two links to the same destination, such as a newsletter hero button and footer link. What they do not prove A UTM tag does not prove that the campaign caused all demand. It says that the measured visit arrived with that label. The URL may have been copied into a private channel, forwarded by a colleague, opened much later or altered by an intermediary. The visitor may have discovered the brand elsewhere first. Reports should therefore describe visits and conversions attributed under the measurement rule, not certain causality. The referrer is conditional browser information When a browser follows a link, it may send the HTTP Referer header to the destination. The historical misspelling remains part of the protocol. What is sent depends on referrer policy, protocol, browser, opening context and the source site's choices. The modern default policy, strict-origin-when-cross-origin, generally sends:the full URL for same-origin navigation; only the origin for HTTPS cross-origin navigation; no referrer when moving from HTTPS to HTTP.A site can apply a stricter policy, an app can open a webview, and redirects or privacy protections can remove the signal. Referrers are useful but never guaranteed. Referrer and UTM can coexist A visit may provide:referrer: linkedin.com; utm_source: linkedin; utm_medium: paid_social; utm_campaign: webinar_june.The analytics platform then applies its own precedence rules. GA4 exposes manual source, medium and campaign dimensions while channel groups follow documented rules that can evolve. Do not compare reports without checking scope. First-user source, session source and key-event attribution answer different questions. Direct means that no better source was assigned In everyday language, “direct” suggests a typed URL or bookmark. Those visits exist, but the channel can also contain visits whose source was lost. Common examples include:untagged links in mobile apps or messaging tools; local documents, PDFs and presentations; redirects that drop parameters; restrictive referrer policies; secure-to-insecure navigation; email campaigns without UTM tags; copied links shared in private channels; analytics deployment errors; URL cleanup before campaign parameters are read.A safer interpretation is:The platform did not assign this visit to a more specific source with the data available.A large direct share is not automatically a problem. It becomes an audit signal when it changes abruptly, concentrates on a campaign landing page, or differs unexpectedly between tools measuring the same scope. Build a controlled UTM taxonomy The main risk is not a missing tag. It is inconsistent naming that fragments reports. 1. Use a closed vocabulary for utm_medium The medium should represent a channel family. Keep a controlled list, for example: email paid_search paid_social organic_social partner affiliate display offlineDo not mix paid-social, paidsocial, cpc_social and social_paid. Platforms may treat case and spelling variants differently, and reports will often show separate rows. 2. Use source for a platform or partner Examples: linkedin google customer_newsletter partner_acme event_parisDo not put the campaign name in the source, or you lose the ability to compare the same source over time. 3. Give campaigns a readable structure A simple convention is: goal_offer_periodExamples: lead_demo_2026q2 launch_product_2026june retention_webinar_2026q3Choose one language, case and separator. Lowercase with underscores is easy to validate. 4. Reserve utm_content for useful variants Examples include:hero_button; footer_link; video_a; creative_02; partner_banner.Never use it for recipient information. 5. Centralise link generation A validated spreadsheet, small internal generator or controlled form removes most variants. Store destination URL, source, medium, campaign, optional content, owner, creation date and status. Never place personal data in UTM tags Query parameters spread across many systems. They may appear in:browser history; web-server and CDN logs; analytics tools; support tools; screenshots; copied links; some referrer data; exports and reports.Do not put an email address, name, phone number, customer ID, token or other person-level identifier in a UTM value. For example: utm_content=customer_12345 utm_campaign=renewal_alice@example.comThese values turn campaign metadata into a personal-data distribution channel. Use a category or creative variant, not a person. Your data collection summary should state which parameters are allowed, retained or removed. Five mistakes that distort reporting Using UTM tags on internal links Internal UTM tags can create new attribution or overwrite prior context depending on the platform. Use an event or internal dimension to compare navigation placements. Tagging everything without a question A tag is unnecessary when the referrer provides enough information and no variant needs to be separated. Campaign metadata should answer a decision, not simply add columns. Changing convention mid-campaign linkedin, LinkedIn and linkedin.com can become three rows. Correct naming at generation time and keep a change log. Cleaning the URL too early Removing visible parameters after capture can produce a cleaner address. Removing them before analytics reads them loses the campaign. Test execution order. Comparing tools without aligning definitions Platforms can differ in session definitions, attribution windows, source lists and precedence rules. A discrepancy is not automatic proof that one tool is broken. Diagnose a rise in direct traffic 1. Locate the change Inspect landing pages, devices, countries and time patterns. A home-page increase differs from a spike on a campaign-only page. 2. Review deployments Look for changes to redirects, routing, CMP behaviour, tag managers, analytics scripts or URL cleanup. 3. Audit live campaign links Open the actual links in emails, ads, profiles, QR codes and documents. Do not rely on the planning sheet. 4. Test the complete journey Follow the link in its real context: app, messenger, embedded browser, PDF or QR code. Inspect the collection request and final report. 5. Accept residual uncertainty Dark social and no-referrer contexts cannot be reconstructed with certainty without more intrusive tracking. Responsible analytics sometimes keeps an unknown bucket rather than manufacturing false precision. A minimal acquisition dashboard For a small B2B team, four views are often enough:visits by source and medium; landing pages by source; meaningful conversions by source; direct and unassigned trends.Add cost and revenue only when definitions and joins are reliable. An apparently precise ROAS built on incomplete identifiers may be less useful than a well-defined cost per qualified request. Review trends over several weeks. Low volumes make daily changes noisy. Conclusion UTM tags, referrers and direct traffic are not three versions of the same field. They are separate mechanisms that complement and sometimes contradict one another. A sound acquisition setup uses:a short, controlled UTM taxonomy; no personal identifiers in URLs; a realistic view of referrer limits; a cautious definition of direct; documented attribution rules; regular checks of the links actually distributed.The goal is not to eliminate all direct traffic. It is to make important campaigns readable without pretending to reconstruct every journey. FAQ What is the difference between utm_source and the referrer? utm_source is deliberately added to a link. The referrer is a signal the browser may send from the previous page. Either, both or neither may be present. Does direct traffic mean people already know the brand? Sometimes, but not exclusively. It also includes visits for which no usable source was assigned, including some apps, documents and untagged campaigns. Which UTM parameters are essential? For most teams, utm_source, utm_medium and utm_campaign are the baseline. Use utm_content for a meaningful variant and add other parameters only for a defined question. Should internal links use UTM tags? Usually not. They can disrupt attribution. Use dedicated events or dimensions for internal navigation. Can UTM parameters be removed after arrival? Yes, once they have been captured correctly. Test execution order and retain the values only according to your collection and retention policy. SourcesGoogle Analytics, Traffic-source dimensions, manual tagging and auto-tagging Google Analytics, Default channel group definitions MDN, Referer header MDN, Referrer-Policy header OWASP, Information exposure through query strings in URL CNIL, The six GDPR principles

Data collection summary: document what your analytics actually collects

Data collection summary: document what your analytics actually collects

Installing analytics can take minutes. Explaining exactly what it collects often takes much longer. The problem is not only volume. Information is scattered across the tracking plan, vendor documentation, consent manager, source code and cloud configuration. When somebody asks, “Do we transmit the full URL?”, “Is the IP address stored?” or “How long do we keep raw events?”, no single person may have a complete answer. A data collection summary is a short operational document that brings those answers together. It describes the collection that is actually deployed, not the collection implied by a marketing page. It is not legal advice, a replacement for a record of processing activities, or a privacy notice. It is the technical layer that helps keep those documents accurate. What a data collection summary is for The document answers one question:For every data point or signal, do we know where it comes from, why it is collected, where it goes, how long it remains and who can access it?Product teams can use it to challenge new events. Marketing teams can see which dimensions genuinely exist. Engineering teams gain a reference for filtering and transformations. A DPO or legal adviser can compare technical reality with compliance records. Management can see the operational debt behind audience measurement. GDPR principles include purpose limitation, data minimisation, transparency and storage limitation. The Regulation also requires information for individuals and, where applicable, records of processing activities. A data collection summary does not create or replace these duties. It makes the underlying facts easier to establish. It is not the record of processing activities The distinction matters. A record of processing activities describes processing at a governance level: purposes, categories of people and data, recipients, transfers, retention and security measures. A data collection summary goes closer to implementation. It may state that:the page URL is stored without its query string; an IP address is used briefly for a technical operation and not retained; a user-agent is reduced to a browser family; only utm_source, utm_medium and utm_campaign are retained; a form event is emitted only after validation; raw events and aggregate reports have different retention periods.A privacy notice translates the relevant facts into language for visitors. It should not become a copy of the technical inventory, but it cannot be accurate without one.Document Main audience Detail level PurposeRecord of processing Internal compliance Processing and categories Govern and demonstrate complianceData collection summary Product and engineering Fields, flows and controls Describe the deployed collectionPrivacy notice Visitors and users Clear public information Explain relevant processingTracking plan Product, marketing and engineering Events and rules Define what should be measuredThese documents complement one another. They should not contradict one another. The ten columns that make the document useful A spreadsheet is enough. The value comes from the columns and the update discipline. 1. Data point or signal Use concrete names: page path, referrer domain, device class, form event, site ID, UTM parameter, derived country or temporary IP address. Avoid broad labels such as “technical data”. They hide design choices. 2. Example value An example removes ambiguity: /pricing/, newsletter or demo_requested. Use synthetic examples, never real personal data. 3. Source State where the signal originates: browser, server, form, CMS, CDN, analytics script or imported system. This reveals indirect collection. A platform may receive a URL or HTTP header before your tracking code transforms it. 4. Operational purpose Connect the field to a decision. “Identify entry pages that lead to a demo request” is more useful than “marketing analysis”. If a field supposedly serves every purpose, the need has probably not been defined well enough. 5. Transformation before storage Document what is removed, truncated, aggregated or derived:stripping unapproved query parameters; normalising paths; reducing the user-agent; deriving coarse geography and discarding the IP address; hashing an identifier, while recognising that hashing is not automatically anonymisation; daily or monthly aggregation.This separates what the system receives from what it keeps. 6. Destination and processors List every relevant destination: collection endpoint, raw storage, aggregate database, BI tool, export, cloud provider and analytics vendor. Record hosting regions and relevant transfers when they are documented. Do not infer a legal location from a cloud region label alone. 7. Retention Separate the layers:technical logs; raw events; pseudonymised records; aggregate statistics; backups; manual exports.A single global period is often misleading. The CNIL notes that retention should follow the purpose and remain limited to what is necessary. For audience-measurement trackers that may fall within the French consent-exemption framework, it recommends a tracker lifetime of thirteen months and a maximum of twenty-five months for collected information. Those benchmarks do not replace an assessment of the actual setup. 8. Access Describe roles rather than only names: administrators, analysts, agency, support or hosting provider. Specify whether access covers aggregate reports, raw events or exports. “Marketing has access” is not enough when a shared account can download the entire dataset. 9. Consent or configuration dependency Keep this factual:collected only after a consent signal; disabled in strict measurement mode; enabled for defined campaigns only; subject to local ePrivacy assessment; used for limited audience measurement, provided every applicable condition is met.Do not write “exempt” without documenting scope, conditions and configuration. 10. Deletion and owner Explain how the field disappears: automated deletion, scheduled job, vendor purge, manual procedure, contract termination or export deletion. Add an internal owner and a last-review date. Without ownership, the summary starts ageing at the next deployment. A minimal SaaS exampleSignal Purpose Before storage Retention AccessPage path Understand content usage Query string removed, path normalised 25 months for reports Product, marketingReferrer domain Understand visit sources Origin only when transmitted 25 months Marketingutm_source Identify a declared campaign Values normalised to a taxonomy 25 months Marketingdemo_requested Measure a B2B conversion No form content transmitted 25 months Product, aggregate sales viewIP address Security and coarse geolocation Used temporarily, not stored in the event Documented technical period Restricted operationsUser-agent Technical distribution Reduced to browser and device categories 25 months ProductThe table proves nothing by itself. It must match the observed network traffic, source code and vendor settings. Start with a real tracker audit and compare the results with your minimal tracking plan. A five-step method Step 1: start with network traffic Open browser developer tools, reload representative pages and inspect requests. Test before and after each consent choice, across several journeys and devices. Record domains, payloads, URL parameters and events. The network panel shows what leaves the browser. It may not expose every server-side transformation, but it gives you a verifiable starting point. Step 2: inspect code and configuration Review the collector, tag manager, CMP rules, environment variables and filters. Generic vendor documentation does not tell you which options your site enabled. Check adjacent capabilities too: session replay, advertising enrichment, CRM connections, user identifiers and exports. Step 3: ask vendors closed questions Request testable answers:Is the full URL received and stored? Can query parameters be removed before storage? Is the IP address logged outside the event dataset? Which backups still contain data after deletion? Does the vendor reuse data for its own purposes? Which subprocessors and transfers apply? Do exports follow the same retention policy?“Privacy-friendly” fills no column. Step 4: reconcile the documents Compare the summary with the processing record, data-processing agreement, privacy notice and consent interface. Contradictions matter more than the writing quality of each document in isolation. A common example is a notice saying that only aggregate statistics are collected while the tag manager sends a user identifier to a third party. Step 5: tie review to change Review the summary whenever you add or change:a tool; an event; a collection domain; an export; a retention period; consent behaviour; a processor; a site or property.A light quarterly review can detect silent drift. Mistakes that make the summary unreliable Copying vendor marketing Vendor documents describe a possible product. Your summary must describe your instance and configuration. Treating pseudonymisation as anonymisation A hashed or rotating identifier may still be personal data if it can distinguish or reconnect a person. Use precise terms and document re-identification risk. Forgetting URLs Full URLs can expose email addresses, order IDs, internal search terms or tokens. Even a minimal analytics tool can receive excessive data when the website places it in the address. Documenting only the dashboard The visible report is only one surface. Logs, raw events, backups, exports and integrations count too. Leaving the document ownerless An accurate but unmaintained inventory can become more dangerous than no inventory because it creates misplaced confidence. Validation checklist Before approval, verify that:every field has a specific purpose; received and stored data are distinguished; URL parameters and free-text fields were audited; retention is defined by layer; recipients and access roles are named; consent and configuration dependencies are explicit; deletion can be tested; the summary matches public and contractual documents; an owner and review date are recorded; every anonymisation claim has technical evidence.Conclusion A data collection summary is not another document for a compliance folder. It is a shared interface between product, marketing, engineering and legal work. Its value comes from precision. A team that knows exactly what it collects can remove unnecessary fields, explain the useful ones, configure tools correctly and answer questions faster. Start with one web property and its ten most important signals. Verify them in the network and code, then expand only when the deployed collection justifies it. FAQ Is a data collection summary legally required? This exact format is not prescribed by the GDPR. It can support documents and processes that are required or necessary, including processing records and transparent information for individuals. Should it be public? Not necessarily. It often contains internal technical details. Relevant information for individuals should be expressed clearly in the privacy notice or another appropriate notice. Should a non-stored IP address be listed? Yes, if it is received or used even briefly. Distinguish receipt, temporary processing, transformation and storage. Can one summary cover multiple sites? Only when their flows and configuration are genuinely identical. For multi-site operations, maintain a common baseline and document property-specific differences. How often should it be reviewed? At every material collection or destination change, plus a periodic review. Quarterly review is a practical operating rhythm for a small team, not a universal legal rule. SourcesRegulation (EU) 2016/679, including Articles 5, 13, 25 and 30 CNIL, Record of processing activities CNIL, Audience-measurement cookies and consent conditions EDPB, Guidelines 4/2019 on data protection by design and by default Chrome for Developers, Network features reference OWASP, Information exposure through query strings in URL