《迈向(更大程度的)“无 Cookie”世界中的消费者监控:当前与未来 Web 追踪机制的对比分析》 原文标题:
内容概要总结
美国 FTC PrivacyCon 2022 论文(马里兰大学 Ido Sivan-Sevilla 与 Patrick T. Parham)。作者提出五类追踪规范类型学,抓取 50 个与 Pubmatic、OpenX、AppNexus、Rubicon 四家 SSP 合作的流行网站,追踪 ID cookie 的持久识别;并对比三种无 cookie 方案(The Trade Desk UID 2.0、LiveRamp RampID、SWAN)。关键发现:SSP 在 77.6%–90% 网站中持久识别用户但不跨 SSP;无 cookie 方案将跨 SSP 识别、依赖 PII,并激励广告主整合一方与三方数据、可能绕过平台定向限制,从而质疑其 GDPR 合规性(第 9 条、序言第 26 条)。
翻译内容
原文内容(English)
迈向(更大程度的)"无 Cookie"世界中的消费者监控:
当前与未来 Web 追踪机制的对比分析
Ido Sivan-Sevilla & Patrick T. Parham(UMD,马里兰大学)
摘要
广告行业预期将弃用第三方 cookie,此举被宣传为"隐私友好"。新的无 cookie 追踪技术正被提出,但这些技术对消费者隐私的影响远未明朗。广告网络会在多大程度上改变其做法,不再依赖跨站消费者监控或历史上丰富的消费者画像来做广告?鉴于针对基于 cookie 的广告机制存在强烈的合规反弹,新的追踪技术能否做到符合 GDPR?
本研究试图评估无 cookie 广告 ID 方案的潜在隐私危害,方法是:(1)构建一种新的追踪规范类型学,以评估追踪技术的隐私影响;(2)对基于 cookie 的追踪机制做演绎分析,通过收集关于持久性用户识别的新数据——样本为 50 个与全部四家主要供给侧平台(SSP)——Pubmatic、OpenX、AppNexus 和 Rubicon——都有合作的流行网站;(3)将这些发现与对来自技术行业文档的数据所做的演绎分析相对照,这些文档涉及三大主要无 cookie ID 架构——The Trade Desk Unified ID 2.0、LiveRamp ID 和 Secure Web Addressability Network(SWAN)。
我们发现,新的追踪架构会使 Web 上的消费者监控在动态上更宽、在时间上更持久,并且更易受到广告主整合一方与三方数据的影响。如果广告主选择基于敏感画像类别对广告出价,这可能导致绕过主要广告平台设定的定向限制。与现有对无 cookie 追踪方案的批评(大多聚焦于同意机制的缺陷或新方案缺乏治理机构)不同,我们的研究强调了广告网络对拟议追踪机制所造成潜在隐私危害的结构性影响。相比基于 cookie 的追踪机制,广告网络的结构与流程可能导致更长久、更宽泛、更丰富的消费者监控。我们的发现也质疑了所提议追踪架构实现 GDPR 合规的能力。尽管公民社会反对当前基于 cookie 的追踪生态,且新追踪方案被营销为"隐私保护广告(privacy-preserving-advertisement)",我们仍展示了这些技术如何实现细粒度的追踪与消费者定向。鉴于广告主如今被激励去构建历史上丰富的消费者画像,"挑出(singling out)"个人对广告主而言可能变得更容易。与关于 Universal ID 方案和 GDPR 合规的现有辩论(大多聚焦于拟议架构中数据控制者和数据处理者角色的缺失)不同,我们强调新方案所允许的画像与定向的程度,以及它们违反 GDPR 关键原则(如第 9 条和序言第 26 条)的方式。
我们的实证发现不仅挑战了"排除第三方 cookie 会减少消费者监控"的假设,还表明消费者监控有扩大的潜力,这严重质疑了这些方案符合 GDPR 的能力。某一种追踪工具的转变——无论该工具多么核心——都无法弥合在线广告市场的结构与激励所导致的消费者隐私固有鸿沟。2
迈向(更大程度的)"无 Cookie"世界中的消费者监控:
当前与未来 Web 追踪机制的对比分析
Ido Sivan-Sevilla & Patrick T. Parham(UMD,马里兰大学)
1 - 引言
基于广告的 Web 内容变现机制即将弃用其主要追踪工具——第三方 cookie(Binns, 2022;Choi et al., 2020;O'Reilly, 2020)。这可以说是一大胜利(对隐私倡导者而言),理论上将禁绝跨站监控,从而增强消费者的在线隐私。广告行业利益相关者广泛讨论新的"无 cookie ID 方案",并将其包装为"隐私保护广告"(PPA)倡议(Thomson & Rescorla, 2021)。然而,这些潜在 ID 方案的隐私与合规影响仍不明确。这些新的"Universal ID 方案"真的像行业所宣传的那样"隐私友好"吗?广告网络会在多大程度上改变其做法,不再依赖跨站消费者监控来做营销?它是否真能阻止个人被未来的在线广告方案"挑出"——如 GDPR 所明确强调的那样?
为更好地理解 Web 追踪与隐私合规的未来,我们将三大主要无 cookie 追踪架构——The Trade Desk Unified ID 2.0、LiveRamp RampID 和 Secure Web Addressability Network(SWAN)——与当前基于 cookie 的追踪机制进行对比,样本是 50 个与全部四家主要供给侧平台(SSP)网络——Pubmatic、OpenX、AppNexus 和 Rubicon——合作的流行网站。我们的对比基于演绎式追踪分析,依据我们提出的追踪规范类型学来评估追踪技术的隐私影响。
我们首先引入一种新的追踪规范类型学,基于五个不同类别,以评估与追踪技术相关的隐私危害。通过这一类型学,我们旨在系统性地衡量并揭示:无 cookie 追踪机制在隐私危害方面彼此之间、以及与当前基于 cookie 追踪相比有多大差异。其次,为刻画当前基于 cookie 的追踪实践,我们从 50 个与业内四家主要 SSP 合作的流行网站收集数据。我们调查这些行为者基于其唯一 ID cookie 在跨站持久识别用户的程度。这些 SSP 有能力跨浏览体验关联用户身份并实施跨站消费者监控。我们抓取排名前百万的流行网站(基于 tranco 排名)根域名下的 ads.txt 文件,以构建围绕同一 SSP 分组的发布者网络。为减少误报,我们检查每个 SSP 网络中 50 个流行网站的 HTML,以核实这些网站确实与该 SSP 合作。随后,我们追踪 ID cookie 在四个 SSP 网络内部及之间的使用情况。我们调查 SSP 对用户身份的构建,并追踪个人跨网站的持久识别。通过多种网络爬虫,我们收集关于经 HTTP 头、URL 参数和浏览器本地存储所安装的追踪 cookie 的数据。我们将其与来自广告行业的技术文档相结合,以进一步 3 理解基于 cookie 追踪的隐私影响。第三,我们将对基于 cookie 追踪的分析,与所提议无 cookie 方案所实现的持久用户识别相对照。我们依靠新闻报道和行业技术文档,考虑业内讨论的三大最主流无 cookie 追踪方案,并基于我们的追踪类型学对其进行演绎评估。
我们的发现揭示了新的追踪架构如何使 Web 上的消费者监控(1)在动态上更宽,不仅在 SSP 网络内部、还跨 SSP 网络关联用户身份;(2)通过依赖确定性识别数据(如 PII)而在时间上更持久;(3)易受一方与三方数据整合用于消费者画像的影响,因为其激励广告主将自己的用户数据与购买的数据关联。这可能允许广告主秘密地基于敏感用户类别出价,从而可能绕过主要广告平台的定向限制。
与现有对无 cookie 追踪方案的批评(大多聚焦于同意机制的缺陷或治理机构的缺失)不同,我们的研究强调了广告网络对拟议追踪机制所造成潜在隐私危害的结构性影响。相比基于 cookie 的追踪,广告网络的结构与流程导致更长久、更宽泛、更丰富的消费者监控。我们的发现也质疑了所提议追踪架构实现 GDPR 合规的能力。尽管公民社会反对当前基于 cookie 的追踪生态(Lomas, 2021),且新追踪方案被营销为"隐私保护广告"(Thomson & Rescorla, 2021),我们仍展示了这些技术如何实现细粒度的追踪与消费者定向。鉴于广告主被激励去构建历史上丰富的消费者画像,"挑出"个人对广告主而言可能变得更容易。与关于 Universal ID 方案和 GDPR 合规的现有辩论(大多聚焦于拟议架构中数据控制者和数据处理者角色的缺失(Schiff, 2021e; Schiff; 2022a),或所提方案中 cookie 的角色(Asim, 2021; Asim, 2022))不同,我们强调新方案所允许的画像与定向的程度,以及它们违反 GDPR 关键原则(如第 9 条和序言第 26 条)的方式。
这些发现凸显了在线追踪如何已成为一种根深蒂固的规范,仅靠某一种追踪工具——第三方 cookie——的转变,无法从根本上改变广告行业行为者的行为和对用户画像的胃口。网站仍将严重依赖第三方请求网络进行广告投放(Gopal et al., 2022),且广告网络的结构预计将保持不变,所有涉及的第三方都有兴趣构建自己关于个人的数据库,以增强其广告出价和投放决策(Martin, 2022)。我们最终主张,在限制数字广告行业持久消费者识别方面的进展尚未取得。跨站监控将继续存在,广告网络的结构与持久用户识别的程度,决定了用户在线隐私的水平,以及广告行业有意义地遵守法律的能力。
为评估未来追踪技术的隐私影响,下一节强调分析未来追踪方案中的现有缺口,以及考虑广告网络的结构性组成与流程之隐私影响的重要性。随后,我们制定五项追踪规范以评估追踪技术的隐私危害。在 3.1 节 4 我们呈现当前基于 cookie 追踪的数据收集与分析。在 3.2 节,我们分析无 cookie 追踪方案的隐私影响。3.3 节总结基于 cookie 与无 cookie 追踪技术的对比。第 4 节讨论我们发现的含义,第 5 节作结,详述研究局限与未来问题。
2 - 广告网络(缺乏)隐私合规
消费者隐私正日益受到重视。关于平台如何使用消费者数据的披露(Mac & Kang, 2021)、联邦与州隐私法案前所未有的浪潮(Lively, 2022),以及广告行业 GDPR 合规持续受到的质疑(Lomas, 2021),都给浏览器厂商施加了压力,使其宣布终止对备受批评的第三方 cookie 的支持(Shein, 2021)。定向广告主要追踪工具的预期转变——该市场预计到 2024 年将增长至 5250 亿美元(Edelman, 2020)——催生了一批替代性用户 ID 方案(Asim, 2021),被行业营销为"隐私友好"。
现有对无 cookie 追踪方案的批评大多聚焦于同意机制的缺陷(Kaye, 2021a)或治理机构的缺失(UnifiedID2, 2022)。Mozilla 最近一份报告揭示了更具体的隐私担忧,指出新追踪方案未提供任何防止访问用户数据的机制(Thomson & Rescorla, 2021)。尽管如此,对这些无 cookie 追踪提案及其隐私影响的讨论,仍未足够重视广告技术网络在继续拼凑持久消费者识别中所将扮演的中介角色(Asim, 2021; Asim, 2022)。相反,我们主张,应仔细分析广告网络对所提议追踪技术潜在隐私危害的结构性影响。
关于所提议追踪技术能否符合 GDPR 的讨论也有限,大多聚焦于拟议架构中数据控制者和数据处理者角色的缺失(Schiff, 2021e; Schiff; 2022a),或所提方案中 cookie 的角色(Asim, 2021; Asim, 2022)。但我们主张,这些方案的核心,是为广告网络行为者所实现的画像与定向能力。应依据 GDPR 关于数据主体对其数据的控制、敏感数据的处理,以及数据控制者"挑出"个人能力的原则,来评估这些能力。
因此,广告网络行为者的结构性组成与流程,应处于消费者追踪与定向评估的中心。令人惊讶的是,这些促成了近一半 Web 变现机制的广告网络(Choi et al., 2020),却极少受到信息系统(IS)文献的实证关注。大多数隐私研究,包括与在线广告相关的研究,都高度针对用户的隐私关切和消费者选择(Aguirre et al., 2015; Dinev, 2014; Dinev et al., 2013; Dinev & Hart, 2006; Goldfarb & Tucker, 2011, 2015; Hui et al., 2007; T. Li & Unger, 2012; Y. Li, 2011, 2012; Malhotra et al., 2004; Xu et al., 2011)。对于消费者在线所经历的隐私威胁、以及他们如何挑战行业合规,我们知之甚少。简言之,关于 Web 隐私"那又如何(so what)"的论述在 IS 文献中仍缺乏实证讨论,本研究旨在填补部分空白。5
具体而言,对基于广告的 Web 变现机制将第三方 cookie 用作规避性追踪工具的转变(Jones, 2020; Lavin, 2006),并未得到批判性分析。我们旨在阐明无 cookie 世界中潜在的用户追踪与定向,质疑后 cookie 方案符合隐私法的能力。我们意在凸显广告网络的运作——它代表一链第三方行为者,监控消费者的在线行为,并为营销目的跨多个网站跟踪消费者(D'Annunzio & Russo, 2020)。虽然发布者收集关于其自身访客的信息,但正是广告网络收集了最多信息,以跨 Web 追踪和定向消费者(Bashir et al., 2016)。这些广告网络影响市场结果,但关于网络行为者决策之影响的研究仍然稀少(Choi et al., 2020)。
为填补这些空白并理解广告网络对消费者隐私的结构性影响,我们基于下文一种新的追踪类型学,对基于 cookie 和无 cookie 的追踪机制做演绎分析,这使我们能够理解追踪技术的隐私影响。
3 - 追踪 Web 追踪
遵循 Binns(2022),我们使用万维网联盟(W3C)追踪保护小组提供的追踪工作定义:即"追踪是指收集关于某一特定用户跨多个不同上下文的活动数据,以及对该活动所衍生数据的保留、使用或共享——且发生在该活动所发生的上下文之外"(W3C Working Group, 2019)。根据这一定义,并非每一次数据收集都被视为"追踪"。只要在一个上下文内收集的数据停留在相同的时空上下文中,我们就不应将该数据收集视为"追踪"。然而,当某一方以某种方式从不同数据源和/或时间戳整理关于该个人的不同数据点时,这就应算作追踪(Binns, 2022)。
追踪的规范差异很大。追踪可通过各种用户识别工具发生——cookie(Jones, 2020)、指纹(Englehardt & Narayanan, 2016)、favicon(Solomos et al., 2021)、个人可识别信息(PII)等等。追踪可由不同行为者(公共或私人)进行,他们可在不同上下文、不同时间戳,为不同目的——营销、执法、国家安全、用户互动、公共卫生等——获得对数据主体的可见性。消费者被追踪的机会和向量十分充足,尤其当追踪行为者拥有消费者直接参与的各种服务时——从搜索引擎和平台到移动软件、硬件、网站,以及本文所关注的广告网络功能。
我们特别关注作为一种对在线消费者不透明的商业实践的追踪:它在网站与第三方之间的共生关系中产生(Gopal et al., 2022),并通过自动化数据捕获实现,使"被动"商业监控对 Web 上的个人几乎不可避免。6
为评估和衡量追踪的程度,我们制定一个追踪规范框架,详述我们认为衡量追踪技术对消费者隐私影响的最重要标准。我们提出以下五项追踪规范:
[1] 用户识别工具:个人如何被识别?追踪者通过何种追踪工具分配用户 ID,并可能跟踪用户、整理数据点以用于画像(即 cookie、token、指纹、favicon)。
[2] 跨上下文用户可见性:哪些公司能跨网站持久识别消费者?我们感兴趣的是了解哪些广告网络行为者能够识别并"享有"跨 Web 对消费者的可见性(即 SSP、DSP、广告主、发布者、广告交易平台)。
[3] 纵向追踪:该追踪技术是否使消费者能够随时间被追踪?
[4] 绕过定向限制:该追踪机制是否使广告主能够/被激励去构建丰富的一方消费者画像,并暗中规避广告平台的定向限制?
[5] 用户数据源:我们能否限制参与画像的数据行为者?可用于数据整理的可能来源有哪些?关于用户的哪些数据点可被收集并用于定向目的?(即当前浏览行为、过往浏览行为、来自数字足迹的线下数据、一方数据、三方数据)。
通过基于上述追踪规范的演绎分析,3.1 和 3.2 节评估基于 cookie 和无 cookie 的追踪机制。3.3 节对比这些机制,揭示无 cookie 追踪替代方案如何可能实现比当前基于 cookie 追踪实践更大程度的消费者监控。
3.1 基于 cookie 的追踪
为将未来无 cookie 追踪机制与广告网络当前基于 cookie 的追踪相对照,我们首先旨在全面理解当前的追踪实践。为此,我们依靠浏览器侧观测,研究广告网络跨网站持久识别并潜在追踪用户的能力。
为评估 SSP 网络内部及之间持久标识符的使用,我们汇编了一份与广告行业全部四家主要 SSP 网络合作的发布者网站清单。所选四家 SSP 基于作者对它们在数字广告行业突出地位的了解。我们制定了两条不同的选择标准以汇编最终的发布者网站清单。第一条选择标准,发布者的纳入基于其将 SSP 列在"ads.txt"文件中。满足该标准后,我们基于"Tranco"流行度指数(Tranco, n.d.)对网站排名。我们最初汇编的、在其 ads.txt 中列出全部四家主要 SSP 的前 100 个流行网站清单,通过访问每个网站的落地页(而非站内页)进行抓取。分析初始结果时,我们看到 7 某些发布者并未登记来自全部四家 SSP 的 cookie。虽然 ads.txt 表明哪些合作伙伴有资格销售某发布者的广告库存,但基于抓取 100 个网站所观察到的实例,我们假定:仅因某合作伙伴被列出,并不意味着该发布者当前正与该合作伙伴合作。在第二条选择标准中,发布者的纳入基于在 Chrome 浏览器开发者工具的 Cookies Storage 表中人工核对,来自 4 家 SSP 各自的 cookie 在落地页的 10 次刷新中至少有 1 次登记。我们得以从核对 Tranco 流行度指数中、其 ads.txt 列出全部四家 SSP 的网站,制定出一份 50 个网站的清单(见附录 #1)。从这一人工核对中——它表明落地页加载时 cookie 并不总被登记——我们决定对 50 个网站分别抓取 10 次并取平均结果。在这 10 次单独抓取中,我们同样仅访问发布者落地页。
为评估四家 SSP 网络中追踪者对持久用户标识符的使用,我们追踪经由 HTTP cookie 的有状态追踪,因为它们仍是跨网站识别在线用户最主要的技术(Roesner et al., 2012; Fouad et al., 2020)。为收集每个发布者网站的追踪信息(即 HTTP cookie、JavaScript 操作和 HTTP 头),我们使用一款基于开源的自动化网络爬虫——OpenWPM——它模拟真实用户活动并记录网站响应、元数据、所用 cookie 和所执行的脚本(Englehardt and Narayanan, 2016)。我们执行了 10 次单独的有状态抓取,并将抓取设置为只使用一个浏览器实例。我们未在配置中设置"Do Not Track",以允许持久识别,并将模拟浏览器配置为接受所有第三方 cookie。我们还使用"机器人检测规避"来随机上下滚动访问的页面。我们将发布者网站之间的睡眠时间设为 5 秒,网站之间的超时设为 100 秒。这 10 次抓取于 2022 年 9 月 20 日和 21 日在本地机器上运行,数据记录在一个 SQLite 数据库中。
在抓取所创建的 SQLite 数据库文件中,分析了 'http_requests'、'javascript' 和 'javascript_cookies' 表。SSP 基于 'host' 字段识别,并基于包含 SSP 网络名称的行进行过滤。发布者网站基于 'top_level_url' 字段确定。持久标识符基于 cookie 名称与 cookie 值字段的拼接来识别。对于发布者网站的抓取,启用了有状态追踪,我们追踪浏览器的 cookie 存储,以分析相同的 cookie ID 如何在同一 SSP 网络内部及跨 SSP 网络的网站之间被使用。我们将那些在两个或更多发布者网站中使用相同 cookie ID 值的 SSP 称为"持久标识符",因为它们在合作的网站上以相同方式识别用户。我们基于每家公司隐私政策中所含的语法识别 SSP ID cookie(见下表 1)(Magnite, 2021a; OpenX, 2022; Pubmatic, 2020; Xandr, 2022)。8
SSP 网络 | ID cookie 名称
Pubmatic | KADUSERCOOKIE
OpenX | i
AppNexus | uuid2
Rubicon | khaos
表 1:用于追踪每个 SSP 网络用户持久识别的 ID cookie 名称
在我们的数据分析中,我们想首先观察某些 SSP 是否在其网络内部的发布者网站之间持久识别用户。我们观察到,SSP 在 77.6-90% 的网站中使用了持久标识符(见下表 2)。同时,我们承认用户的持久识别也在服务端发生(例如 Acar et al., 2014),其方式对研究者而言更难检测。因此,我们预期我们的结果应被视为 SSP 在所抓取网站上持久识别用户量的下界。按网站划分的 SSP 网络内部持久识别模式的完整可视化可在附录 #2 中找到。
SSP 网络 | 我们的浏览器被持久识别的网站平均百分比
Pubmatic | 86.4%
OpenX | 77.6%
AppNexus | 90%
Rubicon | 90%
表 2:每个被抓取 SSP 网络跨网站持久识别的数量
其次,我们想看看持久标识符跨 SSP 的重叠情况。我们能否发现两个或更多 SSP 使用相同的 ID cookie 值,暗示这些 SSP 以相同方式动态识别个人,从而可能跨其合作的网站关联用户数据?有趣的是,从我们的浏览器侧观测中,我们发现在所有情况下,跨 SSP 网络的持久识别都不存在——尽管有一次实例在所有 10 次单独的无状态抓取中都反复出现,我们无法完全解释(见附录 #3)。我们无法追踪到同一个 cookie ID 或任何 cookie 值被两个不同 SSP 投放到我们的浏览器上。我们承认 9 SSP 之间的 cookie 同步可能在服务端发生、远离我们的爬虫,但这些行为者之间的竞争考量使其不太可能。广告主选择与多个 SSP 合作以对用户出价,其选择 SSP 伙伴主要基于通过发布者独立月访客触达不同受众规模的能力,以及投放准确的库存表现(Sluis, 2018; Vargas, 2022a)。这些关于扩大可寻址受众以及 SSP 所提供服务差异化的考量,表明 SSP 之间存在明显竞争,会抑制 cookie 共享。
总之,回到我们的追踪规范,我们观察到用户可被供给侧网络(SSP)投放在发布者网站上的第三方 cookie 识别。这在使用 ID cookie 跨其嵌入的网站持久识别并潜在追踪个人时,为 SSP 提供了跨上下文可见性。不过,我们的数据表明,基于 SSP 的用户识别仍停留在每个 SSP 的发布者网站网络内部,并未跨 SSP 网络。所观察到的追踪方法也实现了随时间追踪,只要消费者能被识别或关联到同一 cookie。
超出我们的数据收集、并依靠我们对行业出版物和技术文档的审阅,第三方 cookie 还使广告主能够了解用户的浏览历史和过往行为,以便跨网站对发布者的广告库存出价。重要的是,cookie 可与其他数据源关联,使广告主能将其数据与 cookie 配对,超出对浏览行为的被动监控所能提供的范围,从而可能绕过广告平台的定向限制。广告主能够将第三方 cookie 标识符与从数据经纪商处购买的数据相匹配,以基于可能敏感的整理类别定向用户(Experian, n.d.; Sherman, 2021)。就用户数据源而言,目前几乎没有限制。尽管我们未直接测量用户追踪的潜在参与者,先前研究已表明,一系列广告网络行为者——线上与线下——都能为营销目的促成消费者画像(Choi et al., 2020; Wei et al., 2020)。
3.2 无 cookie 追踪
继 Google 于 2020 年宣布其数字广告产品将不再支持第三方 cookie 之后,来自发布者网站的一方数据收集开始兴起,成为数字广告的共识,以在程序化竞价过程中保留第三方 cookie 为追踪与定向所提供的部分功能(Schuh, 2020; Southern, 2020)。发布者开始通过订阅、非付费订阅者注册和时事通讯注册等手段,更好地组织和强化一方持有数据中的用户识别(Asim, 2021)。虽然一方用户数据细分已使用由人口统计和站内行为数据构建的定向类别,以在单个发布者网站上直接向用户提供相关广告,但广告主仍有兴趣捕获关于用户跨 Web 行为的信息,这类似于基于 cookie 持久识别模式的追踪。由于保留跨网站跟踪行为所带来的、定向相关受众的精确性,仍是广告主的首要任务,SSP 最初被认为可能成为一个协调机构,能够跨发布者网站聚合一方数据,使广告主仍能跨发布者网站触达相关受众(Joseph, 2021)。然而,被称为"Universal ID"和"Alternative ID"的替代方案已开始大量涌现,它们跨多个 SSP 网络运作,比当前按 SSP 分割的识别更持久地识别个人用户。10
为将提案与现有的基于 cookie 方案对比,我们选择了我们认为三种主要的替代标识符方案——"The Trade Desk Unified ID 2.0(UID 2.0)"、"LiveRamp RampID"和"Secure Web Addressability Network(SWAN)"。我们的选择基于我们对覆盖数字广告行业的行业出版物的阅读,以及关于特定方案的报道频率。我们将这些报道纳入我们所收集的数据,连同不同方案作者公开的信息和技术文档。因此,基于我们在第 3 节提出的追踪规范,我们分析业内讨论的三大主要无 cookie 追踪方案中的每一个。
就用户识别工具而言,三个方案各有不同。对于 SWAN,当个人首次访问采用该方案的发布者网站时即被识别(Asim, 2022; Schiff, 2021d; Thomson & Rescorla, 2021)。加载发布者页面后,用户会看到一个弹窗,请求同意在当前网站及其他采用该方案的网站上展示个性化广告。作为弹窗的一部分,用户还可以选择分享其电子邮件地址,该地址可作为一个标识符。无论用户是否选择接收个性化广告、是否分享其电子邮件地址,最初访问的发布者网站都会在 SWAN 网络内放置一个一方 cookie,创建一个存储在用户浏览器上的假名标识符。对于 UID 2.0——由一家顶级需求侧平台(DSP)撰写的方案——用户身份通过电子邮件登录发布者网站来建立。发布者将电子邮件地址存储在被放置在页面上的一方 cookie 中(Asim, 2022; Thomson & Rescorla, 2021; UnifiedID2, 2022)。电子邮件通过一个相应 token 匹配到 UID2.0,该 token 也被创建并用于加密 UID2.0,只有通过同意 Unified ID 2.0 服务条款而收到解密密钥的合作伙伴才能解密。假名 UID2 的创建由 UID2.0 服务管理。所分析的第三个方案 LiveRamp RampID 与前两者不同,因为它可与其他 ID 方案互操作,包括 UID2.0(Asim, 2022)。RampID 使用用户电子邮件地址,将其与为换取内容而分享给发布者的电子邮件地址相匹配(Asim, 2022; LiveRamp, 2022c)。LiveRamp 不放置一方 cookie,而是通过其专有的认证流量解决方案(ATS)在生态系统中匹配 ID(Asim, 2022; LiveRamp, 2022d)。该方案在生态系统中进一步识别用户的能力,还与其他线下 PII(电话号码、地址历史)相关联,基于广告主可将其与整理的一方、二方和三方数据相匹配的信息。
关于跨上下文用户可见性,我们看到访问 SSP 及其发布者伙伴,仍给广告主提供了跨网站触达用户的机会。某个 ID 方案的取向重现了现有 SSP 对发布者追踪的取向。这些方案在与主要 SSP 合作以协调标识符程序化竞价进行交易的同时,也在重建一个标识符以跨 SSP 识别用户(Asim, 2022)。在 3.1 节中,我们展示了当前基于 cookie 的用户识别似乎按单个 SSP 分割。然而,对于所分析的无 cookie 追踪方案,我们主张:不仅这些网络在 SSP 伙伴关系中跨发布者网站被复制,而且行业领先 SSP 对 ID 方案的整合,可能导致用户跨 SSP 被识别,而不仅是在 SSP 内部。结果,将实现跨发布者网站更大程度的用户识别。所有三个所分析的方案都得到了主要 SSP 的支持和承诺采用,且 The 11 Trade Desk Unified ID 2.0 已宣布与我们作为数据抓取一部分所分析的全部四家 SSP 建立伙伴关系(Asim, 2022; Schiff, 2020a; Schiff, 2020b; Schiff, 2020c; Schiff, 2021a; Schiff, 2021d)。这给那些将被指派治理这些方案者委以大量责任,因为对消费者跨 Web 行为的视野将被扩大。
纵向追踪显然被所有三个无 cookie 追踪方案实现。其持久性源于作为每个方案基础的确定性数据(Asim, 2022; Kaye, 2021b)。每个方案都向用户提供选择退出所有参与伙伴的个性化广告的选项(Asim, 2022; LiveRamp, 2022b; SWAN-community, 2021; UnifiedID2, 2022)。也可以论证,其持久性可能比用户频繁删除的第三方 cookie 持续更久,因为选择退出可能意味着用户失去对发布者内容的访问。虽然广告主能够对分类为某一持久标识符的用户出价,但这些具体方案尚未回答它们是否会支持最持久的定向方法——再定向或再营销(retargeting/remarketing)——即在某个初始行为将用户归类为可能更倾向于完成某项可归因于数字广告的期望行为之后,对该用户的持续跟踪。
有趣的是,所有三个方案都进一步鼓励绕过广告平台的定向限制。广告主可基于其一方数据持久识别用户,轻易覆盖主要广告平台的定向政策。我们此前已暗示,负责治理新追踪方案者负有责任,因为它们为跨 Web 更大程度的持久识别提供了条件。但这里我们要强调的是无 cookie 提案中一个讨论不足的缺口。广告主如今可在持久识别中扮演更大角色,不仅依靠发布者一方数据,还依靠广告主自身上传数据以编码用于定向的能力。例如,在 The Trade Desk Unified ID 2.0 中,公司提到"一方关系(First-Party Relationships)"能力,即广告主可以上传一方数据以编码为 UID2,从而在发布者网站间激活(UnifiedID2, 2022)。同样,LiveRamp 为广告主提供"入驻(onboard)"其数据的机会,其中 PII 可被上传以转换为 RampID 并按细分组织,从而可在 500 多个不同伙伴平台上激活(LiveRamp, 2022a)。SWAN 未提供关于这些能力的太多信息,但在其主页上表示它"与 CRM 数据互补"(Secure Web Addressability Network (SWAN), 2021)。这一由所有三个方案实现的能力,在其他替代方案中也被发现有可能进一步模糊广告主的定向实践。
鉴于广告主在 PII 被身份方案编码用于定向之前所能做的工作,这种能力令人担忧。如 3.1 节所述,广告主也选择超出 cookie 信号来定向用户,并购买第三方数据以扩展用户画像,超出对浏览器行为的被动监控所能提供的范围。无 cookie 方案对一方数据的强调,已鼓励广告主开始利用其现有客户群,通过建模实践找出与当前客户特征相似的其他用户,以创建"相似受众(look-alike)"细分(LiveRamp, 2020a)。与广告主在当前基于 cookie 生态中的努力相比,且可定向信息 ID 所实现的粒度水平仍不明确,广告主如今正加大对能将一方 12 订阅者数据与其他数据集相匹配的技术的投入(Vargas, 2022b)。这项工作在一个"身份图谱(identity graph)"上进行,它允许广告主管理个人层级数据,并通过所选 ID 方法对数据编码,提供一个集中式系统,将线上和线下标识符合并为一个整合的画像,以与购买的第三方数据配对,并与多个伙伴激活选定受众(LiveRamp, 2020c)。身份图谱对个人隐私构成威胁,因为它允许广告主通过将购自数据经纪商的数据配对,构建现有客户——以及开放 Web 上非客户——的丰富画像(Vargas, 2022b)。Unified ID 2.0 和 LiveRamp RampID 如今在这项技术中实现了令人警觉的 ID 关联能力(Schiff, 2021b; Schiff, 2021c; Schiff, 2022b; The Trade Desk, 2021c)。当广告主能够按细分上传 ID 列表时,就有可能不仅上传现有客户 PII 的明确列表,还上传已按特定受众细分类别分组的电子邮件。由于这些细分是使用购自数据经纪商的数据构建的,细分和画像的构建就有可能涉及敏感的类别信息。这些 ID 旨在跨主要平台互操作,而无法核实或指定正被转换为交易的 PII 的分段实践,为广告主违反这些平台特定的定向实践创造了机会(Google, 2022; Meta 2021; Twitter, 2022)。没有任何机制能核实广告主在导入这些系统之前是如何对一方数据分段的,而新的 ID 方案使敏感人群分段和出价对广告主极具吸引力。
例如,一家信用卡公司可能想专门定向非裔美国男性,但无法通过某些主要数字广告平台的受限定向类别做到这一点。相反,该信用卡公司可以基于数据经纪商获取的、能将电子邮件地址与种族关联的数据,汇编一份潜在客户名单。如果该信用卡公司选择用 LiveRamp ID 编码其数据,并在该公司声称合作的 500 个伙伴平台目的地中的任何一个上激活,该公司就只是在没有彻底审查的情况下提供了一个绕过某些公司定向限制的链接(LiveRamp, 2022a, LiveRamp, 2022b)。先前研究已发现广告主上传的名单在定向用户时违反平台政策的实例,使用了种族、宗教、政治、性生活或健康等类别(Wei et al., 2020)。我们主张,新的追踪 ID 方案极度重视利用一方数据,这鼓励广告主寻找通过 ID 方案关联线上和线下数据的方法,从而允许在主要平台几乎没有监督的情况下进行定向实践。这种能力鼓励广告主在与主要平台互动之前就对用户画像,绕过现有的平台限制。
关于用于追踪的用户数据源,我们发现与基于 cookie 的追踪类似,被允许利用和交易这些 ID 的参与者限制仍不明确。负责识别个人的第三方的基本结构几乎完全相同。现在由不同的一方提供持久标识符,广告主仍通过需求侧平台传递 ID 以跨网站对个人出价。在提供关于参与者应如何遵循某些原则的一般条款、承诺将概述行为准则、或对专有治理保密的情况下,尚不清楚当前生态系统的哪些其他成员可以参与(Asim, 2022; LiveRamp, 2022c; Thomson & Rescorla, 2021; SWAN-13 community, 2021; UnifiedID2, 2022)。这一问题尚未解决,即便无 cookie 追踪方案声称通过限制向各方分享持久个人层级信息,来提供很大的个人隐私。我们发现,负责治理这些方案的一方既不明确,参与者被期望遵循的条款也不明确。SWAN 方案将由 SWAN 网络自身治理,该网络概述了一套"示范条款(Model Terms)",详述参与者应如何遵循信息分享实践(Thomson & Rescorla, 2021; SWAN-community, 2021)。The Trade Desk 声称公司计划将控制权移交给一位"管理员(Administrator)",这一角色尚未有人担任(Asim, 2022)。The Trade Desk 文档还提到参与者必须遵循的"行为准则",但目前不可得(UnifiedID2, 2022)。管理员的角色原本应由 Interactive Advertising Bureau(互动广告局)担任,但该组织已选择不再以此身份支持该方案(Katsur, 2022; Mitchell, 2021)。Prebid 也拒绝担任管理员(Shields, 2022)。SWAN 和 Unified ID 2.0 都是开源的,但 LiveRamp ID 是公司自身管理的专有方案(Asim, 2022; LiveRamp, 2022c)。基于透明度的缺乏,尚不清楚这些方案将如何执行标准,以确保参与者不违反尚未完全定义或公开的运营条款。
3.3 对比基于 cookie 与无 cookie 的追踪机制
下表 3 总结了我们对广告行业当前与未来 Web 追踪方案的对比。该表显示消费者监控预期将在以下方面扩大:(1)更宽的持久识别模式;(2)可能跨越更长的时间段追踪;(3)激励广告主绕过定向限制,并基于敏感且更丰富的一方数据画像对消费者出价。
我们的发现表明,Web 上的监控正从当前基于 cookie 的组织方式,向基于 ID 方案交易的方式增长。第一,持久识别已增加,因为识别不再局限于单个 SSP 网络。第二,负责确定身份的数据源源自 PII 和同意机制,使个人难以选择退出以内容换取可启用追踪的交易。第三,上传数据的能力已鼓励广告主寻找在与主要平台互动之前对受众分段的方法。这导致从数据经纪商购买线下数据,这些数据可与现有客户数据配对,或用于基于与现有客户或其他偏好特征的相似性来推断潜在客户。这一做法允许选择可能违反主要平台政策的特征——这些政策禁止将某些敏感类别用于用户定向。违反平台政策的原因在于,没有一种方法能在平台上激活之前,彻底核实个人是如何将生成的细分编码到 ID 方案中的。14
基于 cookie 的追踪 | 无 cookie 的追踪
SSP Cookie ID | UID 2.0 | SWAN | LiveRamp
用户识别工具 | 被动放置第三方 ID cookie | 基于在发布者网站获得的同意并分享电子邮件的一方 cookie | 基于在发布者网站获得的同意并分享电子邮件的一方 cookie | 基于在发布者网站获得的同意并分享电子邮件、并与线下姓名、地址和电话号码合并的专有认证流量解决方案
跨站用户可见性 | 在全部四家 SSP 内,跨网络 77%-90% 的网站,但不跨 SSP | 跨(全部四家 SSP 网络),而不仅是在其内部 | 跨(全部四家 SSP 网络),而不仅是在其内部 | 跨(全部四家 SSP 网络),而不仅是在其内部
纵向追踪 | 是 | 是——依赖更多确定性数据 | 是——依赖更多确定性数据 | 是——依赖更多确定性数据
绕过定向限制 | 是 | 是——且通过将广告主一方数据与购买的三方数据配对,鼓励更大程度的画像 | 是——暗示具备与 CRM 数据配对的能力 | 是——且通过将广告主一方数据与购买的三方数据配对,鼓励更大程度的画像
用户数据源 | 未设限制 | 对参与者和治理机制的限制仍不明确 | 对参与者和治理机制的限制仍不明确 | 对参与者和治理机制的限制仍不明确
表 3 - 通过 cookie 与无 cookie 方案用于营销目的的追踪规范 15
4 - 讨论
基于对四家主要 SSP 行为者关于基于 cookie 追踪的数据收集,以及对广告行业关于无 cookie 追踪方案的技术文档的分析,我们的研究凸显了三大无 cookie 追踪方案预期将加剧 Web 上消费者监控的三项追踪规范。第一,消费者跨网站的持久识别很可能跨越 SSP 网络,从而对消费者产生潜在更大的实时可见性。第二,识别机制预期将依赖 PII,使其更持久,且更可能随时间更好地追踪用户。第三,随着行业向"身份图谱"和一方数据转向,广告主如今被激励基于丰富的一方和第三方数据画像来定向消费者,如果广告主选择基于敏感和受禁类别定向消费者,就可能覆盖现有的定向限制。因此,Web 上的消费者追踪很可能变得在动态上更宽、涉及更丰富的消费者历史,并依赖更多种类的数据源。
与现有对无 cookie 追踪方案的批评不同,我们的研究展示了广告网络对拟议追踪机制所造成潜在隐私危害的结构性影响。我们基于所制定的追踪规范,系统性地将未来方案与当前可运作的追踪机制相对照。我们超越同意机制的缺陷(Kaye, 2021a; Thomson & Rescorla, 2021)或明确治理机构的缺失(UnifiedID2, 2022),来展示相比当前追踪机制,广告网络的结构与流程如何导致潜在更长久、更宽泛、更丰富的消费者监控。
我们的发现对拟议追踪机制实现 GDPR 合规的能力提出了严重质疑。尽管公民社会反对当前基于 cookie 的追踪生态(Lomas, 2021),且新追踪方案被营销为"隐私友好"(The Trade Desk, 2021a; Thomson & Rescorla, 2021),我们展示了这些新机制如何能够实际地基于敏感类别对个人进行细粒度追踪和定向。此外,鉴于广告主被激励去构建历史上丰富的消费者画像,"挑出"个人对广告主而言可能变得更容易(GDPR 序言第 26 条)。与关于 Universal ID 方案和 GDPR 合规的现有辩论(大多聚焦于拟议架构中数据控制者和数据处理者角色的缺失(Schiff, 2021e; Schiff; 2022a),或所提方案中 cookie 的角色(Asim, 2021; Asim, 2022))不同,我们想强调新方案所允许的画像与定向的程度,以及它们可能违反关键 GDPR 原则之处。
第一,所预期的画像策略中涉及的个人信息的潜在丰富程度,可能违背或超出消费者的合理预期,并可能侵犯适用的数据保护原则和规则。例如,当一家广告公司将其自己的一方数据与第三方数据源结合时,如 3.2 节所述,这可能导致个人数据被用于超出其初始目的、且个人无法合理预见的方式。广告主构建的画像可能涉及对个人并未主动披露的兴趣或特征的推断,削弱个人对其个人数据行使控制的能力(EDPS, 2018),或为广告主从其数据中"挑出"个人创造机会(GDPR 序言第 26 条)。此外,GDPR 的 16 透明度要求可能被违反,因为考虑到可能绕过平台的定向限制,不同各方在这一过程中的角色对用户而言很可能不清晰(相关示例见 3.2 节)。
第二,所提议的追踪机制可能鼓励歧视和排斥。广告主的定向可能涉及直接或间接与个人种族或族裔出身、健康状况或性取向、或所涉个人的其他受保护品质相关的、具有歧视效果的准则。定向中潜在的歧视,源于广告主利用大量且多样的个人数据的能力——鉴于新追踪技术围绕确定性识别转向。如 3.2 节所解释,广告主将自身用户分段与 ID 方案关联的能力,使细粒度且可能歧视性的定向成为可能,覆盖现有不可执行的定向限制,从而违反 GDPR 第 9 条关于处理特殊类别个人数据的要求。
5 - 结论
广告网络结构与流程的隐私影响,导致了一个即使在无第三方 cookie 的情况下也令人担忧的广告格局。AdTech 复合体对无 cookie 追踪方案的实现,可能实现对消费者更大的动态可见性、更长的消费者追踪,以及更敏感的消费者画像的组装。
展望未来,尽管这批方案中没有明确的头号追踪方案,且行业领袖并不认为单一方案将独自负责持续的持久识别,但 The Trade Desk Unified ID 2.0 已脱颖而出成为领先方案。尽管由于该提案与现有追踪方法的相似性,The Trade Desk Unified ID 2.0 及其他方案能否符合不断演变的隐私法规存在不确定性,但对 The Trade Desk 的支持仍在持续积累,其基于强调隐私但同时承诺交付第三方 cookie 所提供同样能力的营销语言(Asim, 2021; The Trade Desk, 2021a)。对 The Trade Desk UID 的批评或可被视为被以下因素抵消:不仅来自 IPG、Omnicom 和 Publicis Groupe 等代理商(Bürgi, 2021a; Bürgi, 2021b; The Trade Desk, 2021b; The Trade Desk, 2021c)的行业支持,还有来自 Google 的支持。虽然 Google 最初于 2021 年宣布,因其认为该方案在不断演变的监管环境中不可持续,将不在其生态系统中支持电子邮件或替代性第三方标识符(Temkin, 2021),但该公司通过推出发布者加密信号(ESP)产品改变了方向——该产品允许发布者通过 Google 的 Ad Manager 分享加密的一方数据和替代标识符(Schiff, 2022c; Google Ad Manager Help, 2022)。有了这一支持,The Trade Desk 如今得到了那些在维持程序化广告预算和实践方面有高度利益的相关方的支持。
鉴于本研究中的发现,这是一个令人警觉的趋势。广告行业对 UID 2.0 的实现可能使消费者监控变得更糟。只要广告网络的结构不发生根本改变,消费者画像就会作为一项极受重视的商品持续存在。17
我们的研究有一些局限。第一,对于基于 cookie 机制的评估,我们抓取的是流行网站的 50 个落地页而非其内页,而内页的追踪已知更为普遍。第二,我们基于相同字符串对比 cookie 值,尽管某些行为者可能加密或哈希其 cookie 值,使我们漏掉一些持久识别趋势。第三,我们无法追踪服务端正在分享的信息,也无法完全捕获广告网络行为者关于个人用户所拥有细节的数量。这三个局限使我们假定:我们关于基于 cookie 追踪的发现,暗示了实际持久识别模式的一个下界。先前研究支持我们的假定,承认广告行为者常将一方用户数据与所观察到的 cookie ID 相匹配,以实现更细粒度的定向能力(Trusov et al., 2016)。
未来的后续研究项目可以投入理解对隐私倡议的合规如何被转移到个体广告主层级,如 3.2 节所展示。广告平台制定政策以限制敏感类别被用于定向,但同时又仍赋予广告主精确定向用户的能力。这些宽松可执行的机制是广告网络隐私问题的一部分,亟需结构性变革。18
参考文献
Acar Gunes, Christian Eubank, Steven Englehardt, Marc Juarez, Arvind Narayanan, and Claudia Diaz. (2014). "The Web Never Forgets: Persistent Tracking Mechanisms in the Wild." Proceedings of the ACM Conference on Computer and Communications Security (CCS).
Aguirre, E., Mahr, D., Grewal, D., de Ruyter, K., & Wetzels, M. (2015). Unraveling the Personalization Paradox: The Effect of Information Collection and Trust-Building Strategies on Online Advertisement Effectiveness. Journal of Retailing, 91(1), 34–49.
Asim, A. (2021, November 17). Digiday Media Research: A comprehensive guide to third-party cookie alternatives. Digiday.
Asim, A. (2022, June 9). Digiday+ Research: A guide to the top 10 ID alternatives for publishers. Digiday.
Bashir, M. A., Arshad, S., Wilson, C., & Robertson, W. (2016). Tracing Information Flows Between Ad Exchanges Using Retargeted Ads. Proceedings of the 25th USENIX Security Symposium.
Binns, R. (2022). Tracking on the Web, Mobile and the Internet of Things. Foundations and Trends® in Web Science, 8(1–2), 1–113.
Bürgi, M. (2021a, August 10). Unified ID 2.0 quietly amasses more support from the agency world, but publishers aren't as convinced. Digiday.
Bürgi, M. (2021b, November 17). Omnicom Media Group formally endorses UID 2.0 in a bid to move the post-cookie future forward. Digiday.
Choi, H., Mela, C. F., Balseiro, S. R., & Levy, A. (2020). Online Display Advertising Markets: A Literature Review and Future Directions | Information Systems Research. Information Systems Research, 2(31), 556–575.
D'Annunzio, A., & Russo, A. (2020). Ad Networks and Consumer Tracking. Management Science, 66(11), 5040–5058.
Dinev, T. (2014). Why would we care about privacy? European Journal of Information Systems, 23(2), 97–102.
Dinev, T., & Hart, P. (2006). An Extended Privacy Calculus Model for E-Commerce Transactions. Information Systems Research, 17(1), 61–80.
Dinev, T., Xu, H., Smith, J. H., & Hart, P. (2013). Information privacy and correlates: An empirical attempt to bridge and distinguish privacy-related concepts. European Journal of Information Systems, 22(3), 295–316.
Edelman, G. (2020, October 5). Ad Tech Could Be the Next Internet Bubble. Wired.
EDPS. (2018). EDPS Opinion on online manipulation and personal data.
Englehardt S. and A. Narayanan. (2016). "Online Tracking: A 1-million-site Measurement and Analysis." Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388-1401.
Experian. (n.d.). ConsumerView: Data by the Numbers.
Fouad I., N. Bielova, A. Legout, N. Sarafijanovic-Djukic. (2020). "Missed by Filter Lists: Detecting Unknown Third-Party Trackers with Invisible Pixels."
Goel, Vinay. (2021, June 24). An updated timeline for privacy sandbox milestones. The Keyword.
Google Ad Manager Help. (2022, March). Share encrypted signals with bidders (Beta). Google Ad Manager Help.
Goldfarb, A., & Tucker, C. (2011). Online Display Advertising: Targeting and Obtrusiveness. Marketing Science, 30(3), 389–404.
Goldfarb, A., & Tucker, C. E. (2015). Standardization and the Effectiveness of Online Advertising. Management Science, 61(11), 2707–2719.
Google. (2022). Personalized advertising - Advertising Policies Help.
Gopal, R. D., Hojati, A., & Patterson, R. A. (2022). Analysis of third-party request structures to detect fraudulent websites. Decision Support Systems, 154, 113698.
Hui, K.-L., Teo, H. H., & Lee, S.-Y. T. (2007). The Value of Privacy Assurance: An Exploratory Field Experiment. MIS Quarterly, 31(1), 19–33.
Jones, M. L. (2020). Cookies: A legacy of controversy. Internet Histories, 4(1), 87–104.
Joseph, S. (2021, November 17). 'They will need to use multiple routes': Shifts appear in the publisher-SSP union, as alternative identifiers proliferate. Digiday.
Kaye, K. (2021a, March 24). Third-party cookie replacements fall short of consent and transparency promises. Digiday.
Kaye, K. (2021b, April 1). WTF is the difference between deterministic and probabilistic identity data? Digiday.
Katsur, A. (2022, February 28). Tech Lab Update on UID2.0. IAB Tech Lab.
Lavin, M. (2006). Cookies: What do consumers know and what can they learn? Journal of Targeting, Measurement and Analysis for Marketing, 14(4), 279–288.
Li, T., & Unger, T. (2012). Willing to pay for quality personalization? Trade-off between quality and privacy. European Journal of Information Systems, 21(6), 621–642.
Li, Y. (2011). Empirical Studies on Online Information Privacy Concerns: Literature Review and an Integrative Framework. Communications of the Association for Information Systems, 28.
Li, Y. (2012). Theories in online information privacy research: A critical review and an integrated framework. Decision Support Systems, 54(1), 471–481.
Lively, T. K. (2022, July 7). US State Privacy Legislation Tracker.
LiveRamp. (2020a, March 31). Look-alike Modeling: The What, Why, and How.
LiveRamp. (2020b, August 11). Platform-Specific Distribution Information.
LiveRamp. (2020c, August 28). What's the Difference between a DMP and an Identity Graph?
LiveRamp. (2022a, March 8). Onboarding Your Data.
LiveRamp. (2022b, May 6). Consumer Requests for Opt-Outs, Data Access, or Data Deletions.
LiveRamp. (2022c, June 30). RampID Methodology.
LiveRamp. (2022d, July 21). Authenticated Traffic Solution.
Lomas, N. (2021, June 16). Adtech 'data breach' GDPR complaint is headed to court in EU. TechCrunch.
Mac, R., & Kang, C. (2021, October 3). Whistle-Blower Says Facebook 'Chooses Profits Over Safety.' The New York Times.
Magnite (2021a, August 27). Data Subject Rights Policy. Magnite.
Magnite (2021b, August 27). Platform Cookies statement. Magnite.
Malhotra, N. K., Kim, S. S., & Agarwal, J. (2004). Internet Users' Information Privacy Concerns (IUIPC): The Construct, the Scale, and a Causal Model. Information Systems Research, 15(4), 336–355.
Marotta, V., Wu, Y., Zhang, K., & Acquisti, A. (2022). The Welfare Impact of Targeted Advertising Technologies. Information Systems Research, 33(1), 131–151.
Martin, K. (2022). Finding Consumers, No Matter Where They Hide: Ad Targeting and Location Data. In Ethics of Data and Analytics: Concepts and Cases (pp. 99–111). Auerbach Publications.
Meta. (2021, November 9). Removing Certain Ad Targeting Options and Expanding Our Ad Controls.
Mitchell, J. (2021, January 21). Reviewing Unified ID 2.0 for Long-Term Industry Value – IAB Tech Lab. IAB Tech Lab.
OpenX (2022, March 3). OpenX Ad Exchange Privacy Policy. OpenX.
O'Reilly, L. (2020, January 14). Google plans to kill off third-party cookies in Chrome "within 2 years." Digiday.
Parkin, R. (2021, June 27). What the Delay to the End of Third-Party Cookies Means for Advertisers. AdExchanger.
Pubmatic (2020, July 1). Platform Cookie & Other Similar Technologies Policy. Pubmatic.
Roesner Franziska, Tadayoshi Kohno, and David Wetherall. (2012). "Detecting and defending against third-party tracking on the web." In Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation, NSDI, pages 155–168.
Schiff, A. (2020a, November 17). Magnite Hops Aboard The Unified ID 2.0 Train. AdExchanger.
Schiff, A. (2020b, November 19). PubMatic Is The Latest Ad Tech Company To Join Unified ID 2.0. AdExchanger.
Schiff, A. (2020c, December 16). OpenX Is Latest SSP To Join Unified ID 2.0. AdExchanger.
Schiff, A. (2021a, March 3). Xandr Integrates With Unified ID 2.0 And Outlines Its Identity Roadmap. AdExchanger.
Schiff, A. (2021b, April 21). ID Graph Provider Infutor Joins The Club With Support Unified ID 2.0. AdExchanger.
Schiff, A. (2021c, May 26). LiveRamp Launches Identity Resolution For First-Party Data. AdExchanger.
Schiff, A. (2021d, May 27). SWAN Vs. SWAN: The Differences Between The Two 3P Cookie Alternative Proposals. AdExchanger.
Schiff, A. (2021e, November 16). Unified ID 2.0 Faces Roadblocks In Europe As A Result Of GDPR. AdExchanger.
Schiff, A. (2022a, March 3). IAB Tech Lab Declines To Be The Admin For UID2 (But Hope Springs Eternal). AdExchanger.
Schiff, A. (2022b, March 24). Epsilon Is Making Its Identity Platform Interoperable With Unified ID 2.0. AdExchanger.
Schiff, A. (2022c, March 24). Google's Encrypted Signals Program Just Entered Open Beta, And Here's What You Need To Know About It. AdExchanger.
Schuh, J. (2020, January 14). Building a more private web: A path towards making third party cookies obsolete. Chromium Blog.
Shein, E. (2021, September 14). Third-party cookies are going away: What advertisers, marketers and consumers should know. TechRepublic.
Sherman, J. (2021). Data Brokers and Sensitive Data on U.S. Individuals. Duke University Sanford School of Public Policy.
Shields, R. (2022, January 7). With Prebid at the helm of UID 2.0, indie ad tech marches to a unified beat but not all voices are in harmony. Digiday.
Sluis, S. (2018, November 20). Advertiser Perceptions: How SSPs Can Win Market Share From Google. AdExchanger.
Solomos, K., Kristoff, J., Kanich, C., & Polakis, J. (2021). Tales of Favicons and Caches: Persistent Tracking in Modern Browsers. Proceedings 2021 Network and Distributed System Security Symposium.
Southern, L. (2020, January 29). Beyond ad targeting, the demise of the third-party cookie will hit key digital media functions. Digiday.
Secure Web Addressability Network (SWAN). (2021). Secure Web Addressability Network (SWAN).
SWAN-community (2021). Secure Web Addressability Network (SWAN) - Model Terms Explainer.
Temkin, D. (2021, March 3). Charting a course towards a more privacy-first web. Google.
The Trade Desk. (2021a, February 21). What the Tech is Unified ID 2.0? The Trade Desk.
The Trade Desk. (2021b, April 8). Publicis Groupe all-in on first-party identity solution. The Trade Desk.
The Trade Desk. (2021c, July 28). IPG Leans into Unified ID 2.0 as a Closed Operator. The Trade Desk.
The Trade Desk. (2022). Unified ID 2.0 Partners.
Thomson, M., & Rescorla, E. (2021, August). Comments on SWAN and Unified ID 2.0. Mozilla.
Tranco. (n.d.). A research-oriented top sites ranking hardened against manipulation—Tranco.
Twitter. (2022). Targeting of Sensitive Categories.
UnifiedID2 (2022). uid2docs.
Vargas, A. (2022a, February 4). Google Reigns Supreme In Latest Advertiser Perceptions SSP Report, But Competition Is Tight Among Everyone Else. AdExchanger.
Vargas, A. (2022b, April 12). Goodway Group Stitches Together Identity Graph To Complement Brands' First-Party Data. AdExchanger.
W3C Working Group. (2019, January 22). Tracking Compliance and Scope.
Wei, M., Stamos, M., Veys, S., Retinger, N., Goodman, J., Herman, M., Filipczuk, D., Weinshel, B., Mazurek, M., & Ur, B. (2020). What Twitter Knows: Characterizing Ad Targeting Practices, User Perceptions, and Ad Explanations Through Users' Own Twitter Data. In 29th USENIX Security Symposium (USENIX Security 20) (pp. 145-162).
Xandr (2022, April 29). Digital Platform Cookie Policy. Xandr.
Xu, H., Dinev, T., Smith, J., & Hart, P. (2011). Information Privacy Concerns: Linking Individual Perceptions with Institutional Privacy Assurances. Journal of the Association for Information Systems, 12(12).
附录 #1 – 为评估 SSP 网络上基于 cookie 追踪而抓取的 50 个网站
https://www.yahoo.com, https://www.msn.com, https://www.tumblr.com, https://www.nytimes.com, https://www.sohu.com, https://www.imdb.com, https://www.ebay.com, https://www.forbes.com, https://www.washingtonpost.com, https://www.dailymail.co.uk, https://www.tinyurl.com, https://www.nature.com, https://www.weather.com, https://www.usatoday.com, https://www.cnet.com, https://www.tribunnews.com, https://www.aol.com, https://www.goodreads.com, https://www.time.com, https://www.foxnews.com, https://www.ted.com, https://www.merdeka.com, https://www.wired.com, https://www.independent.co.uk, https://www.latimes.com, https://www.huffingtonpost.com, https://www.kompas.com, https://www.theverge.com, https://www.speedtest.net, https://www.detik.com, https://www.huffpost.com, https://www.nicovideo.jp, https://www.britannica.com, https://www.buzzfeed.com, https://www.usnews.com, https://www.nypost.com, https://www.merriam-webster.com, https://www.9gag.com, https://www.sciencedaily.com, https://www.apnews.com, https://www.youm7.com, https://www.mirror.co.uk, https://www.timeanddate.com, https://www.gsmarena.com, https://www.politico.com, https://www.ndtv.com, https://www.dictionary.com, https://www.chron.com, https://www.vnexpress.net, https://www.thesaurus.com
附录 #2 - 四家 SSP 网络内部的持久识别
(原文为图,此处略)
附录 #3 - 观察到的 Pubmatic 与 Rubicon 之间的信息分享
在我们关于 SSP 网络中基于 cookie 追踪的数据收集里,我们确实看到,在与 Pubmatic 关联的两个唯一 cookie 主机名中,Rubicon 在 10 次无状态抓取中的第 1 次被提及(见下表 4)。我们确实在 10 次无状态抓取中的每一次都看到了其他出现情况,带有不同的持久标识符和相似的主机名语法。
进一步查看 Rubicon(现称"Magnite")的隐私政策,该公司将 Pubmatic 列为"广告 cookie – 第三方技术提供方"合作伙伴,其活动被归类为"DMP、Onboarders、Data Providers(数据提供方)"。
我们仍在试图弄清楚这些行为者之间这种潜在的 cookie 同步对消费者追踪意味着什么。
持久标识符 | 主机名
KADUSERCOOKIE26F46775-238D-433A-B321-8B4D7D2349A9 | https://ads.pubmatic.com/AdServer/js/user_sync.html?gdpr=&gdpr_consent=&us_privacy=1---&predirect=https%3A%2F%2Fprebid-server.rubiconproject.com%2Fsetuid%3Fbidder%3Dpubmatic%26gdpr%3D%26gdpr_consent%3D%26us_privacy%3D1---%26account%3D%7B%7Baccount%7D%7D%26f%3Db%26uid%3Dnull
表 4:在 10 次无状态抓取中的第 1 次里 Pubmatic 与 Rubicon 之间可能的 ID cookie 分享
Toward (Greater) Consumer Surveillance in a ‘Cookie-less’ World:
A Comparative Analysis of Current and Future Web Tracking Mechanisms
Ido Sivan-Sevilla & Patrick T. Parham (UMD)
ABSTRACT
The anticipated shift of the advertising industry away from third-party cookies has been marketed as ‘privacy friendly.’ New cookie-less tracking technologies are being proposed, but the consumer privacy implications of those technologies are far from clear. To what extent ad-networks are going to change their practices, and no longer rely on cross-site consumer surveillance or historically rich consumer profiles for advertising purposes? Can the new tracking technologies become GDPR-compliant, given the significant compliance pushback against cookie-based advertising mechanisms?
Our study seeks to evaluate the potential privacy harms of cookie-less advertising ID solutions by (1) building novel typology of tracking specifications to assess the privacy impacts of tracking technologies; (2) deductively analyzing cookie-based tracking mechanisms by collecting novel data on persistent user identification across 50 popular websites that work with all four main supply side platforms (SSPs) - Pubmatic, OpenX, AppNexus, and Rubicon; and (3) contrasting those findings with deductive analysis of data collected from technical industry documentation on the three main cookie-less ID architectures - The Trade Desk Unified ID 2.0, LiveRamp ID, and Secure Web Addressability Network (SWAN).
We find how the new tracking architectures can make consumer surveillance in the Web dynamically wider, more persistent over time, and extra vulnerable to integration of first- and third-party data by advertisers. This could lead to the circumvention of consumer targeting restrictions posed by the main advertising platforms, in case advertisers choose to bid on ads based on sensitive profile categories. In contrast to existing criticism on cookie-less tracking solutions that mostly focuses on deficiencies in consent mechanisms or the lack of a governing body for the new solutions, our study underscores the structural impact of ad networks on the potential privacy harms caused by the proposed tracking mechanisms. Structures and processes of ad networks might lead to potentially longer, wider, and richer consumer surveillance, compared to cookie-based tracking mechanisms. Our findings also question the ability of suggested tracking architectures to become GDPR compliant. Despite civil society pushback against the current cookie-based tracking ecosystem and the marketing of new tracking solutions as ‘privacy-preserving-advertisement,’ we show how these technologies enable granular tracking and targeting of consumers. ‘Singling out’ individuals might become easier for advertisers, given the historically rich consumer profiles that advertisers are now incentivized to build. In contrast to existing debates about Universal ID solutions and GDPR compliance, that mostly focus on the lack of data controller and data processor roles in the proposed architectures, we are stressing the extent of profiling and targeting that the new solutions allow and the way they violate key GDPR principles such as article 9 and recital 26.
Our empirical findings not only challenge the assumption that excluding third-party cookies would decrease consumer surveillance, but also show the potential for consumer surveillance to expand, seriously questioning the ability of those solutions to comply with the GDPR. A shift in one tracking instrument, as central as that instrument might be, cannot bridge the inherent gaps in consumer privacy caused by the structures & incentives of the online advertising market. 2
Toward (Greater) Consumer Surveillance in a ‘Cookie-less’ World:
A Comparative Analysis of Current and Future Web Tracking Mechanisms
Ido Sivan-Sevilla & Patrick T. Parham (UMD)
1 - Introduction
Ad-based monetization mechanisms of Web content are about to move away from their main tracking instrument – third-party cookies (Binns, 2022; Choi et al., 2020; O’Reilly, 2020) . Arguably a major win for privacy advocates, this would theoretically disable cross-site surveillance, enhancing consumers’ privacy online. New ‘cookie-less ID Solutions’ are extensively discussed by ad industry stakeholders, who frame them as ‘Privacy Preserving Advertisement’ (PPA) initiatives (Thomson & Rescorla, 2021) . The privacy and compliance implications of those prospective ID solutions, however, are still unclear. Are these new ‘Universal ID Solutions’ really ‘privacy-friendly’ as marketed by the industry? To what extent Ad Networks are going to change their practices and no longer rely on cross-site consumer surveillance for marketing purposes? Does it really prevent individuals from being ‘singled-out’ by future online ad solutions as clearly stressed by the GDPR?
To better understand the future of Web tracking and privacy compliance, we compare the three main cookie-less tracking architectures - The Trade Desk Unified ID 2.0, LiveRamp RampID, and Secure Web Addressability Network (SWAN) - to current cookie-based tracking mechanisms in 50 popular websites working with all four main supply-side-platform (SSP) networks - Pubmatic, OpenX, AppNexus, and Rubicon. Our comparison is based on deductive tracking analyses according to our suggested typology of tracking specifications for assessing the privacy implications of tracking technologies.
We first introduce a novel typology of tracking specifications, based on five different categories, to assess the privacy harms associated with tracking technologies. Through this typology we aim to systematically measure and reveal to what extent cookie-less tracking mechanisms differ from one another and from current cookie-based tracking with regards to privacy harms. Second, to capture current cookie-based tracking practices, we collect data from 50 popular websites that work with the four main SSPs in the industry. We investigate the extent to which those actors persistently identify users across sites based on their unique ID cookie. Those SSPs are in a position to link user identities across browsing experiences and conduct cross-site consumer surveillance. We scrape the ads.txt files in the root domains of top million popular websites (based on the tranco rank) to construct publisher networks grouped around the same SSP. To reduce false positives, we inspect the HTMLs of 50 popular websites in each SSP network to verify that those websites practically work with that SSP. Then, we trace the usage of ID cookies within and across the four SSP networks. We investigate the construction of user identities by SSPs and trace persistent identification of individuals across websites. Through various web crawlers, we collect data about tracking cookies installed via HTTP Headers, URL Parameters, and browsers’ local storages. We couple that with technical documentation from the ad industry to further 3
understand the privacy implications of cookie-based tracking. Third, we contrast our analysis of cookie-based tracking with the persistent user identification enabled by suggested cookie-less solutions. We rely on journalistic reports and technical industry documents, considering the three most dominant cookie-less tracking solutions discussed by the industry, and deductively evaluating them based on our tracking typology.
Our findings reveal how the new tracking architectures can make consumer surveillance in the Web (1) Dynamically wider, linking user identities across, not only within SSP networks; (2) More persistent over time by relying on deterministic identification data such as PIIs; and (3) Vulnerable to the integration of first- and third-party data for consumer profiling by Incentivizing advertisers to link their user data with purchased data. This could allow advertisers to secretly bid on sensitive user categories, potentially circumventing targeting restrictions of the main advertising platforms.
In contrast to existing criticism on cookie-less tracking solutions that mostly focus on deficiencies in consent mechanisms or the lack of a governing body, our study underscores the structural impact of ad networks on the potential privacy harms by the proposed tracking mechanisms. Structures and processes of ad networks lead to potentially longer, wider, and richer consumer surveillance, compared to cookie-based tracking. Our findings also question the ability of suggested tracking architectures to become GDPR compliant. Despite civil society pushback against the current cookie-based tracking ecosystem (Lomas, 2021) and the marketing of new tracking solutions as ‘privacy-preserving-advertisement’ (Thomson & Rescorla, 2021) , we show how these technologies enable granular tracking and targeting of consumers. ‘Singling out’ individuals might become easier for advertisers, given the historically rich consumer profiles that advertisers are incentivized to build. In contrast to existing debates about Universal ID solutions and GDPR compliance, that mostly focus on the lack of data controller and data processor roles in the proposed architectures ( Schiff, 2021e; Schiff; 2022a) , or the role of cookies in the suggested solutions (Asim, 2021; Asim, 2022), we are stressing the extent of profiling and targeting that the new solutions allow and the way they violate key GDPR principles such as article 9 and recital 26.
Those findings highlight how online tracking has become a deeply entrenched norm, with a shift in one tracking instrument – third-party cookies – unable to fundamentally change the behavior and user-profiling appetite of ad industry actors. Websites will still heavily rely on third-party request networks for ad delivery (Gopal et al., 2022), and the structure of ad networks is expected to stay the same, with all third-parties involved having an interest in building their own database about individuals to augment their ad bidding and delivery decisions (Martin, 2022). We ultimately argue that progress toward limiting persistent consumer identification in the digital advertising industry has not been made. Cross-site surveillance is here to stay, with the structure of ad networks and the extent of persistent user identification determining the levels of users’ online privacy and the ability of the ad industry to meaningfully comply with the law.
To evaluate the privacy implications of future tracking technologies, the next section highlights existing gaps in analyzing future tracking solutions and the importance of considering the privacy impacts of the structural components and processes of ad networks. Then, we develop five tracking specifications for evaluating the privacy harms of tracking technologies. In Section 3.1 4
we present our data collection & analysis for current cookie-based tracking. In Section 3.2, we analyze the privacy impacts of cookie-less tracking solutions. Section 3.3 summarizes the comparison of cookie-based and cookie-less tracking technologies. Section 4 discusses the implications of our findings and Section 5 concludes, detailing study limitations and future questions.
2 - Ad Networks (lack of) Privacy Compliance
Consumer privacy is gaining momentum. Revelations on how platforms use consumers’ data
(Mac & Kang, 2021) , the unprecedented wave of federal and state privacy bills (Lively, 2022) , and the on-going questioning on GDPR compliance by the advertising industry (Lomas, 2021) created pressure on browser owners to declare their end of support in the muchly criticized third-party cookies (Shein, 2021) . The anticipated shift in the main tracking instrument for targeted Ads, a market projected to grow to $525B by 2024 (Edelman, 2020) , led to a burst of alternative user ID solutions (Asim, 2021), marketed by the industry as ‘privacy friendly.’
Existing criticism on cookie-less tracking solutions mostly focus on deficiencies in consent mechanisms (Kaye, 2021a) or the lack of a governing body (UnifiedID2, 2022). A recent report from Mozilla surfaces more specific privacy concerns, identifying how new tracking solutions provide no mechanism to prevent access to users’ data (Thomson & Rescorla, 2021). Still, discussion of these cookie-less tracking proposals and their privacy implications have not placed enough emphasis on the intermediary role that advertising technology networks will play in continuing to stitch together persistent consumer identification (Asim, 2021; Asim, 2022). In contrast, we argue that the structural impact of ad networks on the potential privacy harms of new tracking technologies should be carefully analyzed.
Discussion over the ability of suggested tracking technologies to become GDPR compliant is also limited, mostly focusing on the lack of data controller and data processor roles in the proposed architectures ( Schiff, 2021e; Schiff; 2022a) , or the role of cookies in the suggested solutions (Asim, 2021; Asim, 2022). But at the center of those solutions, we argue, are the profiling and targeting capabilities enabled for ad-network actors. Those capabilities should be assessed in light of GDPR’s principles regarding data subjects’ control over their data, the processing of sensitive data, and the ability of data controllers to ‘single out’ individuals.
Thus, the structural components and processes of ad network actors should be at the center of consumer tracking and targeting evaluations. Surprisingly, those ad-networks, which facilitate almost half of Web monetization mechanisms (Choi et al., 2020), have received very little empirical attention from the Information Systems (IS) literature. Most works on privacy, including works related to online advertising, are very specific to users’ privacy concerns and consumer choices (Aguirre et al., 2015; Dinev, 2014; Dinev et al., 2013; Dinev & Hart, 2006; Goldfarb & Tucker, 2011, 2015; Hui et al., 2007; T. Li & Unger, 2012; Y. Li, 2011, 2012; Malhotra et al., 2004; Xu et al., 2011). Less is known about the privacy threats that consumers experience online and how they challenge industry compliance. Simply put, the ‘so what’ story about Web privacy remains empirically under-discussed by the IS literature and this study aims to fill some of the gap. 5
Specifically, the shift in the use of third-party cookies as an evasive tracking instrument by ad-based Web monetization mechanisms (Jones, 2020; Lavin, 2006) was not critically analyzed. We aim to shed light on potential user tracking and targeting in a cookie-less world, questioning the ability of post-cookie solutions to comply with privacy law. We intend to highlight the operations of Ad Networks, which represent a chain of third-party actors that monitor consumers’ behavior online and follow consumers across multiple websites for marketing purposes (D’Annunzio & Russo, 2020). While publishers collect information about their own visitors, it is the ad network that collects the most information to track and target consumers across the Web (Bashir et al., 2016). Those Ad Networks affect market outcomes, but research remains sparse regarding the implications of network actors’ decisions (Choi et al., 2020).
To fill those gaps and understand the structural impacts of ad networks on consumers’ privacy, we deductively analyze cookie-based and cookie-less tracking mechanisms based on a novel tracking typology below, that enables us to understand the privacy impacts of tracking technologies.
3 - Tracing Web Tracking
Following Binns (2022), we use a working definition of tracking provided by the World Wide Web Consortium (W3C)’s tracking protection group according to which ‘tracking is the collection of data regarding a particular user’s activity across multiple distinct contexts, and the retention, use, or sharing of data derived from that activity outside the context in which it occurred’ (W3C Working Group, 2019) . According to this definition, not every data collection is considered ‘tracking.’ As long as data collected within one context stays in the same spatial and temporal context, we should not consider such data collection as ‘tracking.’ However, when a party somehow collates different data points about the individual from different data sources and/or time stamps, this should count as tracking (Binns, 2022) .
The specifications of tracking vary considerably. Tracking can take place via various user identification instruments - cookies (Jones, 2020) , fingerprinting (Englehardt & Narayanan, 2016) ,favicons (Solomos et al., 2021) , personally identifiable information (PII), and etc. Tracking can be conducted by different parties (public or private) who can gain visibility on data subjects across different contexts, in different time stamps, and for different purposes - marketing, law enforcement, national security, user engagement, public health, etc. There are ample opportunities and vectors for consumers to be tracked, especially when tracking actors own various services with which consumers directly engage - from search engines and platforms to mobile software, hardware, websites, and for our purposes in this paper, ad network functions.
We focus specifically on tracking as an opaque commercial practice for online consumers on the Web, emerged through symbiotic relationships between websites and third parties (Gopal et al., 2022) , and facilitated through automated data capture, making ‘passive’ commercial surveillance almost inevitable for individuals on the Web. 6
To evaluate and assess the extent of tracking, we develop a framework for tracking specifications that details what we view as the most important criteria for measuring the impact of tracking technologies on consumers’ privacy. We suggest the following five tracking specifications:
[1] User Identification Instrument : How can individuals be identified? What is the tracking instrument through which trackers assign user IDs and potentially follow users and collate data points for profiling purposes (i.e. cookies, tokens, fingerprints, favicons).
[2] Cross-context User Visibility: Which companies can persistently identify consumers across sites? We are interested to understand who are the ad network actors that can identify and ‘enjoy’ visibility over consumers across the Web (i.e., SSPs, DSPs, advertisers, publishers, Ad Exchanges).
[3] Longitudinal Tracking: Does the tracking technology enable consumers to be tracked over time?
[4] Circumvention of Targeting Restrictions: Does the tracking mechanism enable/incentivize advertisers to build rich first-party consumer profiles and hiddenly escape targeting restrictions by advertising platforms?
[5] User Data Sources: Can we limit participating data actors for profiling purposes? What are the possible sources for data collation? Which data points about the user can be gathered and used for targeting purposes? (i.e. current browsing behavior, past browsing behavior, offline data from digital footprints, first party data, third party data).
Through a deductive analysis based on the tracking specifications above, sections 3.1 & 3.2 evaluate cookie-based and cookie-less tracking mechanisms. Section 3.3 compares the mechanisms, revealing how cookie-less tracking alternatives are likely to enable greater consumer surveillance than current, cookie-based, tracking practices.
3.1. Cookie-based Tracking
In order to contrast future cookie-less tracking mechanisms with current cookie-based tracking by ad networks, we first aim to reach a comprehensive understanding of current tracking practices. To do so, we rely on browser-side observations and study the ability of ad networks to persistently identify and potentially track users across websites.
To assess the usage of persistent identifiers within and across SSP networks, we assembled a list of publisher sites that work with all four main SSP networks in the advertising industry. The four SSPs selected were based on the authors’ familiarity with the networks’ prominent position in the digital advertising industry. We developed two different selection criteria to assemble a final list of publisher sites. For the first selection criteria, inclusion of a publisher was based on the publisher site listing the SSP in their ‘ads.txt’ file. With this criterion met, we then ranked sites based on the ‘Tranco’ popularity index (Tranco, n.d.) . Our initial assembled list of the top 100 popular websites that list all four main SSPs in their ads.txt file was crawled by visiting their landing pages (and not inner-site pages) per site. Analyzing the initial results, we saw that 7
certain publishers did not register cookies from all four SSPs. While ads.txt indicate which partners are eligible to sell a publishers ad inventory, based on the instances observed in crawling 100 sites we assume that just because a partner is listed does not mean that the publisher is currently working with the partner. In the second selection criteria, inclusion of a publisher was based on manually checking in the Chrome browser Developer Tools - Cookies Storage table, that a cookie from each of 4 SSPs registered in at least 1 of 10 refreshes of the landing page. We were able to develop a list of 50 sites from checking sites in the ‘Tranco’ popularity index that list all four SSPs in their ads.txt file (see appendix #1). From this manual check that indicated that a cookie does not always register when a landing page is loaded, we decided to crawl the 50 sites 10 separate times and average the results. In the 10 separate crawls, we again visited only the publisher landing pages.
To evaluate the usage of persistent user identifiers by trackers among the four SSP networks we traced stateful tracking via HTTP cookies, as they are still the most dominant technique to identify online users across websites (Roesner et al., 2012; Fouad et al., 2020). To collect tracking information (i.e HTTP cookies, JavaScript Operations, and HTTP headers) of each publisher site we used an open source-based automated web crawler – OpenWPM - that simulates real users’ activity and records website responses, metadata, cookies used, and scripts executed (Englehardt and Narayanan, 2016). We performed 10 individual stateful crawls and set the crawl to use only one browser instance. We did not set the “Do Not Track” setting in the configuration to allow for persistent identification and configured the simulated browser to accept all 3rd-party cookies. We also used the “bot detection mitigation” to scroll randomly up and down visited pages. We set sleep time between publisher sites to five seconds, and timeout between websites to 100 seconds. The 10 crawls were run on a local machine on September 20 th & 21 st , 2022, and data were recorded in a SQLite database.
Within the SQLite database file created by the crawl, the ‘http_requests’, ‘javascript’, and ‘javascript_cookies’ tables were analyzed. SSPs were recognized based on the ‘host’ field and filtered based on rows containing the SSP network name. Publisher site was established based on the ‘top_level_url’ field. The persistent identifier was recognized based on a concatenation of the cookie name and cookie value fields. For the crawl of publisher sites, stateful tracking was enabled and we traced the cookie storage of our browser to analyze how the same cookie IDs were used across websites within and across SSP networks. We refer to the SSPs that use the same cookie ID value in two or more publisher sites as ‘persistent identifiers,’ as they are identically identifying the user across the sites they work with. We identified SSP ID cookies based on the syntax included in each company’s privacy policy (See Table 1 below) (Magnite, 2021a; OpenX, 2022; Pubmatic, 2020; Xandr, 2022). 8
SSP Network ID Cookie Name
Pubmatic KADUSERCOOKIE
OpenX i
AppNexus uuid2
Rubicon khaos
Table 1: ID Cookie Names Used for Tracing Persistent Identification of Users Per SSP Network
In our data analysis, we wanted to first observe whether certain SSPs persistently identify users across publisher sites within their networks. We observed that SSPs utilized a persistent identifier in 77.6-90% of sites (See Table 2 below). At the same time, we acknowledge that the persistent identification of users is happening on the server-side as well (e.g., Acar et al., 2014), in ways that are more challenging for researchers to detect. Hence, we expect our results to be considered as a lower bound on the amount of persistence identification of users by SSPs across the crawled websites. A full visualization of persistent identification patterns within SSP networks, by site, can be found in appendix #2.
SSP Network
Average percentage of websites across which our browser was persistently identified
Pubmatic 86.4%
OpenX 77.6%
AppNexus 90%
Rubicon 90%
Table 2: Amount of Persistent Identification Across Websites Per Crawled SSP Network
Second, we wanted to see the overlap of persistent identifiers across SSPs. Can we spot the same ID cookie value by two or more SSPs, hinting that those SSPs are dynamically identifying individuals in the same way to potentially link user data across the sites they work with? Interestingly, from our browser-side observations, we saw how in all cases, despite one instance recurring across all 10 individual stateless crawls we could not fully explain (see appendix #3), persistent identification across SSP networks was non-existent. We could not trace the same cookie ID or any cookie value being dropped on our browser by two different SSPs. We acknowledge that cookie-syncing between SSPs might happen on the server-side, away from our crawler, but competition considerations between those actors make it unlikely. Advertisers choose to work with multiple SSPs to bid on users and base the selection of SSP partners primarily on ability to reach different audience sizes by publisher unique monthly visitors and delivery of accurate inventory performance (Sluis, 2018; Vargas, 2022a). These considerations related to 9
expanding the addressable audience and differentiation in services offered by SSPs demonstrate clear competition between SSPs that would disincentivize cookie sharing.
In summary, going back to our tracking specifications, we observed how users can be identifie dvia third-party cookies placed by supply-side networks (SSPs) on publisher sites. This provides SSPs with cross-context visibility when using ID cookies to persistently identify and potentially track individuals across the websites they are embedded in. Still, our data show that SSP-based identification of users remains within the network of publisher sites per SSP, and not crossing SSP networks. The observed tracking method also enables tracking over time , as long as consumers can be identified or linked to the same cookie.
Going beyond our data collection and relying on our review of the industry’s publications and technical documents, third-party cookies also enable advertisers to learn about users’ browsing history and past behavior for bidding on publishers’ ad inventory across sites. Importantly, cookies can be linked to other data sources, allowing advertisers to pair their data with cookies, beyond what passive surveillance of browsing behavior can provide, potentially circumventing targeting restrictions by advertising platforms. Advertisers are able to match third-party cookie identifiers to data purchased from data brokers to target users based on potentially sensitive assembled categories (Experian, n.d.; Sherman, 2021). In terms of user data sources , minimal limitations are currently in place. Even though we have not directly measured potential participants in user tracking, previous studies showed how a range of ad networks actors, online and offline, can contribute to profiling consumers for marketing purposes (Choi et al., 2020; Wei et al., 2020) .
3.2. Cookie-less tracking
Following Google’s announcement in 2020 that the company would no longer support third-party cookies within its digital advertising products, first-party data collection from publisher sites began to arise as the digital advertising consensus to preserve some of the functionality provided by third-party cookies to tracking and targeting in the programmatic bidding process (Schuh, 2020; Southern, 2020). Publishers began to better organize and intensify user identification in first-party held data through means such as subscriptions, non-paying subscriber registration, and newsletters sign-ups (Asim, 2021). While first-party user data segments have used targeting categories constructed from demographic and on-site behavioral data to provide users with relevant advertisements directly on individual publisher websites, advertisers are still interested in capturing information about user behavior across the web that resembles tracking from cookie-based persistent identification patterns. As the preservation of precision in targeting relevant audiences that occurs from being able to follow behavior across sites is still a top priority for advertisers, SSPs were initially thought to potentially be a coordinating body that could aggregate first-party data across publisher sites so that advertisers could still have access to relevant audiences across publisher sites (Joseph, 2021). However, alternative solutions, labeled under the umbrella terms ‘Universal IDs’ and ‘Alternative IDs,’ have started to proliferate that work across multiple SSPs networks, enabling the identification of individual users more persistently than current, SSP-partitioned identification. 10
In order to compare proposals to existing cookie-based solutions, we have selected what we consider to be the three primary alternative identifier solutions - ‘The Trade Desk Unified ID 2.0 (UID 2.0),’ ‘LiveRamp RampID,’ and ‘Secure Web Addressability Network (SWAN).’ Our selection was informed by our reading of trade publications covering the digital advertising industry and the frequency of reporting on specific solutions. We have incorporated the reporting into the data we have collected along with messaging and technical documentation that the authors of the different solutions have made public. Hence, based on our suggested tracking specifications in section 3, we analyze each of the three main cookie-less tracking solutions discussed by the industry.
In terms of the instrument of user identification , the three solutions vary. For SWAN, an individual is identified when they first visit a publisher site that has adopted the solution (Asim, 2022; Schiff, 2021d; Thomson & Rescorla, 2021). Upon loading the publisher page, the user is presented with a pop-up asking for consent to show personalized advertising on the current site and other sites that have adopted the solution. As part of the pop-up, the user also has the option to share their email address which can serve as an identifier. Regardless if the user accepts the option to receive personalized advertising with or without sharing their email address, a first-party cookie is placed by the initially visited publisher site within the SWAN network that creates a pseudonymous identifier that is stored on the user’s browser. For UID 2.0, the solution authored by a top demand-side platform (DSP), user identity is established by logging via emails into publisher sites. Publishers store the email address in a first-party cookie placed on the page (Asim, 2022; Thomson & Rescorla, 2021; UnifiedID2, 2022). The email is matched to a UID2.0 through a corresponding token is also created and is used for encryption of the UID2.0 and can be decrypted only by partners that receive a decryption key through agreeing to the Unified ID 2.0’s terms of service. The creation of the pseudonymous UID2 is managed by the UID2.0 service. The third solution under analysis, The LiveRamp RampID, is distinct from the other previous two, in that it is interoperable with other ID solutions, including UID2.0 (Asim, 2022). The RampID uses user email addresses that are matched with email addresses that are shared with publishers in exchange for content (Asim, 2022; LiveRamp, 2022c). Instead of placing a first-party cookie, LiveRamp matches IDs in the ecosystem through its proprietary authenticated traffic solution (ATS) (Asim, 2022; LiveRamp, 2022d). The ability of this solution to further identify users in the ecosystem is also attached to other offline PII (phone number, address history), based on information that advertisers can match with assembled first, second and third-party data.
Regarding cross-context user visibility, we see that access to SSPs and their publisher partners still gives advertisers the opportunity to reach users across sites. The orientation of an ID solution recreates that of the existing SSP to publisher tracking. While partnering with primary SSPs to coordinate the identifier programmatic bidding is transacting on, these solutions are also recreating an identifier to recognize users across SSPs (Asim, 2022). In Section 3.1, we showed how current cookie-based identification of users seems partitioned by individual SSPs. For the analyzed cookie-less tracking solutions, however, we argue that not only are these networks replicated across publisher sites in SSP partnerships, but also the integration of ID solutions by industry leading SSPs could potentially lead to identification of users across SSPs, not only within SSPs. As a result, greater identification of users across publisher sites will be enabled. All three analyzed solutions have been supported and committed adoption from primary SSPs, and The 11
Trade Desk Unified ID 2.0 has announced partnerships with all four SSPs we analyzed as part of our data crawl (Asim, 2022; Schiff, 2020a; Schiff, 2020b; Schiff, 2020c; Schiff, 2021a; Schiff, 2021d). This delegates a large amount of responsibility to those that will be assigned with governing these solutions, as the view of consumer behavior across the web will be expanded.
Longitudinal tracking is clearly enabled by all three cookie-less tracking solutions. The persistence derives from deterministic data serving as the basis for each solution (Asim, 2022; Kaye, 2021b). Each offers users the option to opt out of personalized advertising universally from all participating partners (Asim, 2022; LiveRamp, 2022b; SWAN-community, 2021; UnifiedID2, 2022). An argument can also be made that persistence can last for a greater duration than third-party cookies, which users frequently delete, as the choice to opt out could mean the user losing access to publisher content. While advertisers are able to bid on users that are classified to a persistent identifier, these specific solutions have not yet answered whether they will support the most persistent targeting method, retargeting or remarking - the following of a user after an initial action that classifies them as potentially more likely to carry out a desired action attributed to a digital advertisement.
Interestingly, all three solutions further encourage circumvention of targeting restrictions by advertising platforms. Advertisers can persistently identify users based on their first-party data, easily overriding targeting policies by main advertising platforms. We have previously alluded to the responsibility of those in charge of governing the new tracking solutions as they provide for greater persistent identification across the web. But what we highlight here is an under-discussed gap in cookie-less proposals. Advertisers can now enjoy an increased role in persistent identification through not just relying on publisher first-party data but also on the ability of the advertisers themselves to upload data to be encoded for targeting. Within The Trade Desk Unified ID 2.0, for instance, the company mentions ‘First-Party Relationships’ capabilities, where an advertiser is able to upload first-party data to be encoded to the UID2 for activation across publisher sites (UnifiedID2, 2022). Similarly, LiveRamp offers advertisers the opportunity to ‘onboard’ their data where PII can be uploaded in order to be converted into RampIDs and organized by segment so that they can be activated in more than 500 different partner platforms (LiveRamp, 2022a). SWAN does not provide a great deal of information related to these capabilities, but states on its homepage that it is ‘Complementary to CRM data’ (Secure Web Addressability Network (SWAN), 2021). This capability, enabled by all three solutions, was found in other alternatives to create the potential to further obfuscate advertiser targeting practices.
Such capability is concerning based on the work that can be done by advertisers before PII is encoded by identity solutions for targeting. As stated in Section 3.1, advertisers also choose to target users beyond the signals from cookies and purchase third-party data to expand profiling of users beyond what passive surveillance of browser behavior can provide. The emphasis on first-party data in cookie-less solutions has encouraged advertisers to begin leveraging their existing customer base to find other users that resemble the traits of current customers through modeling practices to create ‘look-alike’ segments (LiveRamp, 2020a). Compared to advertisers’ efforts in the current cookie-based ecosystem, and with the level of granularity enabled by targetable information IDs still unclear, advertisers are now investing more in technology that can match first-12
party subscriber data to other datasets (Vargas, 2022b). This work is carried out on an ‘identity graph’ that allows advertisers to manage individual-level data and encode the data through an ID method of choice, providing a centralized system to merge online and offline identifiers into a consolidated profile to pair with purchased third-party data and activate selected audiences with multiple partners (LiveRamp, 2020c). An identity graph poses a threat to individual privacy as it allows advertisers to develop rich profiles of existing customers through pairing data purchased from data brokers but also non-customers across the open web (Vargas, 2022b). The Unified ID 2.0 and LiverRamp RampID now enable alarming ID linkage capabilities in this technology (Schiff, 2021b; Schiff, 2021c; Schiff, 2022b; The Trade Desk, 2021c). When an advertiser is able to upload a list of IDs by segment, there is the potential that advertisers are not only uploading explicit lists of existing customer PII but also emails that have been grouped according to specific audience segmentation categories. As these segments are constructed using data purchased from data brokers, there is a chance the construction of segments and profiles could involve sensitive categorical information. These IDs aim to be interoperable across major platforms, and the inability to verify or specify the segmentation practices of PII being converted to transact in major digital advertising platforms creates the opportunity for advertisers to violate targeting practices specific to these platforms (Google, 2022; Meta 2021; Twitter, 2022). There is no mechanism to verify how advertisers have segmented first-party data before importing into these systems, and new ID solutions makes sensitive population segmentation and bidding very attractive for advertisers.
For example, a credit card company might want to target African-American men specifically, but is unable to do so though the restricted targeting categories on certain primary digital advertising platforms. Instead, the credit card company could assemble a list of prospective customers based on data broker acquired data that can tie email address to race. If the credit card company chose to encode their data with the LiveRamp ID, for instance, and activate on any of the 500 partner platform destinations the company claims to partner with, the company just provides a link to certain companies’ targeting restrictions without a thorough review (LiveRamp, 2022a, LiveRamp, 2022b). Previous research has identified instances where advertiser-uploaded lists have violated platform policies when targeting users, using categories such as race, religion, politics, sex life, or health (Wei et al., 2020). We argue that new tracking ID solutions place a significant amount of importance on leveraging first-party data that encourages advertisers to look for ways to link online and offline data through ID solutions, allowing for targeting practices with little to no oversight by primary platforms. This capability encourages advertisers to profile users before even interacting with primary platforms, circumventing existing platform restrictions.
Regarding user data sources used for tracking , we found that similar to cookie-based tracking, the limitation of participants allowed to utilize and transact on these IDs remains unclear. The basic structure of third-parties responsible for identifying individuals remains almost identical. A different party now provides the persistent identifier and advertisers still pass IDs to bid on individuals across sites through demand-side platforms. Providing general terms detailing how participants are expected to adhere to certain principles, promises of outlining a code of conduct, or keeping proprietary governance privileged, it is unclear what other members of the current ecosystem can participate (Asim, 2022; LiveRamp, 2022c; Thomson & Rescorla, 2021; SWAN-13
community, 2021; UnifiedID2, 2022). This issue has not been resolved, even though cookie-less tracking solutions claim to provide great individual privacy through limiting the sharing of persistent individual-level information to various parties. We found that parties responsible for governing the solutions are not clear nor are the terms that participants are expected to follow. The SWAN solution will be governed by the SWAN Network itself which has outlined a set of ‘Model Terms’ that details how participants are expected to adhere to information sharing practices (Thomson & Rescorla, 2021; SWAN-community, 2021). The Trade Desk claims that the company plans to turn over control to an ‘Administrator,’ a role that has not yet been filled (Asim, 2022). The Trade Desk documentation further mentions a ‘code of conduct’ that participants must follow, but is not currently available (UnifiedID2, 2022). The role of the administrator was originally supposed to be maintained by the Interactive Advertising Bureau, but the organization has chosen to no longer pursue supporting the solution in this capacity (Katsur, 2022; Mitchell, 2021). Prebid has also declined to serve as the administrator (Shields, 2022). Both SWAN and the Unified ID 2.0 are open source, but LiveRamp ID is a proprietary solution that is managed by the company itself (Asim, 2022; LiveRamp, 2022c). Based on the lack of transparency, it is not clear how these solutions will enforce standards to ensure participants do not violate terms of operation that have not been fully defined or made public.
3.3. Comparing cookie-based and cookie-less tracking mechanisms
Table 3 below summarizes our comparison between current and future Web tracking solutions by the advertising industry. The table shows how consumer surveillance is expected to expand in terms of (1) wider persistent identification patterns, (2) potentially spanning tracking across larger time periods, and (3) incentivizing advertisers to circumvent targeting restrictions and bid on consumers based on sensitive and richer first-party data profiles.
Our findings show how surveillance on the web is in the process of increasing from its current cookie-based organization to one transacting on ID solutions. First, persistent identification has increased as identification is not limited to individual SSP networks. Second, the source of data responsible for determining identity is derived from PII and consent mechanisms, making it difficult for individuals to opt out of an exchange for content that enables tracking. Third, the capability to upload data has encouraged advertisers to find ways to segment audiences before interacting with primary platforms. This has resulted in the purchase of offline data from data brokers that can be paired with existing customer data or used to derive potential customers based on similarity to those current customers or other preferred traits. This practice allows for the selection of traits that can violate the policies of primary platforms that prohibit certain sensitive categories to be used for targeting of users. The violation of platform policy is a result of not having a method to thoroughly verify how individuals are encoding generated segments to ID solutions before activating in platforms. 14
Cookie-based Tracking
Cookie-less Tracking
SSP Cookie IDs UID 2.0 SWAN LiveRamp
User Identification Instrument
Passive placing of third-party ID cookies
First-party cookie based on consent obtained on publisher site & sharing of email
First-party cookie based on consent obtained on publisher site & sharing of email
Proprietary authenticated traffic solution based on consent on publisher site & sharing of email merged with offline name, address, and phone number
Cross-site User Visibility
Within all four SSPs, across 77%-90% of sites in a network, but not across SSPs
Across, not only within all four SSP networks
Across, not only within all four SSP networks
Across, not only within all four SSP networks
Longitudinal Tracking Yes Yes - with reliance on more deterministic data
Yes - with reliance on more deterministic data
Yes - with reliance on more deterministic data
Circumvention of Targeting Restrictions
Yes Yes - and encourage greater profiling by pairing advertiser 1st-party data to purchased 3rd-party data
Yes - Alludes to the ability to pair CRM data
Yes - and encourage greater profiling by pairing advertiser 1st-party data to purchased 3rd-party data
User Data Sources No limitations are in place
Limitations on participants and governance mechanisms still unclear
Limitations on participants and governance mechanisms still unclear
Limitations on participants and governance mechanisms still unclear
Table 3 - Tracking Specifications for marketing purposes via cookies and cookie-less solutions 15
4 - Discussion
Based on data collection from the four main SSP actors on cookie-based tracking, and analysis of technical documentation from the advertising industry on cookie-less tracking solutions, our study highlights three tracking specifications through which the main cookie-less tracking solutions are expected to increase consumer surveillance on the Web. First, persistent identification of consumers across sites is likely to cross SSP networks, creating potentially greater real-time visibility on consumers. Second, identification mechanisms are expected to rely on PIIs, making them more persistent and likely to better track users over time. Third, with the pivoting of the industry toward ‘identity graphs’ and first-party data, advertisers are now incentivized to target consumers based on rich first- and third-party data profiles, potentially overriding existing targeting restrictions in case advertisers choose to target consumers based on sensitive and forbidden categories. Hence, consumer tracking on the Web is likely to become dynamically wider, involve richer consumer histories, and rely on a greater variety of data sources.
In contrast to existing criticism on cookie-less tracking solutions, our study shows the structural impact of ad networks on the potential privacy harms by the proposed tracking mechanisms. We systematically contrast future solutions with current, operational, tracking mechanisms based on our developed tracking specifications. We look beyond deficiencies in consent mechanisms (Kaye, 2021a; Thomson & Rescorla, 2021) or lack of a clear governing body (UnifiedID2, 2022) to show how structures and processes of ad networks lead to potentially longer, wider, and richer consumer surveillance, compared to current tracking mechanisms.
Our findings pose serious questions on the ability of the suggested tracking mechanisms to become GDPR compliant. Despite civil society pushback against the current cookie-based tracking ecosystem (Lomas, 2021) and the marketing of new tracking solutions as ‘privacy-friendly’ (The Trade Desk, 2021a; Thomson & Rescorla, 2021), we show how the new mechanisms can practically enable granular tracking and targeting of individuals based on sensitive categories. Moreover, ‘singling out’ individuals might become easier by advertisers, given the historically rich consumer profiles that advertisers are incentivized to build (GDPR’s recital 26). In contrast to existing debates about Universal ID solutions and GDPR compliance, that mostly focus on the lack of data controller and data processor roles in the proposed architecture (Schiff, 2021e ; Schiff; 2022a ), or the role of cookies in the suggested solutions (Asim, 2021; Asim, 2022), we would like to stress the extent of profiling and targeting that the new solutions allow, and their possible violations of key GDPR principles.
First, the potential richness of personal information involved in the projected profiling tactics may go against or beyond consumers’ reasonable expectations and might infringe applicable data protection principles and rules. When an advertising company, for instance, joins its own first-party data with third-party data sources, as described in Section 3.2, this may result in personal data being used beyond their initial purpose and in ways the individual could not reasonably anticipate. The profiles built by the advertiser might involve an inference of interests or characteristics which individuals had not actively disclosed, undermining the ability of individuals to exercise control over their personal data (EDPS, 2018) , or creating an opportunity for advertisers to ‘single out’ individuals from their data (GDPR’s recital 26). Moreover, GDPR’s 16
transparency requirements might be violated as the role of different parties in this process is probably unclear to the user given the potential circumvention of platforms’ targeting restrictions (see Section 3.2 for an example).
Second, the suggested tracking mechanisms might encourage discrimination and exclusion. Targeting by advertisers can involve criteria that, directly or indirectly, have discriminatory effects relating to an individual’s racial or ethnic origin, health status or sexual orientation, or other protected qualities of the individual concerned. The potential for discrimination in targeting arises from the ability for advertisers to leverage the extensive quantity and variety of personal data given the pivoting of new tracking technologies around deterministic identification. As explained in Section 3.2, the ability of advertisers to link their own segmentation of users with ID solutions enables granular and possibly discriminatory targeting that overrides existing, non-enforceable, targeting restrictions, and thus, violates requirements in GDPR’s Article 9 about processing special categories of personal data.
5 - Conclusion
The privacy implications of ad networks structures and processes lead to a privacy-concerning ad landscape even without third-party cookies. The implementation of cookie-less tracking solutions by the AdTech complex potentially enables greater dynamic visibility on consumers, longer consumer tracking, and the assembling of more sensitive consumer profiles.
Looking ahead, even though there is not a clear primary tracking solution in the group of offerings, and with industry leaders not believing that one solution will be responsible for continued persistent identification alone, the Trade Desk Unified ID 2.0 has emerged as a leading solution. Despite uncertainty about whether the Trade Desk Unified ID 2.0 and other solutions will be compliant with evolving privacy regulation due to the similarities between the proposal and existing tracking methods, support continues to amass for the Trade Desk based on marketing language that emphasizes privacy but also promises to deliver on the same capabilities provided by third-party cookies (Asim, 2021; The Trade Desk, 2021a). The criticism of the Trade Desk UID could be viewed as outweighed by not only industry support from agencies such as IPG, Omnicom, and Publicis Groupe (Bürgi, 2021a; Bürgi, 2021b; The Trade Desk, 2021b; The Trade Desk, 2021c) but also by Google. While Google initially announced in 2021 that it would not support email or alternative third-party identifiers in its ecosystem due to their belief that the solution would not be sustainable in the evolving regulatory environment (Temkin, 2021), the company has reserved course through the introduction of encrypted signals from publishers (ESP) product that allows publishers to share encrypted first-party data and alternative identifiers via Google’s Ad Manager (Schiff, 2022c; Google Ad Manager Help, 2022). With this backing, the Trade Desk now has support from the parties with high interests in maintaining budgets and practices in programmatic advertising.
Given our findings in this study, this is an alarming trend. The implementation of UID 2.0 by the advertising industry might make consumer surveillance even worse. As long as the structure of ad networks will not fundamentally change, consumers’ profiling will persist as a highly valued commodity. 17
Our study has a few limitations. First, for our evaluation of cookie-based mechanisms, we crawled 50 landing pages of popular websites and not their inner pages, where tracking is known to be more pervasive. Second, we compare cookie values based on identical strings, even though some actors might encrypt or hash their cookie values, making us miss some persistent identification trends. Third, we cannot trace information being shared on the server side and fully capture the amount of detail ad network actors have about the individual user. The three limitations lead us to assume that our findings on cookie-based tracking suggest a lower bound for actual persistent identification patterns. Previous studies support our assumption, acknowledging that advertising actors often match first-party user data with the observed cookie IDs to achieve more granular targeting capacities (Trusov et al., 2016).
Future follow-up research projects could invest in understanding how compliance to privacy initiatives is being shifted to the individual advertiser level, as demonstrated in Section 3.2. Advertising platforms create policies to restrict sensitive categories from targeting, but at the same time still giving advertisers the capability to target users with precision. Those loosely-enforceable mechanisms are part of the privacy problem of ad-networks and a structural change is much needed. 18
References
Acar Gunes, Christian Eubank, Steven Englehardt, Marc Juarez, Arvind Narayanan, and Claudia Diaz. (2014). “The Web Never Forgets: Persistent Tracking Mechanisms in the Wild.”
Proceedings of the ACM Conference on Computer and Communications Security (CCS).
Aguirre, E., Mahr, D., Grewal, D., de Ruyter, K., & Wetzels, M. (2015). Unraveling the Personalization Paradox: The Effect of Information Collection and Trust-Building Strategies on Online Advertisement Effectiveness. Journal of Retailing , 91 (1), 34–49. https://doi.org/10.1016/j.jretai.2014.09.005
Asim, A. (2021, November 17). Digiday Media Research:A comprehensive guide to third-party cookie alternatives. Digiday https://digiday.com/media/a-digiday-media-guide-to-third-party-cookie-alternatives/
Asim, A. (2022, June 9). Digiday+ Research: A guide to the top 10 ID alternatives for publishers.
Digiday. https://digiday.com/media/digiday-research-a-guide-to-the-top-10-id-alternatives-for-publishers/
Bashir, M. A., Arshad, S., Wilson, C., & Robertson, W. (2016). Tracing Information Flows Between Ad Exchanges Using Retargeted Ads. Proceedings of the 25th USENIX Security Symposium August 10–12, 2016 • Austin, TX , 17.
Binns, R. (2022). Tracking on the Web, Mobile and the Internet of Things. Foundations and Trends® in Web Science , 8(1–2), 1–113. https://doi.org/10.1561/1800000029
Bürgi, M. (2021a, August 10). Unified ID 2.0 quietly amasses more support from the agency world, but publishers aren’t as convinced. Digiday. https://digiday.com/media/unified-id-2-0-quietly-amasses-more-support-from-the-agency-world-but-publishers-arent-as-convinced/
Bürgi, M. (2021b, November 17). Omnicom Media Group formally endorses UID 2.0 in a bid to move the post-cookie future forward. Digiday. https://digiday.com/marketing/omnicom-media-group-formally-endorses-uid-2-0-in-a-bid-to-move-the-post-cookie-future-forward/
Choi, H., Mela, C. F., Balseiro, S. R., & Levy, A. (2020). Online Display Advertising Markets: A Literature Review and Future Directions | Information Systems Research. Information Systems Research , 2(31), 556–575. https://doi.org/10.1287/isre.2019.0902
D’Annunzio, A., & Russo, A. (2020). Ad Networks and Consumer Tracking. Management Science , 66 (11), 5040–5058. https://doi.org/10.1287/mnsc.2019.3481
Dinev, T. (2014). Why would we care about privacy? European Journal of Information Systems ,
23 (2), 97–102. https://doi.org/10.1057/ejis.2014.1 19
Dinev, T., & Hart, P. (2006). An Extended Privacy Calculus Model for E-Commerce Transactions. Information Systems Research , 17 (1), 61–80. https://doi.org/10.1287/isre.1060.0080
Dinev, T., Xu, H., Smith, J. H., & Hart, P. (2013). Information privacy and correlates: An empirical attempt to bridge and distinguish privacy-related concepts. European Journal of Information Systems , 22 (3), 295–316. https://doi.org/10.1057/ejis.2012.23
Edelman, G. (2020, October 5). Ad Tech Could Be the Next Internet Bubble. Wired .
https://www.wired.com/story/ad-tech-could-be-the-next-internet-bubble/
EDPS. (2018). EDPS Opinion on online manipulation and personal data .
https://edps.europa.eu/sites/edp/files/publication/18-03-19_online_manipulation_en.pdf
Englehardt S. and A. Narayanan. (2016). “Online Tracking: A 1-million-site Measurement and Analysis.” Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security , pp. 1388-1401.
Experian. (n.d.). ConsumerView: Data by the Numbers .https://www.experian.com/content/dam/marketing/na/assets/ems/marketing-services/documents/infographics/consumerview.pdf
Fouad I., N. Bielova, A. Legout, N. Sarafijanovic-Djukic. (2020). “Missed by Filter Lists: Detecting Unknown Third-Party Trackers with Invisible Pixels.” Available in: https://arxiv.org/abs/1812.01514
Goel, Vinay. (2021, June 24). An updated timeline for privacy sandbox milestones. The Keyword . https://blog.google/products/chrome/updated-timeline-privacy-sandbox-milestones/
Google Ad Manager Help. (2022, March). Share encrypted signals with bidders (Beta). G oogle Ad Manager Help . https://support.google.com/admanager/answer/10488752
Goldfarb, A., & Tucker, C. (2011). Online Display Advertising: Targeting and Obtrusiveness.
Marketing Science , 30 (3), 389–404. https://doi.org/10.1287/mksc.1100.0583
Goldfarb, A., & Tucker, C. E. (2015). Standardization and the Effectiveness of Online Advertising. Management Science , 61 (11), 2707–2719. https://doi.org/10.1287/mnsc.2014.2016
Google. (2022). Personalized advertising - Advertising Policies Help .https://support.google.com/adspolicy/answer/143465?hl=en
Google Ad Manager Help. (2022, March). Share encrypted signals with bidders (Beta). G oogle Ad Manager Help . https://support.google.com/admanager/answer/10488752
Gopal, R. D., Hojati, A., & Patterson, R. A. (2022). Analysis of third-party request structures to detect fraudulent websites. Decision Support Systems , 154 , 113698. https://doi.org/10.1016/j.dss.2021.113698 20
Hui, K.-L., Teo, H. H., & Lee, S.-Y. T. (2007). The Value of Privacy Assurance: An Exploratory Field Experiment. MIS Quarterly , 31 (1), 19–33. https://doi.org/10.2307/25148779
Jones, M. L. (2020). Cookies: A legacy of controversy. Internet Histories , 4(1), 87–104. https://doi.org/10.1080/24701475.2020.1725852
Joseph, S. (2021, November 17). ‘They will need to use multiple routes’: Shifts appear in the publisher-SSP union, as alternative identifiers proliferate. Digiday. https://digiday.com/media/they-will-need-to-use-multiple-routes-shifts-appear-in-the-publisher-ssp-union-as-alternative-identifiers-proliferate/
Kaye, K. (2021a, March 24). Third-party cookie replacements fall short of consent and transparency promises . Digiday. https://digiday.com/media/third-party-cookie-replacements-fall-short-of-consent-and-transparency-promises/
Kaye, K. (2021b, April 1). WTF is the difference between deterministic and probabilistic identity data?. Digiday. https://digiday.com/media/wtf-is-the-difference-between-deterministic-and-probabilistic-identity-data/
Katsur, A. (2022, February 28). Tech Lab Update on UID2.0. IAB Tech Lab .https://iabtechlab.com/blog/tech-lab-update-on-uid2-0/
Lavin, M. (2006). Cookies: What do consumers know and what can they learn? Journal of Targeting, Measurement and Analysis for Marketing , 14 (4), 279–288. https://doi.org/10.1057/palgrave.jt.5740188
Li, T., & Unger, T. (2012). Willing to pay for quality personalization? Trade-off between quality and privacy. European Journal of Information Systems , 21 (6), 621–642. https://doi.org/10.1057/ejis.2012.13
Li, Y. (2011). Empirical Studies on Online Information Privacy Concerns: Literature Review and an Integrative Framework. Communications of the Association for Information Systems , 28 .https://doi.org/10.17705/1CAIS.02828
Li, Y. (2012). Theories in online information privacy research: A critical review and an integrated framework. Decision Support Systems , 54 (1), 471–481.
https://doi.org/10.1016/j.dss.2012.06.010
Lively, T. K. (2022, July 7). US State Privacy Legislation Tracker .
https://iapp.org/resources/article/us-state-privacy-legislation-tracker/
LiveRamp. (2020a, March 31). Look-alike Modeling: The What, Why, and How.
https://liveramp.com/blog/look-alike-modeling-the-what-why-and-how/
LiveRamp. (2020b, August 11). Platform-Specific Distribution Information .https://docs.liveramp.com/connect/en/platform-specific-distribution-information.html 21
LiveRamp. (2020c, August 28). What’s the Difference between a DMP and an Identity Graph?.
https://liveramp.com/blog/difference-between-dmp-data-management-platform-identity-g
raph/
LiveRamp. (2022a, March 8). Onboarding Your Data .https://docs.liveramp.com/connect/en/onboarding-your-data.html
LiveRamp. (2022b, May 6). Consumer Requests for Opt-Outs, Data Access, or Data Deletions .https://docs.liveramp.com/connect/en/consumer-requests-for-opt-outs,-data-access,-or-data-deletions.html
LiveRamp. (2022c, June 30). RampID Methodology .https://docs.liveramp.com/connect/en/rampid-methodology.html
LiveRamp. (2022d, July 21). Authenticated Traffic Solution . https://docs.liveramp.com/privacy-manager/en/authenticated-traffic-solution.html
Lomas, N. (2021, June 16). Adtech ‘data breach’ GDPR complaint is headed to court in EU.
TechCrunch .
https://social.techcrunch.com/2021/06/16/adtech-data-breach-gdpr-complaint-is-he aded-to-court-in-eu/
Mac, R., & Kang, C. (2021, October 3). Whistle-Blower Says Facebook ‘Chooses Profits Over
Safety.’ The New York Times .
https://www.nytimes.com/2021/10/03/technology/whistle-blower-facebook-frances-haugen.html
Magnite (2021a, August 27). Data Subject Rights Policy . Magnite. https://www.magnite.com/legal/data-subject-rights-policy/
Magnite (2021b, August 27). Platform Cookies statement . Magnite. https://www.magnite.com/legal/platform-cookie-statement/
Malhotra, N. K., Kim, S. S., & Agarwal, J. (2004). Internet Users’ Information Privacy Concerns (IUIPC): The Construct, the Scale, and a Causal Model. Information Systems Research , 15 (4), 336–355. https://doi.org/10.1287/isre.1040.0032
Marotta, V., Wu, Y., Zhang, K., & Acquisti, A. (2022). The Welfare Impact of Targeted Advertising Technologies. Information Systems Research , 33 (1), 131–151. https://doi.org/10.1287/isre.2021.1024
Martin, K. (2022). Finding Consumers, No Matter Where They Hide: Ad Targeting and Location Data. In Ethics of Data and Analytics: Concepts and Cases (pp. 99–111). Auerbach Publications. https://doi.org/10.1201/9781003278290-16
Meta. (2021, November 9). Removing Certain Ad Targeting Options and Expanding Our Ad Controls . https://www.facebook.com/business/news/removing-certain-ad-targeting-options-and-expanding-our-ad-controls 22
Mitchell, J. (2021, January 21). Reviewing Unified ID 2.0 for Long-Term Industry Value – IAB Tech Lab. IAB Tech Lab . https://iabtechlab.com/blog/reviewing-uid2-for-long-term-industry-value/
OpenX (2022, March 3). OpenX Ad Exchange Privacy Policy . OpenX. https://www.openx.com/privacy-center/ad-exchange-privacy-policy/
O’Reilly, L. (2020, January 14). Google plans to kill off third-party cookies in Chrome “within 2
years.” Digiday .
https://digiday.com/media/google-plans-kill-off-third-party-cookies-chrome-within-2-years/
Parkin, R. (2021, June 27). What the Delay to the End of Third-Party Cookies Means for Advertisers . AdExchanger. https://www.adexchanger.com/data-driven-thinking/what-the-delay-to-the-end-of-third-party-cookies-means-for-advertisers/
Pubmatic (2020, July 1). Platform Cookie & Other Similar Technologies Policy . Pubmatic. https://pubmatic.com/legal/platform-cookie-policy/
Roesner Franziska, Tadayoshi Kohno, and David Wetherall. (2012). “Detecting and defending against third-party tracking on the web.” In Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation, NSDI , pages 155–168.
Schiff, A. (2020a, November 17). Magnite Hops Aboard The Unified ID 2.0 Train . AdExchanger. https://www.adexchanger.com/online-advertising/magnite-hops-aboard-the-unified-id-2-0-train/
Schiff, A. (2020b, November 19). PubMatic Is The Latest Ad Tech Company To Join Unified ID 2.0 . AdExchanger. https://www.adexchanger.com/publishers/pubmatic-is-the-latest-ad-tech-company-to-join-unified-id-2-0/
Schiff, A. (2020c, December 16). OpenX Is Latest SSP To Join Unified ID 2.0 . AdExchanger. https://www.adexchanger.com/online-advertising/openx-is-latest-ssp-to-join-unified-id-2-0/
Schiff, A. (2021a, March 3). Xandr Integrates With Unified ID 2.0 And Outlines Its Identity Roadmap . AdExchanger. https://www.adexchanger.com/online-advertising/xandr-integrates-with-unified-id-2-0-and-outlines-its-identity-roadmap/
Schiff, A. (2021b, April 21). ID Graph Provider Infutor Joins The Club With Support Unified ID 2.0. AdExchanger. https://www.adexchanger.com/online-advertising/id-graph-provider-infutor-joins-the-club-with-support-unified-id-2-0/
Schiff, A. (2021c, May 26). LiveRamp Launches Identity Resolution For First-Party Data. AdExchanger . https://www.adexchanger.com/data-exchanges/liveramp-launches-identity-resolution-for-first-party-data/
Schiff, A. (2021d, May 27). SWAN Vs. SWAN: The Differences Between The Two 3P Cookie Alternative Proposals . AdExchanger. https://www.adexchanger.com/online-advertising/swan-vs-23
swan-the-differences-between-the-two-3p-cookie-alternative-proposals/
Schiff, A. (2021e, November 16). Unified ID 2.0 Faces Roadblocks In Europe As A Result Of GDPR . AdExchanger. https://www.adexchanger.com/privacy/unified-id-2-0-faces-roadblocks-in-europe-as-a-result-of-gdpr/
Schiff, A. (2022a, March 3). IAB Tech Lab Declines To Be The Admin For UID2 (But Hope Springs Eternal) . AdExchanger. https://www.adexchanger.com/online-advertising/iab-tech-lab-declines-to-be-the-admin-for-uid2-but-hope-springs-eternal/
Schiff, A. (2022b, March 24). Epsilon Is Making Its Identity Platform Interoperable With Unified ID 2.0 . AdExchanger. https://www.adexchanger.com/data-exchanges/epsilon-is-making-its-identity-platform-interoperable-with-unified-id-2-0/
Schiff, A. (2022c, March 24). Google’s Encrypted Signals Program Just Entered Open Beta, And Here’s What You Need To Know About It . AdExchanger. https://www.adexchanger.com/ad-exchange-news/googles-encrypted-signals-program-just-entered-open-beta-and-heres-what-you-need-to-know-about-it/
Schuh, J. (2020, January 14). Building a more private web: A path towards making third party
cookies obsolete. Chromium Blog .
https://blog.chromium.org/2020/01/building-more-private-web-path-towards.html
Shein, E. (2021, September 14). Third-party cookies are going away: What advertisers,
marketers and consumers should know . TechRepublic.
https://www.techrepublic.com/article/third-party-cookies-are-going-away-what-advertisers-marke ters-and-consumers-should-know/
Sherman, J. (2021). Data Brokers and Sensitive Data on U.S. Individuals . Duke University
Sanford School of Public Policy. https://sites.sanford.duke.edu/techpolicy/report-data-
brokers-and-sensitive-data-on-u-s-individuals/
Shields, R. (2022, January 7). With Prebid at the helm of UID 2.0, indie ad tech marches to a unified beat but not all voices are in harmony . Digiday. https://digiday.com/media/with-prebid-at-the-helm-of-uid-2-0-indie-ad-tech-marches-to-a-unified-beat-but-not-all-voices-are-in-harmony/
Sluis, S. (2018, November 20). Advertiser Perceptions: How SSPs Can Win Market Share From
Google . AdExchanger.
Solomos, K., Kristoff, J., Kanich, C., & Polakis, J. (2021). Tales of Favicons and Caches:
Persistent Tracking in Modern Browsers. Proceedings 2021 Network and Distributed System
Security Symposium . https://doi.org/10.14722/ndss.2021.24202 24
Southern, L. (2020, January 29). Beyond ad targeting, the demise of the third-party cookie will
hit key digital media functions . Digiday. https://digiday.com/media/cookie-collateral-damage/
Secure Web Addressability Network (SWAN). (2021). Secure Web Addressability Network (SWAN) . https://swan.community/
SWAN-community (2021). Secure Web Addressability Network (SWAN) - Model Terms Explainer . https://github.com/SWAN-community/swan/blob/main/model-terms-explainer.md
Temkin, D. (2021, March 3). Charting a course towards a more privacy-first web. Google .
https://blog.google/products/ads-commerce/a-more-privacy-first-web/
The Trade Desk. (2021a, February 21). What the Tech is Unified ID 2.0? . The Trade Desk. https://www.thetradedesk.com/us/news/what-the-tech-is-unified-id-2-0
The Trade Desk. (2021b, April 8). Publicis Groupe all-in on first-party identity solution . The Trade Desk. https://www.thetradedesk.com/us/news/publicis-groupe-all-in-on-first-party-identity-solution
The Trade Desk. (2021c, July 28). I PG Leans into Unified ID 2.0 as a Closed Operator . The Trade Desk. https://www.thetradedesk.com/cn/news/ipg-leans-into-unified-id-2-0-as-a-closed-operator
The Trade Desk. (2022). Unified ID 2.0 Partners . https://www.thetradedesk.com/us/about-us/industry-initiatives/unified-id-solution-2-0/unified-id-2-partners
Thomson, M., & Rescorla, E. (2021, August). Comments on SWAN and Unified ID 2.0 . Mozilla. https://mozilla.github.io/ppa-docs/swan_uid2_report.pdf
Tranco. (n.d.). A research-oriented top sites ranking hardened against manipulation—Tranco .
Retrieved July 26, 2022, from https://tranco-list.eu/
Twitter. (2022). Targeting of Sensitive Categories . https://business.twitter.com/en/help/ads-policies/campaign-considerations/targeting-of-sensitive-categories.html
UnifiedID2 (2022). uid2docs . https://github.com/UnifiedID2/uid2docs
Vargas, A. (2022a, February 4). Google Reigns Supreme In Latest Advertiser Perceptions SSP
Report, But Competition Is Tight Among Everyone Else . AdExchanger.
Vargas, A. (2022b, April 12). Goodway Group Stitches Together Identity Graph To Complement Brands’ First-Party Data . AdExchanger. https://www.adexchanger.com/agencies/goodway-group-stitches-together-identity-graph-to-complement-brands-first-party-data/ 25
W3C Working Group. (2019, January 22). Tracking Compliance and Scope . Tracking
Compliance and Scope. https://perma.cc/3HXA-F47U
Wei, M., Stamos, M., Veys, S., Retinger, N., Goodman, J., Herman, M., Filipczuk, D., Weinshel, B., Mazurek, M., & Ur, B. (2020). What Twitter Knows: Characterizing Ad Targeting Practices, User Perceptions, and Ad Explanations Through Users’ Own Twitter Data. In 29th USENIX Security Symposium (USENIX Security 20) (pp. 145-162).
Xandr (2022, April 29). Digital Platform Cookie Policy . Xandr. https://www.xandr.com/privacy/cookie-policy/
Xu, H., Dinev, T., Smith, J., & Hart, P. (2011). Information Privacy Concerns: Linking Individual Perceptions with Institutional Privacy Assurances. Journal of the Association for Information Systems , 12 (12). https://doi.org/10.17705/1jais.00281 26
Appendix #1 – 50 Sites Crawled to Evaluate Cookie-Based Tracking on SSP Networks
https://www.washingtonpost.com,
https://www.independent.co.uk,
https://www.huffingtonpost.com,
https://www.merriam-webster.com,
Appendix #2 - Persistent Identification Within the Four SSP Networks 28
Appendix #3 - observed sharing of information between Pubmatic and Rubicon
In our data collection on cookie-based tracking in SSP networks we did see that in two unique cookies’ hostnames associated with Pubmatic, Rubicon was named in stateless crawl 1 of 10 (See Table 4 below). We did see other occurrences with different persistent identifiers and similar host name syntax in each of the 10 stateless crawls.
Looking further into Rubicon’s (now known as ‘Magnite’) privacy policy, the company lists Pubmatic as an ‘Advertising Cookies – Third Party Technology Providers’ partner with ‘DMP, Onboarders, Data Providers’ activities associated with this classification.
We are still trying to figure out what this potential cookie syncing between those actors means for consumer tracking.
Persistent Identifier Host Name
KADUSERCOOKIE26F46775-238D-433A-B321-8B4D7D2349A9
https://ads.pubmatic.com/AdServer/js/user_sy nc.html?gdpr=&gdpr_consent=&us_privacy=1 ---&predirect=https%3A%2F%2Fprebid-server. rubiconproject.com %2Fsetuid%3Fbi dder%3Dpubmatic%26gdpr%3D%26gdpr_co nsent%3D%26us_privacy%3D1---%26account%3D%7B%7Baccount%7D%7D %26f%3Db%26uid%3Dnull
https://ads.pubmatic.com/AdServer/js/user_sy nc.html?gdpr=0&gdpr_consent=&us_privacy= 1NNN&predirect=https%3A%2F%2Fprebid-server. rubiconproject.com %2Fsetuid%3Fbi dder%3Dpubmatic%26gdpr%3D0%26gdpr_c onsent%3D%26us_privacy%3D1NNN%26ac count%3D%7B%7Baccount%7D%7D%26f% 3Db%26uid%3Dnull
Table 4: Possible ID cookie sharing between Pubmatic and Rubicon in stateless crawl 1 of 10