面向移动应用的指纹识别 SDK 及其发现途径:理解设备指纹市场 原文标题:Fingerprinting SDKs for Mobile Apps and Where to Find Them: Understanding the Market for Device Fingerprinting
内容概要总结
本文(arXiv:2506.22639v1,作者 Michael A. Specter 等,Georgia Tech 与 Google)是迄今最大规模的 SDK 指纹识别行为分析。数据集含来自 9 个 Maven 仓库的 228,598 个 SDK 与来自 Google Play 的 3,025,417 个 APK(最终分析 178,054 个应用),静态分析流水线检测超过 500 个信号的外传。核心结论:Ads SDK 只是 30.56% 指纹识别行为的来源,23.92% 来自目的未知/不明的 SDK,安全与认证 SDK 仅关联 11.7%;信号/API 使用分布稀疏(仅 2% 的外传 API 被超过 75% 的 SDK 使用),故难以靠权限控制。方法上以自我宣称指纹识别的 14 个 SDK 为种子集(唯一信号 504 个,平均每 SDK 外传 75.5 个),用静态污点分析(CoFlow)选出 723 个可能指纹识别的扩展集,并按五类目的标签人工标注(Krippendorff's alpha 0.804)。结果:39.4% 热门应用含至少一个指纹识别 SDK,Game/Comics/Dating 类别占比最高,23 个类别中此类应用比其他应用受欢迎 10 倍,存在跨 App 跟踪风险。论文还提出用细粒度代码相似度做 SDK 识别(平均精度 65.07%)。
翻译内容
原文内容(English)
Michael A. Specter
Georgia Tech & Google LLC, Atlanta, GA, USA Email address: mobile-fingerprinting-sdks-paper@google.com // specter@gatech.edu
Mihai Christodorescu — Google LLC, Mountain View, CA, USA
Abbie Farr — Google LLC, Mountain View, CA, USA
Bo Ma — Google LLC, Mountain View, CA, USA
Robin Lassonde — Google LLC, Mountain View, CA, USA
Xiaoyang Xu — Google LLC, Mountain View, CA, USA
Xiang Pan — Google LLC, Mountain View, CA, USA
Fengguo Wei — Google LLC, Mountain View, CA, USA
Saswat Anand — Google LLC, Mountain View, CA, USA
Dave Kleidermacher — Google LLC, Mountain View, CA, USA
摘要
本文对移动应用生态中类似指纹识别(fingerprinting-like)的行为进行了大规模分析。我们采取基于市场的方法,聚焦于由应用普遍使用第三方 SDK 所促成的第三方跟踪。我们的数据集包含来自流行 Maven 仓库的逾 228,000 个 SDK、从 Google Play 商店采集的 178,000 个 Android 应用,我们的静态分析流水线检测到超过 500 个独立信号的外传。据我们所知,这是迄今对 SDK 行为进行的最大规模分析。
我们发现,Ads SDK(广告 SDK,Apple 的 App Tracking Transparency 和 Google 的 Privacy Sandbox 等业界努力表面上关注的焦点)似乎只是 30.56% 的指纹识别行为的来源。令人意外的 23.92% 源自目的未知或不明晰的 SDK。此外,Security and Authentication SDK(安全与认证 SDK)只与 11.7% 的可能指纹识别实例相关联。这些结果表明,仅在广告这类特定细分市场情境中处理指纹识别,可能收益不完整。执行反指纹识别政策也很复杂,因为我们观察到可能进行指纹识别的 SDK 所使用的信号和 API 分布稀疏。例如,只有 2% 的外传 API 被超过 75% 的 SDK 使用,这使得难以依靠用户权限来控制指纹识别行为。
1. 引言
设备指纹识别是一种通过收集关于设备特定硬件、软件和配置设置的大量信息来识别和跟踪用户设备的技术。这些属性的组合为该设备创建一个唯一的、或近乎唯一的数字「指纹」。这一过程有明确的隐私隐忧——指纹识别标识符可以在用户无法控制或不知情的情况下被收集,并在设备的整个生命周期内持续存在,无论用户采取多少寻求隐私的操作(例如清除 cookie、轮换广告 ID 或启用隐私浏览)。
两大移动操作系统厂商都已做出重大努力来限制设备指纹识别的隐私影响。Apple 引入了要求应用在收集与跟踪相关的设备数据前须征得用户同意的政策(5),并要求为使用特定的高熵「required reason」(必需理由)API 提供人类可读的解释(4)。Google 和 Apple 的移动平台现在都要求开发者向用户提供营养标签式的隐私信息(3, 25),或作为提交给各自应用商店的元数据,或作为应用本身的一部分附加。Google 正在为 Web 开发隐私沙盒(1),并在 Android 上开发一个新的沙盒,限制第三方广告库访问应用其余部分可访问的敏感信息(26)。这些干预措施很有前景——提供了急需的透明度和问责制。
此类反指纹识别努力的成功取决于应用开发者的技术实现、市场契合度和意图。例如,Apple 的反跟踪和应用透明度政策明确允许为反欺诈目的收集指纹识别数据(5),而 Android 的 Privacy Sandbox 仅专注于将代码与第三方广告商隔离。Apple 的「required-reason APIs」方法也有局限;它目前仅适用于 30 个 API,其有效性取决于在更广泛的指纹识别生态中所收集的其他数据点。因此,刻画现实世界中指纹识别的技术实现并理解其中的利益相关方,将为了解这些干预措施的有效性提供宝贵洞见。
本文对 Android 应用生态内的设备指纹识别实践进行了全面的、大规模的分析。我们采用以经验为基础的方法,核心是识别集成在移动应用中的第三方软件开发工具包(SDK),衡量其市场覆盖、跟踪方法和隐私影响。据我们所知,本研究是关于设备指纹识别等侵犯隐私实践的 SDK 行为的最全面分析,拥有超过 228,000 个唯一 SDK 和 178,000 个 Android 应用的庞大数据集。我们的方法论力求对移动生态中设备指纹识别的规模与范围形成更细致的理解;我们避免套用自己可能有偏或狭隘的指纹识别行为定义,并采用多种技术,在主观分析不可避免时提供可靠性与一致性。
尽管已有许多研究衡量了指纹识别的影响(相关工作总结见 §2),但围绕指纹识别行为的目的仍知识匮乏。例如,此前没有研究尝试理解哪些类型的第三方收集了足以对设备进行指纹识别的信息,以及这如何与第一方开发者的需求相契合。从开发者视角刻画整体问题,有助于判断这些技术为何被使用,为更好地保护用户隐私所需的条件提供宝贵洞见,并解读当前及拟议执法方法背后假设的价值。
有若干挑战显著地使我们的研究复杂化。任何指纹识别检测机制都可能不完整,因为从设备收集熵的方法很多(且可能是隐蔽的),包括计时信息、指令执行特性以及其他硬件特定来源。SDK 的分类与分析也是一项困难任务——尽管应用会自我标注其用途和市场契合度,但当前的 SDK 分发方式并不要求 SDK 作者对其代码提供有意义的描述。我们在 §3 中深入描述这些挑战及其他挑战的解决方案。
一个重要的挑战是定义上的:声称某项服务在进行指纹识别,暗示了代码作者的意图,而这在实践中最往往无法确知。应用可能出于诸多原因收集足以唯一识别设备的信息,包括分析、崩溃报告、反欺诈,或应用自身正常运行所需。我们强调,本研究纯粹是观察性的——我们度量指纹识别行为,不对代码作者赋予任何动机。我们还要强调,我们选择考察 Android 生态完全是因为便利,我们的结果很可能也适用于 iOS。正如此前工作(35)所指出的,Android 的开放生态允许可扩展的分析,而 iOS 的数字版权管理方案则主动阻碍同样的分析。
我们回答以下研究问题:
- RQ1:自我标识为指纹识别的 SDK 表现出哪些类型的行为?
- RQ2:具有可能指纹识别行为的 SDK 所声明的用途是什么?
- RQ3:哪些类型的应用使用具有可能指纹识别行为的 SDK,这些 SDK 在真实世界应用中的普遍程度如何?
我们发现,许多类型的 SDK 都收集足以跟踪用户的信息(每个 SDK 至少外传 20 个信号),且所收集的信号高度多样(在 SDK 数据集中观察到的 504 个唯一信号里,SDK 平均外传 75.5 个信号)。尽管广告确实占表现出指纹识别行为的 SDK 的显著部分(约 30%),但常见 Android 应用中使用的、数量惊人的类似指纹识别 SDK 功能不明、缺乏可供分类的充分描述(约 24%)。反欺诈和分析服务在我们的数据集中也很普遍,表明需要更多研究来为这类功能中所用的指纹识别创建保护隐私的替代方案。最后,表现出可能指纹识别行为的 SDK 受欢迎程度不成比例地高——安装量约为非指纹识别替代品的 10 倍——且单个 SDK 很可能跨越多个应用细分市场(例如健康与约会)存在。
图 1. 我们分析流水线的概览。我们首先 ① 从一系列 Maven 仓库和 Google Play 应用商店获取应用、SDK 及相关元数据,仅选择安装在超过 10k 活跃设备上的应用(§3.1)。接着 ② 提取一个种子集(Seed Set),其文案表明它们在进行指纹识别(§3.2)。然后 ③ 使用静态污点分析确定哪些 SDK 外传这些信号(§3.3),并 ④ 人工标注所得 SDK 以确定其市场契合度(§3.4)。最后 ⑤ 进行另一轮静态分析以确定哪些应用包含哪些 SDK(§3.5)。
2. 背景与相关工作
据我们所知,我们的工作是首个对真实世界中基于原生应用的设备指纹识别的大规模研究。此前没有研究尝试理解这一现象为何普遍,或围绕这些工具使用的市场。现有文献大多源于将指纹识别作为一种攻击来考察,聚焦于新颖的指纹识别方法。
Android 应用与 SDK
Android 应用可以用任何语言编写,并可从任意来源安装,包括 Play 商店、二级应用商店、侧载,或由制造商预装在设备上。因此,Android 的许多安全模型都围绕对应用进行沙盒化,结合使用 SELinux 的 SEPolicy 和标准 Linux UID 式访问控制机制。除沙盒外,对某些敏感数据的访问由应用提供的元数据声明,并通过安装时和运行时的权限检查强制执行(42)。除非被手动沙盒化,Android 第三方库(称为 SDK)在第一方应用的上下文中执行,因此享有与第一方应用相同的权限。
SDK 可以作为原始代码分发,也可以在构建时从任意数量的仓库或构建系统自动下载。一个常用的构建系统是 Maven 格式,这是一种用于 Java 依赖解析的开放标准。在实践中,SDK 的分发通常使用名为 Gradle 的构建工具完成,它从任意数量的公共 Maven 仓库加载 SDK。
指纹识别信号
在本文中,我们将信号(signal)定义为从设备收集的单个数据点。不同论文对指纹中的唯一数据点有不同的术语,Eckersley(18)称之为变量(variable)。信号可以从许多来源获得,包括 API 调用、平台上的常见文件、系统属性、硬件特性或运行时环境值。
信号的多样性
尽管指纹识别通常在浏览器的语境下被讨论(19, 18),但围绕通过原生代码进行指纹识别的众多方法有丰富的文献。可指纹识别的组件包括麦克风和扬声器(14, 70, 9)、加速度计、陀螺仪和磁力计(69, 68, 54, 16, 9, 40, 57)、硬件时钟(34, 51)、摄像头(8, 49)、GPU 时钟(44)以及电池(13, 47)。对系统或应用的指纹识别范围从特定 API(48)到系统配置(例如通过 procfs(52, 55))、用户设置(37)和浏览器配置(18, 53)。指纹识别还可以通过与邻近设备通过蓝牙等短距离无线电协议通信来完成(36)。不同的信号可以组合以提高准确度(2, 12)。
检测与预防
检测和预防方法包括对 API 调用进行机器学习分类(7, 21, 31)、重新校准传感器(15)、更改系统设置(32)、为收集的数据添加随机噪声(45, 15),或使用污点跟踪来识别可指纹识别数据的外传(39)。权限系统对指纹识别不能提供充分的保护(60, 17)。
此前的测量研究
已有若干研究测量了指纹识别的使用,不过大多数聚焦于 Web(19, 46)。一个显著的挑战似乎是不断增长的 API 表面,它们已被指纹识别者迅速采用(7)。
对 Android 原生系统的现有分析相对稀少。纵向分析不仅凸显了这种独特的、移动特有的基于 SDK 的跟踪形态,还表明隐私风险随时间和应用版本变化很大,与现有执法方法几乎无关联(50)。Han 等人(29)发现,隐私高风险行为(包括指纹识别)的存在似乎并不随应用的成本而变化,免费应用和付费应用共享相似的第三方 SDK 或危险权限集合。
与我们的研究最接近的是 Torres 等人 2018 年关于识别应用中指纹识别的工作(22)。他们发现移动设备上的指纹识别者更多地依赖分类信号而较少依赖侧信道,并主张移动端指纹识别的检测与预防不同于 Web 浏览器。我们的工作在信号、SDK 和应用规模上显著更大(30k 对比我们的 178k),呈现了对生态更完整的理解,此外还有 SDK 标签和更多统计。
3. 方法论
在本节中,我们深入讨论我们的分析流水线。我们在图 1 中描绘了整体过程,并在下面概述:
- 数据集收集(§3.1):我们通过爬取流行 Maven 仓库获取 SDK 数据集,并从 Google Play 商店获取应用数据集。
- 种子集与人工信号提取(§3.2):我们利用每个 SDK 发布的广告文案生成一个种子集(Seed Set),其中的 SDK 在公开声明中自我宣称是指纹识别者。然后我们人工逆向工程这些 SDK,提取每个 SDK 外传的信号。
- 自动化信号外传检测(§3.3):将静态污点分析与由我们的种子集生成的数据配合使用,我们确定一个 SDK 是否外传相关的指纹识别信号。我们把所有外传足够多、从而表现出指纹识别行为的 SDK 称为扩展集(Extended Set)。
- SDK 标注与分析(§3.4):与应用程序不同,SDK 是未标注的,不携带与市场、用例或目标受众相关的元数据。为提供充分的统计和市场信息,我们人工标注扩展集中的所有 SDK。为避免偏见和无意义的标签,我们借鉴 HCI 社区的编码技术,迭代地制定编码手册(codebook)并就 SDK 标签定义和分配达成共识。
- 应用-SDK 匹配(§3.5):为提供每个 SDK 使用情况的统计,我们必须首先确定哪个 SDK 存在于哪个应用中。我们通过一系列静态分析技术来实现这一点。
我们聚焦 SDK 而非整体应用有若干原因。人们可能期望一个 SDK 是自包含的,并且有明确公开的功能。与任何现代开发环境一样,SDK 在应用中的使用极为普遍,大多数应用都使用第三方库来支持各种核心功能。最后,强调 SDK 也使第三方的潜在危害更加清晰:用户更可能理解并信任自己主动安装的应用,但可能并未意识到他们对应用所使用的第三方 SDK 和服务所施加的传递性信任。
3.1 数据集收集
应用数据集
我们在近 18 个月内(2023 年 1 月至 2024 年 5 月)收集了 Google Play 商店上发布的 3,025,417 个 APK2。我们用每个应用的总受众规模——单个 APK 已安装到的活跃设备数量——来补充这组 APK。活跃设备是指在前 30 天内至少开机过一次的设备(24)。
为避免样本集偏向缺乏显著用户基础的应用,我们把分析限制在 2024 年 4 月 13 日至 2024 年 5 月 13 日期间在 Google Play 商店上活跃、总受众规模超过 10,000 的应用。总计覆盖 178,054 个应用。虽然我们无法估计这些应用被启动的频率,但受众规模这一指标确保安装这些应用的设备处于活跃使用中。我们通过对这些应用所有 30 天活跃安装量求和来近似一个应用的市场覆盖——我们注意到这可能重复计算那些安装、移除后又重新安装同一应用的用户,以及同一用户在一台设备上以多个配置文件(例如个人和工作)或在多台设备(例如手机和平板)上的安装。
SDK 数据集
SDK 没有单一的信息来源,开发者改为向其构建系统提供特定 Maven 仓库的 URL 和库 ID。我们使用自定义爬虫整合数据集,从 9 个独立的大型 Maven 仓库中提取所有 SDK:JCenter、Maven Central、Google、Sonatype、Spring.io、Jitpack、Bintray 和 Artifactory。从这些仓库中,我们获取了 228,598 个 SDK 的数据集以及每个 SDK 的相关元数据。我们排除了版本标签包含 {"alpha"、"beta"、"test"、"dev"、"debug"、"qa"} 之一的所有 SDK。
3.2 种子集与人工信号提取
从我们的 SDK 数据集中,我们选择了一个种子集(Seed Set),其中的 SDK 在其广告文案或其他元数据中公开承认出于指纹识别目的收集信息。然后我们人工逆向工程每个 SDK,确认该 SDK 在收集一组非平凡的设备数据,并提取该 SDK 上传到服务器的信号列表。为避免错误标注,每个候选 SDK 随后由第二位分析师再次逆向工程,独立确认信号列表。总计,我们的种子集包含 14 个 SDK,报告了超过 500 个不同的信号。这一工作的结果在 §4.1 中更详细地报告。
3.3 自动化信号外传检测
我们收集了一个类似指纹识别 SDK 的扩展集(Extended Set),由在收集数据方面与种子集相似的 SDK 组成。为获得该集合,我们开发了一套静态分析套件,对 SDK 及其依赖进行信息流分析,并选出与我们的种子集有足够信号重叠的 SDK。注意,扩展集中的 SDK 可能收集超出种子集中所发现的信号。
为避免过度声称我们数据集中存在指纹识别行为,我们只纳入外传信号数超过我们种子集中任一 SDK 所收集最低信号数的 SDK。换句话说,只有当某个 SDK 外传的量超过一个公开承认进行指纹识别的 SDK 所上传的量时,我们才认为它表现出指纹识别行为。注意这是一个保守估计——更精细的熵估计很可能表明,一个 SDK 能用比我们在此考察更少的信息唯一识别用户。这里的权衡是有意为之的;我们的目标是在不陷入更复杂分析的情况下给出上限估计。此步骤得到 723 个不同的 SDK 家族,每个有多个版本,共 14,178 个 SDK 版本。
我们构建了一套用于 Android APK 和 SDK 分析的静态分析套件,并部署了一种过程间、上下文敏感、字段敏感、对象敏感的污点流跟踪算法用于指纹识别检测。它通过用元信息污点标记所有与指纹识别相关的数据,然后在指令级传播污点来工作,从而可以实现流信息的透明性,允许我们重建污点流路径以独立验证外传。我们对这一分析不声称任何新颖性,不过一些实现细节可能具有独立价值,因此我们将其列入附录 B。
表 1. SDK 标签定义。每个 SDK 基于其 Maven 元数据和网站描述获得一个标签,而非基于其代码。更正式的描述见附录 D 的表 6。
| SDK 标签 | 描述 |
|---|---|
| Advertising(广告) | 支持展示广告、广告竞价、广告定向、广告中介,或出于变现或转化目的的分析(例如:AppLovin、Teads) |
| Analytics(分析) | 监测并报告应用健康状况(例如:TOAST Logger、RichAPM Agent),或收集用户在应用中的行为(例如:Pushwoosh、Acoustic Tealeaf) |
| Security & Authentication(安全与认证) | 实现用户认证(例如:Passbase、Ondato),检测欺诈及相关安全异常(例如:Incognia、SEON),以及支付功能(例如:Alipay、PayPal)。 |
| Tools / Other(工具/其他) | 提供导航(例如:Tencent Map Nav、Radar)、物体或人员跟踪(例如:BeaconsInSpace、Foursquare Movement)、与社交网络的通信(例如:Facebook、Chat SDK),或其他定义明确的功能(例如:iZooto App Push、GameUp) |
| Unclear / Not Found(不明/未找到) | 无法从在线元数据确定其目的或功能 |
3.4 SDK 标注
尽管开发者在向 Google Play 商店提交时会对其应用的分类用例作出证明,但 SDK 没有等同的流程。事实上,Maven 仓库通常只提供 SDK 的名称、简短说明和一个指向代码来源的链接。这些元数据往往不完整,进一步掩盖了 SDK 的用途。其他数据集更完整,包括 Google Play SDK Index(27),但只提供一个小型的 SDK 数据库。
这里我们借鉴 HCI 社区的技术,把问题当作人工标注任务处理。标签定义由一个五名专家编码员组成的团队,基于 SDK 的元数据(包括 SDK 在 Maven 中的描述及其开发者网站的内容),从我们数据集中 100 个随机 SDK 的样本协作制定。为简洁起见,这些类别的非正式描述见表 1,完整解释(包括子类别)见表 6。在达到饱和后,我们把剩余的 SDK 分配给评审员,使每个 SDK 都被独立检查两次。
为提高效率,我们把标注工作限制在我们的应用数据集中检测到的 723 个 SDK 家族,假设同一 SDK 的所有版本具有等同的用例并应共享同一标签。所得定义是稳健的,评审员通常在 SDK 标签上达成一致;独立标注步骤的 Krippendorff's alpha 评分者间信度分数为 0.804。最后,所有标签分歧都在全体小组会议中解决并重新标注,这意味着任何标注分歧都通过比较全部五位编码员的标签来处理。
3.5 应用-SDK 匹配
受 Android 应用 SDK 识别的大量工作(65, 67, 66, 6, 41, 38, 58, 38, 30, 64, 61)启发,我们创建了一个 SDK 识别流水线,使用一种细粒度的代码相似度度量,它可以跨代码单元(例如类、模块)和打包单元(例如 SDK、SDK 版本)聚合。该相似度度量依赖系统 API 的标识符(例如操作系统调用、标准库调用)、操作码频率、框架 API 和字符串常量。
在我们的设计中,我们为该相似度度量选择参数,包括判定 SDK 存在于某 APK 中所需的与该已知 SDK 相似代码的百分比,以确保我们的结果限制误报,代价是一些漏报。换句话说,我们可能漏掉某个 SDK 在某个 APK 中的存在,因此论文其余部分的统计分析为指纹识别 SDK 的普遍程度提供了下界。我们在附录 C 中包含了我们方法的详细描述。
4. 结果
我们分析了种子集、扩展集及其在我们的应用数据集中的普遍程度,以回答我们的三个研究问题(§1)。
4.1 RQ1:自我标识的指纹识别者
图 2. 种子集中已知指纹识别 SDK 所收集信号的分布图,每个外传的 API 用一个点表示。上方的图显示外传该 API 的种子集 SDK 百分比;80% 的 API 被不到 50% 的 SDK 外传,只有 2% 的 API 被 75% 的 SDK 外传。
表 2. 种子集 SDK。通过广告文案、开发者文档或其他文案公开承认进行指纹识别的 SDK 列表。它们收集的原始信号数在右侧。「唯一信号」不是总数,而是所有信号去重后的集合并集。
| 名称 | 信号数 |
|---|---|
| Seon | 43 |
| Forter | 69 |
| Kaspersky AntiVirus SDK | 20 |
| Accertify (InAuth) | 213 |
| Castle | 31 |
| Microsoft Dynamics 365 | 128 |
| IP Quality Score | 58 |
| Fingerprint.js | 30 |
| Shield | 148 |
| ThreatMetrix (Lexus Nexus) | 94 |
| Ravelin | 30 |
| TransUnion TruValidate | 55 |
| Socure | 43 |
| Incognia | 81 |
| 唯一信号 | 504 |
自我标识的指纹识别 SDK 表现出哪些类型的行为?
我们在真实世界中找到 14 个承认进行指纹识别的 SDK,列于表 2。我们的人工分析发现,指纹识别库至少外传 20 个唯一信号,平均 75.5 个。值得注意的是,这些 SDK 使用的技术很直接,没有任何 SDK 试图收集超出框架级 API 调用所能获得的信息。这与 Web 语境中此前所测量的(例如音频上下文指纹识别(19))以及学术界讨论的更先进的、以硬件为重点的技术有所不同。
图 3. 种子集 SDK 之间的余弦相似度,每个 SDK 用其使用特定 API 的独热编码向量表示。
我们考察种子集中 SDK 所收集信号的分布,并在图 2 中给出概览。尽管收集的信号总数为 1043,但只有 504 个唯一信号,不到一半的信号被至少两个 SDK 收集。所收集的指纹识别信号集相对稀疏,因为 SDK 选择收集的各个信号有些不同——图 3 显示不同指纹识别 SDK 之间的余弦相似度,只有两个 SDK(Transunion 和 Ravelin)得分高于 0.5。在收集的 504 个唯一信号中,我们发现只有 21 个独立 API 被我们种子集中超过一半的 SDK 收集,最多被 11 个 SDK 收集。我们得出结论:单个信号的余弦相似度(如此前工作(22)所用)不太可能是一种有效的检测机制。
一些 SDK 以独特的 API 使用模式脱颖而出。例如,Forter、Accertify、Microsoft Dynamics 和 Shield 使用的 API 范围似乎比其他的宽得多。这可能表明这些 SDK 更复杂,或服务于更广泛的功能。相反,Kaspersky AntiVirus SDK 使用的 API 集似乎非常有限,可能反映了它更聚焦的安全用途。在假设所有这些 SDK 进行指纹识别的效果相当的前提下,这种 API 使用模式的宽度差异意味着某些 API 提供的信号比其他更有价值(即更可指纹识别)。
4.2 RQ2:可能指纹识别者的目的
具有可能指纹识别行为的 SDK 所声明的用途是什么?
自动外传检测(§3.3)得到 723 个表现出与种子集中已知指纹识别行为相似行为的 SDK。每个 SDK 可能有多个版本,我们的 SDK 数据集为这 723 个 SDK 识别出 14,178 个版本,平均每个 SDK 19.60 个版本。在按第 3.3 节所述标注扩展集后,我们考察表现出类似指纹识别行为的 SDK 中各种目的的普遍程度,并在图 4 中绘出所得分布。
图 4. 723 个可能指纹识别 SDK 的扩展集中各目的的普遍程度(数值:Analytics 77、Security and Authentication 85、Tools / Other 167、Unclear / Not found 173、Ads 221)。
从图中可以立即得出几点观察。首先,「Ads」类别、「Tools / Other」类别与其余类别之间有明显的区分。「Tools / Other」类别较大是意料之中的,因为它涵盖从云存储到图像视频处理、再到同意管理等各种功能。以「Ads」为目的的 SDK 数量较多,可能反映了该特定行业的复杂结构:广告网络提供自己的 SDK,广告平台作为应用与广告网络之间的聚合者和中介,并配有相应的移动 SDK 来反映这些关系。广告中介系统往往由 5–15 个独立 SDK 组成,每个对应中介可连接的一个广告网络,很容易抬高该垂直领域 SDK 的总数。我们选择把这些连接器 SDK 分开计数——从安全和隐私的角度看,它们是不同的产物。
第二点观察是,有大量 SDK 的目的无法从在线公开信息中看出。「Unclear / Unfound」类别是可能指纹识别行为的第二大来源。我们知道这些 SDK 被各种应用使用(正如我们稍后在 §4.3 所见),所以问题不仅是这些 SDK 服务于什么目的,还在于如何让这一目的信息可供安全和隐私执法机制使用。一种选择是逆向工程每个 SDK 的代码并从中推断其目的,但这不太可能是一种可扩展的长期方案(也超出本文范围)。
关于采用保护隐私的替代方案,另一个考虑因素是「Tools / Other」类别中存在的功能长尾,它占可能指纹识别 SDK 总数的 23%。尽管「Ads」、「Security and Authentication」和「Analytics」在隐私方面得到了相当充分的理解和研究,但「Tools / Other」SDK 涵盖广泛的算法和数据类型,可能没有现成的保护隐私的替代方案。
SDK 目的与行为。
如果能足够有区分度,识别某一类可能指纹识别 SDK 的目的这一难题,可以通过聚焦其对指纹识别信号的使用来缓解。例如,如果 Ads SDK 的指纹识别行为(在其外传的信号和 API 数据方面)与 Security and Authentication SDK 的不同,人们就可以适当地识别和控制指纹识别,而不必依赖 SDK 的目的声明或其非指纹识别功能——这两者都可能被对抗性地操纵。我们通过把每个 SDK 的指纹识别行为表示为高维空间中的点来评估这一假设,该空间由外传 API 的独热编码定义。每个相关 API 是该空间中的一个独立维度,如果 SDK 不外传相应 API 则置于该维度位置 0,如果外传则置于 1。这得到一个 504 维空间(对应扩展集 SDK 中观察到的每个 API),我们把 723 个 SDK 定位其中。
在图 5 中,我们使用 t 分布随机邻域嵌入(t-SNE)(56)图,提供了扩展集中不同 SDK 类型之间相似度的可视化表示。t-SNE 允许我们基于 SDK 在指纹识别行为上的「自然」相似度对它们聚类,方法是通过非线性变换把高维空间(我们的 504 个信号)映射到低维空间中的忠实表示,同时保留数据点之间的局部和全局关系。根据 Wattenberg、Viégas 和 Johnson(59)的建议,我们将 t-SNE 的困惑度设为 25、学习率为 10、迭代次数为 5,000,得到最终 KL 散度为 0.614789。所得的 t-SNE 输出如图 5 所示,每个 SDK 一个点,位置由 t-SNE 决定,颜色根据我们的五个 SDK 目的标签编码。
图 5. 扩展集中 SDK 的可能指纹识别行为分布图,使用 t-SNE 在我们通过对外传 API 独热编码构造的嵌入上计算得出。SDK 的邻近表明它们从相似的 API 集外传数据。
t-SNE 图展示了指纹识别行为中信号/API 使用的多样性,形成了大量小簇。至少这让我们相信,相应数量更多的小型、聚焦的基于权限的策略或许能解决指纹识别问题,但伴随执法性能(由于维护和评估这么多策略的成本)和低可用性(由于让用户基于看似相似但在隐私上有区别的权限做决定)的风险。
把业界提出的反指纹识别/反跟踪策略(广告:不允许跟踪,反欺诈:允许跟踪)自动化的可行性,可以归结为「Ads」SDK(图 5 中标记为相应符号)是否容易与「Security and Authentication」SDK(图 5 中标记为相应符号)分开。t-SNE 图的右侧三分之一包含大多数「Security and Authentication」SDK,而「Advertising」SDK 在左侧。然而图左侧也有许多「Security and Authentication」SDK,更不用说「Analytics」和「Tools / Other」SDK,它们看起来与「Ads」SDK 有相似的类似指纹识别行为。因此,任何需要区分「Ads」SDK 和「Security and Authentication」SDK 的自动执法,都将需要依赖比权限更具表达力的、非平凡的分类器。
最后我们观察到,「Unclear / Unfound」SDK(在 APK 中声明、存在于 Maven 仓库中,但缺乏任何描述性信息)的类似指纹识别行为与所有其他 SDK 类别都相似。这支持了对 SDK 建立稳健分类和标注机制的必要性,也支持了这样一个结论:如果没有额外的带外(非代码)信息,行为分析可能不够。
敏感信号的使用。
我们人工识别了 24 个可用于精确或近似获取位置数据的 API,然后检查扩展集中有多少可能指纹识别行为依赖这些 API。我们对应用使用信号(基于我们识别出的三个提供关于用户已安装应用、正在使用的应用或已安装应用使用统计信息的 API)和账户列表信号(基于两个获取设备上注册的个人账户列表的 API)做了类似分析。我们发现,在可能指纹识别的 SDK 中,72% 收集粗粒度位置信号,71.6% 收集细粒度位置,86.29% 收集至少其中一种。只有 6.15% 记录账户列表信号,38.46% 收集应用使用信息。
4.3 RQ3:市场覆盖
哪些类型的应用使用具有指纹识别行为的 SDK,这些 SDK 在真实世界应用中的普遍程度如何?
为回答这个问题,我们考虑 RQ2 结果中描述的 SDK 类别,以及 Google Play 商店分配的应用类别(23)。
我们感兴趣的是理解指纹识别 SDK 在移动应用市场中的存在。为此,我们测量了每个应用类别中有多少应用包含指纹识别 SDK、哪些类别的指纹识别 SDK 使用最多,以及哪些指纹识别 SDK 最常在同一应用中共同出现。第一项测量旨在确定是否存在指纹识别普遍程度特别高、因而应优先进行任何减少指纹识别干预的应用类别。第二和第三项测量为任何用保护隐私的替代方案取代基于指纹识别的方案的技术努力提供信息。
图 6 显示每个应用类别中包含指纹识别 SDK 的应用的普遍程度。6(a) 表明,使用指纹识别 SDK 的应用的原始数量在各应用类别中平均为 3.2%,介于 0.8%(「Events」应用)和 10%(「Video Players」应用)之间。6(b) 考虑了每个此类应用在 30 天内的安装量,表明指纹识别功能的存在严重偏向热门应用。
平均而言,一个类别中 39.4% 的应用至少包含一个指纹识别 SDK,虽然最低普遍程度为 5.5% 的应用(「Libraries and Demo」类别),但有若干应用类别普遍程度 >50%:「Maps and Navigation」、「Beauty」、「Shopping」、「Sports」、「House and Home」、「Social」、「Food and Drink」、「Dating」、「Game」和「Comics」。从用户的角度看,这表明随机安装一个热门应用有 39.4% 的概率被指纹识别,而如果他们选择约会、游戏或漫画类应用,被指纹识别的概率将达 80% 以上。
「Comics」、「Game」和「Dating」类别以包含可能指纹识别 SDK 的应用数量最多而脱颖而出,分别按此顺序。这可以归因于若干因素,例如依赖定向广告或应用内购买的免费使用功能(例如免费游戏)的普遍,或在线环境中对(无摩擦的)用户识别的需求。
图 6. 带有指纹识别 SDK 的应用数量其实相当少,平均不到一个类别中应用总数的 5%(6(a)),然而这些应用是安装量最高的一批(6(b)),使得可能指纹识别 SDK 在市场中的存在感异常之大。在 23 个应用类别中(6(c) 中高亮),带指纹识别 SDK 的应用比其他应用受欢迎 10 倍。
(图 6 各子图:(a) 按应用类别统计的应用数(蓝色)和带类似指纹识别 SDK 的应用数(红色);(b) 按应用类别统计的安装量(蓝色)和带指纹识别 SDK 的应用的安装量(红色);(c) 按总数(左)和安装量(右)统计包含类似指纹识别 SDK 的应用百分比。)
指纹识别在各应用类别中的普遍程度
为进一步理解可能指纹识别 SDK 在移动应用生态中的普遍程度,我们使用在 §4.2 中制定的目的标签把应用类别映射到 SDK 类别。这得到图 7(a) 所示的热力图,其中越深的红色表示该行的 SDK 类别在该列的应用类别中越占主导。例如,「Art and Design」应用所使用的任何可能指纹识别 SDK 主要来自「Ads」SDK 类别,而「Finance」应用中的可能指纹识别 SDK 最主要来自「Analytics」SDK 类别。
图 7. 可能指纹识别 SDK 在应用类别内部及之间的普遍程度。在 7(a) 中,颜色梯度按应用类别计算,允许在应用类别之间比较 SDK 普遍程度。在 7(b) 中,颜色梯度按 X 轴和 Y 轴坐标给出的每一对应用类别计算。原始数据见附录 A。
「Ads」SDK 在几乎所有应用类别中都作为可能指纹识别行为的来源占主导,「Unclear / Unfound」SDK 是第二常见的。我们注意到,在许多情况下「Unclear / Unfound」SDK 的绝对数量接近「Ads」SDK(例如在「Business」应用类别中,16,097 个 SDK 的标签不明,17,780 个 SDK 带有「Ads」标签),因此任何从「Unclear / Unfound」到「Ads」的转变都只会进一步巩固「Ads」SDK 作为可能指纹识别行为来源的主导地位。
从这张热力图得出的第二点观察是,有若干应用类别(「Finance」、「Food and Drink」、「Shopping」)中「Analytics」可能指纹识别 SDK 比其他类别的可能指纹识别 SDK 更普遍。我们假设,在这些应用类别中,指纹识别较少用于跟踪用户身份(身份可从用户账户信息得知),而更多用于理解用户对应用内所售商品的偏好。
通过指纹共享实现跨应用类别的可能跟踪
指纹识别带来的一个重大隐私风险是第三方有可能跨应用跟踪用户活动。当包含在多个应用中的某个 SDK 对用户进行指纹识别,从而允许将用户活动归属为同一个用户时,这种情况就会发生。某服务可能由此得知,例如,一个使用某种风格约会应用的用户,同时也使用某个特定的医疗或金融应用。这种跨应用跟踪可能发生在设备上或服务器上,两种情况下都由从共享 SDK 获得的指纹识别数据驱动。
为估计跨应用跟踪风险的下界,我们通过计算从每个应用类别中随机选择两个应用共享至少一个可能指纹识别 SDK 的概率,来分析存在于不同类别中的可能指纹识别 SDK 的普遍程度。结果在图 7(b) 中以热力图显示(只呈现下三角,因为热力图对称)。图中越深的蓝色表示共享指纹识别 SDK 的普遍程度越高,例如(「Game」,「Entertainment」)条目相比(「Travel and Local」,「Comics」)条目。我们对每个应用类别按总受众规模(如 §3.1 所定义)排名前 1000 的应用计算这些概率,无论这些应用是否包含可能指纹识别 SDK。因此,普遍程度(以及 7(b) 中所示的相关热力图)既反映了应用的受欢迎程度,也反映了可能指纹识别 SDK 在此类热门应用中的分布。
对这张热力图的分析清楚地表明,少数几个应用类别与许多其他应用类别共享可能指纹识别 SDK。例如,「Game」应用与「Art and Design」应用共享可能指纹识别 SDK,也与「Beauty」、「Books and Reference」、「Comics」、「Communication」、「Dating」、「Entertainment」、「Health and Fitness」、「Libraries and Demo」、「Lifestyle」、「Maps and Navigation」、「Music and Audio」、「News and Magazines」、「Personalization」、「Photography」、「Productivity」、「Social」、「Tools」、「Video Players」和「Weather」类别中的应用共享。类似地,「Personalization」应用与 33 个应用类别中的 7 个共享 SDK。从用户的角度看,这意味着如果他们同时安装「Game」和「Comics」类别的应用,跨应用跟踪的风险更高。
另一方面,某些类别的应用很少与其他应用类别共享可能指纹识别 SDK。我们特别指出「Finance」和「Medical」应用类别,因为这类应用常处理高度敏感的数据。若不做进一步研究,人们无法判断为何指纹识别在这里不那么普遍,我们注意到这类应用往往要求用户认证后才能访问其银行或投资账户或医疗记录,因此可能不需要通过间接信号对用户进行指纹识别。
5. 讨论与局限
我们的结果表明,指纹识别生态比此前估计的更复杂,无论是指纹识别行为还是指纹识别目的。我们在下面解读结果,讨论我们方法论的局限,并提出对应用安全机制的启示。
针对特定行业解决方案的挑战。
本文的一个核心观察是,当前移动生态已演化为通过若干类型的 SDK,在各种各样的应用中自发地部署类似指纹识别行为。我们的分析揭示,虽然广告 SDK 对指纹识别有贡献,但它们并非唯一的元凶:很大一部分类似指纹识别行为源自用于分析和反欺诈目的的 SDK,而相当一部分(23.9%)缺乏关于其目的或功能的充分公开信息以辨别其类别。这些 SDK 往往因直接与变现无关的原因而被集成——理解用户行为以改进应用,或防止机器人和欺诈——尽管它们最终收集了足以创建指纹的设备信息。
这一发现挑战了「指纹识别主要由应用开发者通过广告变现的需求驱动」这一流行观念,并凸显了以更广阔的视角看待保护隐私的替代方案的机会。例如,关于如何激励开发者采用保护隐私的分析的研究可能有用。沙盒化努力(例如 Android 的 Privacy Sandbox(1))也很可能为针对过度收集 SDK 的检测和执法提供额外收益——不过需要针对支持这一场景的轻量级沙盒化进行系统研究,因为把当前的基于进程的技术扩展到非广告 SDK 需要对于受限的移动环境而言不可承受的开销。
针对特定 API 或其他行为防御的挑战。
我们对种子集(自我标识的指纹识别 SDK)的探索得出的一个意外结果是,所使用的 API 空间是稀疏的;这些 SDK 从不同的 API 集收集信息。无论 SDK 的用例如何,或它在我们的应用数据集中的普遍程度如何,这一点都成立。一个可能的解释是 API 代理(API Proxying)(33),它可能进一步使任何反指纹识别执法复杂化,不过确定特定 API 的联合熵或共享熵超出了本工作的范围。
无论如何,针对特定 API(类似于 Apple 的 required reasons(4))似乎是一种脆弱、易被绕过的防御。开发者手头有大量信号,可以轻易转向其他熵来源。此外,尽管我们(意外地)在种子集中没有发现非 API、基于硬件的指纹识别的证据,但如果引入了 API 层面的全面执法,人们可能会预期开发者转向更先进的方法。
针对特定行业分析 & 定向的潜力。
值得注意的是,某些敏感的应用垂直领域似乎有更好的隐私立场。按安装量归一化后,医疗类别中只有 30% 的应用使用了指纹识别 SDK,并且(假设样本集之间服从正态分布)其中只有 19.5% 使用了广告 SDK——大部分可识别的指纹识别行为似乎来自分析。医疗应用跨应用跟踪的潜力也似乎较低。这一令人欣慰的结果在金融类别中大体重演,凸显了未来工作需要聚焦于针对特定敏感市场垂直领域的解决方案。
多平台分析的必要性。
我们的结果很可能也适用于 iOS 生态——事实上,我们种子集中的所有 SDK 似乎都有现成的 iOS 版本——这一发现与此前关于跨平台跟踪的工作(35)一致。然而,在 iOS 上很难进行此类分析,因为 Apple 的应用及操作系统级 DRM 限制了第三方对其 App Store 中应用进行可扩展的静态和动态分析的能力。未来研究 iOS 生态将为理解两个操作系统之间设计选择的有效性提供宝贵洞见。
5.1 局限
任何经验研究,包括本文,都是对现实世界状况和趋势的有限视角,因此评估威胁其有效性的因素很重要。遵循「坎贝尔传统」(Campbell Tradition)(11),我们考虑四种有效性——内部、统计、构念和外部——及其对本研究的影响。
内部有效性(internal validity)指的是所测量的效应(可指纹识别的 API)是否真正对应于所关注的结果(指纹识别行为)。一个风险是,使用 API 获取高熵数据可能并非由有意的指纹识别行为引起,而是应用必要功能的结果。我们通过聚焦于记录可指纹识别数据收集的目的来规避这一点。第二个局限存在于我们种子集中隐含的选择偏差,它由自我标识为指纹识别的 SDK 组成。可能那些出于隐藏原因进行指纹识别的 SDK 使用替代技术,而这不会在我们的后续分析中被捕获。我们假设种子集 SDK 的自我报告是诚实的,不对 SDK 的意图做进一步推断。
统计有效性(statistical validity)指的是统计功效不足实验的风险,即缺乏充分的统计支持。我们 228,598 个 SDK 和 3,025,417 个应用的大样本量减轻了这一风险。
构念有效性(construct validity)指的是衡量可指纹识别 API 和行为存在与否的度量选择。我们聚焦于 API 数量作为指纹识别行为的高效度量,不过我们注意到并非所有 API 对指纹识别都同等有用。目前我们做一个简化假设:现实世界中的技术大体相当,且所收集的信号之间没有关系。使用更复杂的度量(如碰撞熵(collision entropy)(10))需要在大量设备和用户上进行实验,我们将其留待未来工作。
外部有效性(external validity)指的是我们的结果对现实世界的可推广性。我们从流行 Maven 仓库选择真实 SDK、从 Google Play 商店选择移动应用,确保将这一风险降到最低。然而,我们并未尝试编目所有移动指纹识别生态,而是把自己限制为属于 Android/AOSP 框架的 Java 语言 SDK(排除非平台 API 或来自 OEM 的 API)。通过网页搜索手工挑选的指纹识别 SDK 种子集可能无法代表现实世界中的所有指纹识别行为,需要进一步研究以确保全面的视角。
6. 结论
在本文中,我们呈现了迄今所进行的最大规模的 SDK 行为分析,考察了超过 228,000 个 SDK 和 178,000 个 Android 应用,以理解类似指纹识别行为的普遍程度和目的。我们的发现揭示,除那些明确为广告设计的 SDK 之外,大量 SDK 收集了足以潜在地跟踪用户的信息。这包括用于分析和反欺诈的 SDK,凸显了在这些领域需要保护隐私的替代方案。令人意外的是,表现出类似指纹识别行为的一大部分 SDK 缺乏明确标识,凸显了 SDK 生态中需要更大的透明度。此外,我们观察到这些具有类似指纹识别行为的 SDK 受欢迎程度不成比例地高,且常常跨越各种应用类别被集成。这些结果凸显了 Apple 和 Google 持续努力增强用户隐私的重要性,并强调需要继续研究以确保此类业界努力方向得当。
附录 B:我们静态分析的细节
Android SDK 的依赖分析
目标软件开发工具包(SDK)依赖其他 SDK 是一种常见做法。例如,一个广告 SDK 可能为指纹识别收集用户数据,然后用另一个 SDK(例如 OkHttp)将数据外发共享。我们把目标 SDK 称为主 SDK(main SDK),把主 SDK 所使用的 SDK 称为它的依赖 SDK。准确推断并将每个主 SDK 的依赖整合进分析中,对于彻底检测至关重要。
为推断每个主 SDK 的依赖,我们分析仓库中所有 SDK 的 Maven 项目对象模型(POM)文件,提取每个文件中引用的 SDK。从每个 SDK 的 POM 文件信息,我们构建一个依赖图,每当一个 SDK 的 POM 文件引用另一个 SDK 的 POM 文件时就有向边。在依赖图中,一个不同的 SDK 版本由三元组表示:指示开发者的 group ID、指定 SDK 名称的 artifact ID,以及版本号,通常写作 X:Y:Z,表示来自开发者 X 的 SDK Y 的版本 Z。在解析依赖图时,可能发生版本冲突。例如,一个 SDK M:A 可能同时引用 SDK N:B:1 和 P:C:1。然而,SDK N:B:1 可能需要 SDK P:C:2。这意味着对于 SDK P:C,版本 1 和 2 都被列为 SDK M:A 的依赖。由于把两个版本都导入静态分析可能导致一个版本覆盖另一个、造成非确定性行为,我们为每个 SDK 只保留一个版本。基于奥卡姆剃刀原则,我们优先选择在依赖图中距主 SDK 路径最短的版本,在上述例子中即 P:C:1。
在分析过程中,我们把主 SDK 中的公共方法配置为分析的入口点,这意味着只有由这些公共方法之一触发的执行路径和行为才会被报告。
静态污点分析
我们采用静态污点流分析来发现潜在的指纹识别实例。该过程首先用元数据污点标记 PII 数据,使系统能够在程序代码中跟踪其移动。为实现这一点,静态污点流分析包含几个关键阶段。首先,系统扫描输入程序,以精确定位可能的污点源(程序访问 PII 的 API 方法调用)和汇点(sink,数据离开设备的 API 方法调用)。
一旦确定了源和汇点,就传播污点标志以构建污点流图。该图使用节点表示程序元素(例如寄存器或字段),边则封装污点数据的可能传输。图增量扩展,直到达到不动点。一条连接源与汇点的检测到的污点流路径标志着潜在的 PII 外传。
图 8. 静态污点分析 —— 识别源与汇点
图 9. 静态污点分析 —— 污点传播
CoFlow 分析
CoFlow 分析是一种专门的污点流分析,旨在基于「交叉流」(crossover flows)识别可疑的应用行为模式。与传统污点流分析聚焦于数据源与汇点之间一对一关系不同,CoFlow 分析检测多个源汇聚到单个汇点的场景。这种对多对一关系的关注使它有助于揭示指纹识别或 ID 桥接(ID bridging)等行为。
CoFlow 分析构建于污点流跟踪能力之上,利用我们的静态污点分析过程跟踪数据在整个应用中的移动。它应用一组约束和规则来精确定位那些匹配所配置的可疑行为定义的流模式。CoFlow 分析使用一个专用配置文件来定义其行为检测规则。该文件指定源 API 集合和对应的汇点 API,并包含细粒度约束以最小化误报。
分析过程的核心是确保对于每种配置的行为,每个已定义的源组中至少有一个源参与该行为。系统检查潜在的汇点,向后追踪数据的流动,并将发现的源与规则集比较。如果一条规则中的每个源组都找到匹配,系统就将其标记为该可疑行为的一个实例。存在优化来精简这一过程并提高效率。当 CoFlow 分析检测到可疑行为时,它会生成输出,包含一组源(每个配置的源组各一个源)及关联的汇点。
图 10. CoFlow 分析
指纹识别检测
指纹识别检测利用 CoFlow 分析来识别 SDK 中的指纹识别行为。一个「指纹识别行为」被定义为 N 个或更多指纹识别源流向一个共同的关注汇点。该分析从我们在种子集 SDK 中观察到的指纹识别 API 列表出发,检查至少有 N 个来自指纹识别 API 的数据项在可能被组合成新数据对象之后被外传。指纹识别的 CoFlow 源组配置为自我报告的指纹识别 SDK 所收集的信号集合。源的数量相当大,我们预期大多数指纹识别 SDK 的标识符只会包含这些源的一个子集。指纹识别的汇点配置为两组:网络 API(造成指纹通过网络外传的可能性)和加密函数。还存在其他类别的汇点,但未包含在本工作中。例如,指纹识别源可能被收集到一个 map 或 JSON 对象中,并简单地由一个公开可见的 SDK 方法返回。类似地,一个 SDK 可能用指纹识别源填充一个通过引用传递给公共方法的参数。
指纹识别与其他 CoFlow 用例略有不同之处在于:检测到的 CoFlow 越多,我们对行为存在的信心越高。因此,对于指纹识别行为,在输出中包含所有检测到的源(而不是通常的每组一个源)是有用的。包含所有检测到的源也允许对 SDK 语料库中的指纹识别行为进行更彻底的分析,并为调试和逆向工程以确认行为提供有用信息。
挑战
我们在与其依赖、以及依赖的依赖等打包在一起的 SDK 上运行分析。我们第一版指纹识别检测包含了来自整个 SDK 与依赖包络的污点源。由于这些包络可能变得相当大,而污点跟踪随着源数量的增加表现出超线性行为,我们发现我们的分析对于非常大的 SDK 以及依赖数量多的 SDK 而言开销高得难以承受。
此外,跟踪源自依赖的污点源会导致识别出完全包含在某个依赖内、且可能根本未被主 SDK 使用的指纹识别行为。这种「过度检测」非常嘈杂,我们看到许多 SDK 因把相同的流行指纹识别 SDK 作为依赖包含进来而被标记。识别哪些 SDK 把指纹识别 SDK 作为依赖包含是有价值的,但用昂贵的静态分析技术来做这件事并不高效,而且噪声可能掩盖主 SDK 所进行的指纹识别。
为解决这些问题,我们把污点源的范围缩小到仅源自主 SDK 的那些,排除所有源自依赖的源。值得注意的是,如果汇点在依赖中,仍会包含该汇点,这使我们能够捕获 SDK 收集指纹识别源但使用依赖来哈希或外传它们的情况。我们承认这一权衡意味着我们会漏掉一些指纹识别情形,主要是某个 SDK 使用多个依赖、每个依赖获取的源数低于指纹识别阈值,但组合到主 SDK 中时达到阈值的情况。能够检测出 SDK 中不感兴趣、可排除在分析之外的边界,以及一些动态分析技术,可能有助于填补这一空白。
图 11. 使用 CoFlow 进行指纹识别检测。本例中有 2 个指纹识别行为。源 {1,2,3} → SinkA,源 {3} → SinkB。源 4、5 和 6 未包含在结果中,因为它们处于某个依赖中。
附录 C:我们 SDK 识别方法的细节
任何 SDK 识别方法都必须不依赖对源代码的访问(因为应用主要以二进制格式分发),必须不假设所有 SDK 代码都存在于应用中(因为编译和链接常常会压缩、优化或移除 SDK 代码),必须处理代码混淆(在 Android 应用中普遍存在,超过 75% 的应用使用混淆),必须处理 SDK 依赖(在我们分析的 30,000 个 Java SDK 中,第 80 百分位的 SDK 依赖 17 个其他 SDK),并且必须可规模化运行(正如 modulecounts.com 的数据所示,SDK 和版本的数量在 10^4–10^7 量级,每天新增 10^2–10^3 个新 SDK)。
我们对 SDK 识别的解决方案使用一种细粒度的代码相似度度量,它可以跨代码单元(例如方法、类、包)和分发单元(例如 SDK、SDK 家族)聚合。方法的相似度度量基于以下特征集的相似度:方法的类型签名、方法体中所调用的 Java 和 Android 框架 API、方法体所使用的字符串常量,以及方法体中存在的指令类型直方图。
- 方法的类型签名捕获其参数类型和返回值类型。由于 Java 和 Android 框架之外的所有类型都由开发者定义,其名称不可信,不能用于相似度比较。我们消除所有程序员选择的类型名称,以获得方法的「匿名化」类型签名。所得类型签名被独热编码为一个(布尔值)特征,嵌入我们的向量空间。
- 方法体中所调用的 Java 和 Android 框架 API 各自成为一个特征,计数值为方法体中出现的调用次数。这样的特征不考虑间接调用框架 API 的方式,例如通过 Java 反射、通过其他语言的代码(例如原生或 JavaScript 代码),或通过运行时加载的动态代码。
- 方法体中使用的字符串常量各自形成一个(布尔值)特征,类似于方法的匿名化类型签名。
- 方法体中存在的指令类型各自形成一个特征,度量该方法体中该特定指令类型的频率。我们考虑 34 种指令类型,与 Dalvik VM 规范中的指令紧密对应,范围从 ASSIGN、LOAD_INSTANCE 到 THROW。
方法 m 的特征向量 fv(m) ∈ ℝ^(1×d) 是通过组合上述所有特征并将它们嵌入一个 2^64 维空间得到的,嵌入方式是用一个 64 位哈希函数对每个特征名称做哈希。该哈希函数需要抗碰撞,以确保对手不能轻易创建与某个目标 SDK 代码相似的代码来规避 SDK 识别,但不一定需要抗原像,因此不必是密码学哈希。这一嵌入步骤为每个 Java 方法构造一个稀疏向量表示,因为 Java 方法通常只有数百个非零特征。
一项初步数据分析表明,这些特征在不同 SDK 之间是非均匀分布的。图 12 最左侧的柱子表明,在约 37,000 个 SDK 的语料库中,约 11,000,000 个特征恰好出现在一个 SDK 中,因此可以用于使这些 SDK 在应用代码中被唯一地重新识别。因此我们为每个特征添加一个权重因子,以计入这种非均匀性:
wfv(m) = fv(m)^T × weights。
图 12. SDK 代码的特征在大量 SDK 中并非均匀分布。只出现在单个 SDK 中的特征对 SDK 匹配特别有用,因此在任何用于 SDK 识别的相似度度量中需要被赋予更大权重。
我们把 Java 方法表示到其中的高维空间,自然允许我们处理代码组合。一组方法(无论是组织为一个类、一个包、一个 SDK 还是其他代码结构)的向量表示是对应向量的向量和。我们对两个向量 fv1 和 fv2 使用余弦相似度作为距离度量:
δ(fv1, fv2) = (fv1 · fv2) / (‖fv1‖ ‖fv2‖)。
每个类被赋予相应的向量表示,即其组成方法向量的向量和,且为简写我们记 fv(c) 表示类 c 的向量,使得 fv(c) = Σ_{m∈c} fv(m)。然后,对于应用中的每个类,我们在所有 SDK 类的集合中搜索与该校相似度最小的类。最后,我们匹配在应用中存在性得到充分支持的 SDK,其度量依据是 SDK 中与应用类相似的类的数量。算法 1 给出了我们技术的伪代码版本。
SDK 匹配算法由一个阈值参数化,以确保应用类和 SDK 类满足最小相似度;并由第二个阈值参数化,以确保如果 SDK 有足够数量的类出现在应用中,则该 SDK 匹配。在我们的实验中,我们使用 0.2 作为相似度阈值,0.55 作为类计数阈值。
把方法和类表示为高维空间中的稀疏向量,使我们能够借用近似最近邻(ANN)问题的解决方案,这在信息检索领域得到了充分研究。特别是一些高性能库,如 Faiss [20]、SPTAG [43]、ScaNN [28]、Hnswlib [63] 和 NGT [62],提供了在数百万 SDK 类的数据库中搜索典型应用中数千个类的匹配项的能力。
我们在一组 302,397 个 SDK 上评估了该技术的准确性,对于这些 SDK,我们能找到至少一个在其 Gradle 构建配置中声明该 SDK 为依赖的移动应用。Gradle 配置数据提供了我们据以衡量 SDK 识别平均精度的真实基准,结果为 65.07%。在限制到较小的 9,683 个 SDK 集合时,SDK 识别在 46.16% 的召回率下达到 99.89% 的精度。这意味着在这个选定的 SDK 集上,该识别技术能准确检测出某个 SDK 在应用中不存在,尽管它有时可能无法检测出 SDK 在应用中的存在。
在论文其余部分的测量语境中,这种 SDK 识别技术的应用为市场覆盖(参见 RQ3)提供了下界,因为它低估了所关注 SDK 的存在。
算法 1 SDK 匹配算法
inputs : A // an app
S1,…,Sm // a set of SDKs
η // class-dissimilarity upper bound
γ // SDK-similarity lower bound
outputs: 𝒯 // map from app classes to SDKs,
// 𝒯 : A ↦ 2^(S1∪⋯∪Sm)
begin
let A = {c1,…,cn} be the set of app classes
let 𝒞 = S1 ∪ ⋯ ∪ Sm be the set of SDK classes
// compute candidate SDKs for each app class
foreach ci ∈ A, 1 ≤ i ≤ n do
𝒯(ci) ← {c ∈ 𝒞 : δ(fv(ci), fv(cyx)) < η}
// filter out SDKs with insufficient presence
for j ← 1 to m do
if |Sj ∩ (𝒯(c1) ∪ ⋯ ∪ 𝒯(cn))| ≤ γ then
for i ← 1 to n do 𝒯(ci) ← 𝒯(ci) \ Sj
end
附录 D:长式定义
表 6. 我们标注过程所创建类别的编码手册定义。
| 类别 | 定义 |
|---|---|
| Ads | SDK 的主要目的是支持展示广告、广告竞价、广告定向、广告中介,或出于变现或转化明确目的的用户分析。示例:中介、广告适配器、广告推送通知 |
| Analytics | Analytics(应用健康):SDK 的主要目的是跟踪应用的系统性能。示例:崩溃日志、电池跟踪与使用、UI 延迟。Analytics(用户行为分析):SDK 的主要目的是为商业目的(例如增长、留存、参与、流失)在设备上跟踪用户行为,且没有直接变现或把所收集数据反馈进广告的证据。例如:应用参与度测量 |
| Security & Authentication | SDK 的主要目的是保护应用或用户免受恶意软件和欺诈。Security(反欺诈):为应用提供欺诈检测能力的 SDK,以处理账户欺诈、支付欺诈或身份欺诈。欺诈检测可在设备上或远程服务器上执行。Security(支付):提供支付功能的 SDK,可能管理支付流程的所有方面(授权、清算、结算、冲正)。此处涵盖所有支付方式。Security(认证):使用应用针对本地或远程账户或身份对用户进行认证的 SDK。此处涵盖所有形式的认证机制(密码、生物识别、硬件令牌、基于知识)。Security(反恶意软件):尝试扫描恶意软件、bug 或其他设备上安全问题的 SDK。Security(其他):不能明确归入上述之一的 SDK。 |
| Tools / Other | Location:SDK 的主要目的是访问位置数据并将其发送到远程服务器。Location Tracking (Person):SDK 的主要目的是出于收集商业数据的目的(例如某人访问某地点的频率,如 Foursquare)或出于定位家庭成员的目的跟踪用户位置。Location Tracking (Object):特定物体跟踪(例如帮助跟踪 beacon/标签以寻找丢失物品的 SDK)。Location Tracking (Maps):任何形式的导航功能,包括显示地图和地图路线,以及基于位置检索兴趣点。Location Tracking (Other):不能明确归入上述之一的 SDK。Social:SDK 的主要目的是将用户连接到更大的社交网络,或提供发现他人的手段,或提供用于社交互动的联系人列表。客服聊天 SDK 不归入社交类别。Other:SDK 的主要目的是支持或提供与任何其他已定义类别无关的功能。该 SDK 必须在信息来源中有明确定义的目的。 |
| Unclear / Not found | 使用任何允许的信息来源中的数据,SDK 的主要目的都不明确,或者我们找不到关于该 SDK 的任何信息。注意此类别与「Tools / Other」不同,后者捕获的是目的明确定义的 SDK。 |
表 6 提供了我们标注过程中所制定和使用的定义的完整列表。该表基于软件开发工具包(SDK)的主要功能定义了一套分类系统,将其分为五个主要类别:Ads、Analytics、Security and Authentication、Tools/Other 和 Unclear/Not Found。「Ads」类别涵盖与广告相关的 SDK;「Analytics」包括用于跟踪应用健康和用户行为的 SDK;「Security and Authentication」涵盖用于反欺诈、支付、认证和反恶意软件的 SDK;「Tools / Other」包括位置跟踪(人、物体、地图)、社交网络以及具有明确但未归类目的的其他 SDK;而「Unclear / Not Found」用于无法确定 SDK 目的的情况。一些类别进一步细分为具有特定定义和示例子类别,例如「Analytics (App Health)」或「Security (Anti-Fraud)」,以确保精确分类。
我们的编码员使用这些信息来对 SDK 进行分类:首先审阅关于该 SDK 目的和功能的可用信息。然后他们把这一信息与表中提供的定义比较,从主要类别开始,再到子类别。通过把 SDK 的特征与表中的描述匹配,评分者分配最合适的标签。如果 SDK 的目的明确定义但不匹配任何现有类别,他们使用「Other」子类别。如果找不到任何信息,或可用信息不足以确定 SDK 的目的,评分员将其归类为「Unclear / Not Found」。这一系统化方法确保了基于 SDK 主要功能的一致而准确的标注。
【译注】本字段完整翻译了论文的摘要、正文第 1–6 节(引言、背景与相关工作、方法论、结果、讨论与局限、结论)以及附录 B、C、D。论文的参考文献列表(约 2.8 万字符的书目条目)与附录 E(表 7,指纹识别行为中观察到的信号 API 清单,均为 Java 类名与方法/属性名等标识符)按翻译规则第 5 条「代码/API 名与文献标题保留原文」保留英文原文,未在本字段重复;两者均完整保存在 _work/raw 的 content_text 中。附录 A(应用类别原始数据)为图表数据,亦保留于原文。
arXiv is now an independent nonprofit!
Learn more
×
Back to arXiv
License: CC BY-NC-ND 4.0
arXiv:2506.22639v1 [cs.CR] 27 Jun 2025
Fingerprinting SDKs for Mobile Apps and Where to Find Them: Understanding the Market for Device Fingerprinting
Michael A. Specter
Georgia Tech & Google LLC, Atlanta, GA, USA Email address: mobile-fingerprinting-sdks-paper@google.com // specter@gatech.edu
Mihai Christodorescu
Google LLC, Mountain View, CA, USA
Abbie Farr
Google LLC, Mountain View, CA, USA
Bo Ma
Google LLC, Mountain View, CA, USA
Robin Lassonde
Google LLC, Mountain View, CA, USA
Xiaoyang Xu
Google LLC, Mountain View, CA, USA
Xiang Pan
Google LLC, Mountain View, CA, USA
Fengguo Wei
Google LLC, Mountain View, CA, USA
Saswat Anand
Google LLC, Mountain View, CA, USA
Dave Kleidermacher
Google LLC, Mountain View, CA, USA
Abstract.
This paper presents a large-scale analysis of fingerprinting-like behavior in the mobile application ecosystem. We take a market-based approach, focusing on third-party tracking as enabled by applications’ common use of third-party SDKs. Our dataset consists of over 228,000 SDKs from popular Maven repositories, 178,000 Android applications collected from the Google Play store, and our static analysis pipeline detects exfiltration of over 500 individual signals. To the best of our knowledge, this represents the largest-scale analysis of SDK behavior undertaken to date.
We find that Ads SDKs (the ostensible focus of industry efforts such as Apple’s App Tracking Transparency and Google’s Privacy Sandbox) appear to be the source of only
30.56
% of the fingerprinting behaviors. A surprising
23.92
% originate from SDKs whose purpose was unknown or unclear. Furthermore, Security and Authentication SDKs are linked to only
11.7
% of likely fingerprinting instances. These results suggest that addressing fingerprinting solely in specific market-segment contexts like advertising may offer incomplete benefit. Enforcing anti-fingerprinting policies is also complex, as we observe a sparse distribution of signals and APIs used by likely fingerprinting SDKs. For instance, only
2
% of exfiltrated APIs are used by more than
75
% of SDKs, making it difficult to rely on user permissions to control fingerprinting behavior.
1.Introduction
Device fingerprinting is a technique used to identify and track user devices by collecting a wide range of information about device-specific hardware, software, and configuration settings. The combination of these attributes creates a unique, or near-unique, digital “fingerprint” for that device. This process has clear privacy concerns—fingerprinting identifiers can be collected without user control or notice, and persist over the device’s lifetime regardless of most user privacy-seeking actions (e.g., clearing one’s cookies, rotating advertising IDs, or enabling private browsing).
Both major mobile operating system vendors have undertaken significant efforts to limit the privacy impact of device fingerprinting. Apple has introduced policies that require apps to request user consent to collect tracking-relevant device data (5), and provide human readable explanations for the use of specific high entropy “required reason” APIs (4). Both Google and Apple’s mobile platforms now require developers to provide nutrition label-style privacy information to the user (3, 25), either as metadata submitted to their respective application stores or attached as part of the application itself. Google is developing a privacy sandbox for the web (1), and, on Android, a new sandbox that restricts third-party advertising libraries from accessing sensitive information available to the rest of the application (26). These interventions are promising—providing much needed transparency and accountability.
The success of such anti-fingerprinting efforts depend on the technical implementation, market fit, and intention of the application’s developer. For example, Apple’s anti-tracking and app transparency policies explicitly allow the collection of fingerprinting data for anti-fraud purposes (5), and Android’s Privacy Sandbox focuses solely on isolating code from third-party advertisers. Apple’s “required-reason APIs” approach also has limitations; it currently applies to just 30 APIs, and its effectiveness depends on what other data points are collected across the wider fingerprinting ecosystem. Therefore, characterizing the technical implementation of fingerprinting in the wild and understanding stakeholders involved would provide invaluable insight into the effectiveness of these interventions.
This paper presents a comprehensive, large-scale analysis of device fingerprinting practices within the Android application ecosystem. We adopt an empirical approach centered on the identification of third-party Software Development Kits (SDKs) integrated in mobile applications, measuring their market reach, tracking methodologies, and privacy impact. To the best of our knowledge, this research represents the most comprehensive analysis of SDK behavior regarding privacy-invasive practices like device fingerprinting, with an extensive dataset of over 228,000 unique SDKs and 178,000 Android applications. Our methodology attempts to enable a more nuanced understanding of the scale and scope of device fingerprinting in mobile ecosystems; we avoid applying our own potentially biased or narrow definitions on fingerprinting behavior, and adopt a number of techniques to provide reliability and consistency when subjective analysis is unavoidable.
While many studies have measured the impact of fingerprinting (see §2 for related work), there is a dearth of knowledge surrounding the purpose of fingerprinting behavior. For example, no prior study has attempted to understand what kinds of third parties collect sufficient information to fingerprint a device, and how this fits in with the needs of the first party developer. Characterizing the overall problem from the perspective of developers can help determine why these techniques are used, provide invaluable insight into what is required to better preserve user privacy, and interpret the value of assumptions underlying current and proposed enforcement methods.
There are a number of challenges that significantly complicate our study. Any fingerprinting-detection mechanism will likely be incomplete, as there are many (potentially stealthy) methods of collecting entropy from a device, including timing information, instruction execution quirks, and other hardware-specific sources. Categorization and analysis of SDKs is also a difficult task—while applications self-label their use and market-fit, current SDK distribution methods do not require SDK authors to provide significant descriptions of their code. We describe the solutions to these and other challenges in depth in §3.
One important challenge is definitional: The claim that a service is fingerprinting suggests intent of the author of the code, which is most often practically unknowable. Applications may collect sufficient information to uniquely identify a device for any number of reasons, including analytics, crash reporting, anti-fraud, or through normal operation of the application itself. We emphasize that our study is purely observational—we measure fingerprinting behavior, and ascribe no motivation to the authors of the code. We also emphasize that our choice to examine the Android ecosystem is entirely due to convenience, and that our results are likely to extend to iOS as well. As noted in prior work (35), whereas Android’s open ecosystem allows for scalable analysis, iOS’s digital rights management scheme actively hinders the same.
We answer the following research questions:
RQ1::
What types of behaviors do self-identifying fingerprinting SDKs exhibit?
RQ2::
What are the stated purposes of SDKs with likely fingerprinting behavior?
RQ3::
What kinds of apps use SDKs with likely fingerprinting behavior, and how prevalent are these SDKs in real world apps?
We find that many kinds of SDKs collect sufficient information to track a user (at least
20
signals exfiltrated per SDK), and that there is a large diversity in signals collected (SDKs exfiltrate
75.5
signals on average, out of a total of
504
unique signals observed across the SDK dataset). Though ads do make up a significant portion (
≈
30
%) of the SDKs that exhibit fingerprinting behavior, a surprising number of fingerprinting-like SDKs used in common Android applications have unclear functionality and lack significant description for categorization (
≈
24
%). Anti-fraud and analytics services were also prevalent in our dataset, indicating that more research must be done to create privacy-preserving alternatives to fingerprinting as used in such functionality. Finally, SDKs that exhibit likely fingerprinting behavior are disproportionately popular—roughly
10
×
more installs than non-fingerprinting alternatives—and individual SDKs are likely to exist across multiple application market segments (e.g., health and dating).
Figure 1. An overview of our analysis pipeline. We begin by \raisebox{-0.5pt}{\sffamily\small1}⃝ fetching apps, SDKs, and associated metadata from a series of Maven repositories and the Google Play app store, only selecting applications installed on
10
𝑘
active devices (§3.1). We continue by \raisebox{-0.5pt}{\sffamily\small2}⃝ extracting a seed set of SDKs whose copy indicates that they are fingerprinting (§3.2). We then \raisebox{-0.5pt}{\sffamily\small3}⃝ use static taint analysis to determine which SDKs exfiltrate these signals (§3.3), and \raisebox{-0.5pt}{\sffamily\small4}⃝ manually label the resulting SDKs to determine their market fit (§3.4). Finally, we \raisebox{-0.5pt}{\sffamily\small5}⃝ perform another round of static analysis to determine which applications contain which SDKs (§3.5).
2.Background & Related Work
To the best of our knowledge, our work is the first large-scale study of native app-based device fingerprinting in the wild.1 No prior research has attempted to understand why this phenomenon is common, or the market surrounding the use of these tools. Much of the existing literature stems from examining fingerprinting as an attack, with a focus on novel methods of fingerprinting.
Android applications & SDKs
Android applications may be written in any language, and can be installed from arbitrary sources including the Play store, secondary app stores, side-loading, or may come pre-loaded on-device from the manufacturer. As a result, much of Android’s security model revolves around sandboxing applications using a combination of SELinux’s SEPolicy and standard Linux UID-style access control mechanisms. In addition to sandboxing, access to certain sensitive data is declared via metadata provided by the app, and enforced via both install and run-time permissions checks (42). Unless manually sandboxed, Android third party libraries (called SDKs) execute in the first-party application context, and therefore enjoy the same permissions as the first-party application.
SDKs may be distributed as either raw code or automatically downloaded at build time from any number of repositories or build systems. A commonly used build system is the Maven format, an open standard for Java dependency resolution. In practice, distribution of SDKs is commonly done using a build tool called Gradle, which loads SDKs from any number of public Maven repositories.
Fingerprinting signals
In this paper we define a signal to be an individual data point collected from a device. Different papers have different terms for a unique datapoint in a fingerprint, Eckersley (18) calls it a variable. Signals may be obtained from many sources including API calls, common files on the platform, system properties, hardware quirks, or runtime environment values.
Diversity of signals
Though fingerprinting is commonly discussed in the context of browsers (19, 18), there is a rich literature surrounding the many methods of fingerprinting via native code. Fingerprintable components include the microphones and speakers (14, 70, 9), the accelerometer, the gyroscope, and the magnetometer (69, 68, 54, 16, 9, 40, 57), the hardware clock (34, 51), the camera (8, 49), the GPU clock (44), and the battery (13, 47). Fingerprinting of the system or the apps ranges from specific APIs (48), to system configuration (e.g., via procfs (52, 55)), user settings (37), and browser configuration (18, 53). Fingerprinting can also be done by communicating with neighboring devices via short-range radio protocols such as Bluetooth (36). Distinct signals can be combined to increase accuracy (2, 12).
Detection & prevention
Detection and prevention methods include ML classification of API calls (7, 21, 31), re-calibrating sensors (15), changing system settings (32), adding random noise to collected data (45, 15), or using taint tracking to identify exfiltration of fingerprintable data (39). Permissions systems do not offer adequate protection against fingerprinting (60, 17).
Prior measurement studies
There have been a number of studies that measure the use of fingerprinting, though the majority focus on the web (19, 46). A significant challenge appears to be the ever-growing surface of APIs, which have been quickly adopted by fingerprinters (7).
Existing analyses of Android native systems are comparatively rare. Longitudinal analyses not only highlighted this distinctive, mobile-specific flavor of SDK-based tracking in general, but also showed that the privacy risk across time and app versions varies greatly with little correlation to existing enforcement approaches (50). Han et. al. (29) find that the presence of privacy-risky behavior (including fingerprinting) does not seem to vary by the cost of the app, with free apps and paid apps sharing similar sets of third-party SDKs or dangerous permissions.
The closest study to ours is Torres et al’s 2018 work (22) on identifying fingerprinting in applications. They find that fingerprinters on mobile devices rely more on categorical signals and less on side channels, and argue that detection and prevention of fingerprinting on mobile is distinct from web browsers. Our work uses a significantly increased scale in terms of signals, SDKs, and applications (30k vs our 178k), presenting a more complete understanding of the ecosystem, in addtion to SDK labels and further statistics.
3.Methodology
In this section, we provide an in-depth discussion of our analysis pipeline. We depict our overall process in Figure 1, and outline a summary below:
(1)
Dataset Collection (§3.1): We fetch a dataset of SDKs by crawling popular Maven repositories and a dataset of applications from the Google Play Store.
(2)
Seed Set & Manual Signal Extraction (§3.2): We use the advertising copy published by each individual SDK to generate a Seed Set that in public statements self-announce as fingerprinters. We then manually reverse engineer these SDKs to extract what signals each exfiltrate.
(3)
Automated Signal Exfiltration Detection (§3.3): Using static taint analysis in tandem with data generated from our Seed Set, we determine if an SDK exfiltrates relevant fingerprinting signals. We call all SDKs that perform enough exfiltration to be exhibiting fingerprinting behavior the Extended Set.
(4)
SDK Labeling & Analysis (§3.4): Unlike applications, SDKs are unlabeled, and do not carry metadata associated to market, use-case, or intended audience. To provide adequate statistics and market information, we manually label all SDKs in the Extended Set. To avoid bias and meaningless labels, we borrow coding techniques from the HCI community, iteratively developing a codebook and reaching consensus on SDK label definitions and assignments.
(5)
App-SDK Matching (§3.5): To provide statistics on the use of each SDK, we must first determine which SDK exists in which application. We accomplish this through a series of static analysis techniques.
There are a number of reasons we focus on SDKs rather than applications as a whole. One may expect an SDK to be self-contained and have explicitly publicized functionalities. As with any modern development environment, SDKs use in applications is incredibly common, with the majority of apps using third-party libraries to support a variety of core functions. Finally, an emphasis on SDKs also makes the potential harm from third parties far clearer: users are more likely to understand and trust an application they have actively installed, but may be unaware of the transitive trust they have placed on the third party SDKs and services used by an app.
3.1.Dataset Collection
Application Dataset
We collected 3,025,417 APKs2 published on the Google Play store over almost 18 months (from January 2023 to May 2024). We supplement this set of APKs with each application’s total audience size – the number of active devices that an individual APK has been installed on. An active device is a device that has been turned on at least once in the previous 30 days (24).
To avoid biasing our sample set with applications that lack a significant user base, we limit our analysis to applications active on the Google Play store with a total audience size of over 10,000 from April 13, 2024 to May 13, 2024. In total, this covers 178,054 applications. While we cannot estimate how often these applications were launched, the audience-size metric assures that devices on which the apps were installed were in active use. We approximate the market reach of an app by summing all active 30-day installs of these apps — we note that this may double-count users who install, remove, and then reinstall the same app, as well as installs by the same user on one device under multiple profiles (e.g., personal and work) or on multiple devices (e.g., phone and tablet).
SDK Dataset
There is no single source of information for SDKs, instead developers supply their build system with a URL of a particular Maven repository and library ID. We collate our dataset using a custom crawler, extracting all SDKs from 9 separate large-scale Maven repositories: JCenter, Maven Central, Google, Sonatype, Spring.io, Jitpack, Bintray, and Artifactory. From these repositories, we fetched a dataset of 228,598 SDKs as well as each SDK’s associated metadata. We excluded all SDKs whose version label included one of the words {“alpha”, “beta”, “test”, “dev”, “debug”, “qa”}.
3.2.Seed Set and Manual Signal Extraction
From our dataset of SDKs we select a Seed Set of SDKs that, in their advertising copy or other metadata, openly admit to collecting information for the purpose of fingerprinting. We then manually reverse engineered each SDK, confirmed that the SDK was collecting a nontrivial set of device data, and extracted a list of signals that the SDK uploaded to a server. To avoid mislabeling, each candidate SDK is then reverse engineered again by a second analyst, who independently confirms the list of signals. In total, our Seed Set contains 14 SDKs, reporting over 500 distinct signals. The results from this effort are reported in more detail in §4.1.
3.3.Automated Signal Exfiltration Detection
We collected an Extended Set of fingerprinting-like SDKs, consisting of SDKs that are similar in terms of collected data to the Seed Set. To obtain this set, we developed a static analysis suite that performs information-flow analysis on SDKs and their dependencies, and selected SDKs with sufficient signal overlap with our Seed Set. Note that SDKs in the Extended Set may collect signals beyond those found in the Seed Set.
To avoid over-claiming the existence of fingerprinting behavior in our dataset, we only include an SDK if it exfiltrates more than the lowest number of signals collected by any SDK in our Seed Set. Put another way, we only consider an SDK to be exhibiting fingerprinting behavior if it exfiltrates more than what is uploaded by an SDK that openly admits to fingerprinting. Note that this is a conservative estimate—it is likely that more sophisticated estimates of entropy would indicate that an SDK could uniquely identify a user using less information than what we examine here. The tradeoff here is intentional; our goal is to provide upper-bar estimates without indulging in more complicated analyses. This step results in 723 distinct SDK families, each with multiple versions, for a total of 14,178 SDK versions.
We built a static-analysis suite for Android APK and SDK analysis, and deployed an interprocedural, context-, field-, object-sensitive taint-flow tracking algorithm for fingerprinting detection. It works by tainting all fingerprinting-related data with meta information and then propagating taint at the instruction level, so that the transparency of the flow information can be achieved, allowing us to reconstruct the taint flow path to independently verify exfiltration. We claim no novelty for this analysis, though some implementation details may be of independent interest, so we include these in Appendix B.
Table 1.SDK Label Definitions. Each SDK receives one label based on its Maven metadata and website description, not based on its code. See Table 6 in Appendix D for more formal descriptions used in our labeling process.
SDK Label
Description
Advertising
Supports displaying ads, ads bidding, ads targeting, ads mediation, or analytics for the purpose of monetization or conversion (Ex: AppLovin, Teads)
Analytics
Monitors and reports on app health (examples: TOAST Logger, RichAPM Agent), or collects the user’s behavior in app (Ex: Pushwoosh, Acoustic Tealeaf)
Security &
Authentication
Implements user authentication (E: Passbase, Ondato), detects fraud and related security anomalies (Ex: Incognia, SEON), and payment functionality (Ex: Alipay, PayPal).
Tools / Other
Provides navigation (Ex: Tencent Map Nav, Radar), object or person tracking (Ex: BeaconsInSpace, Foursquare Movement), communication with social networks (Ex: Facebook, Chat SDK), or other well-defined functionality (Ex: iZooto App Push, GameUp)
Unclear /
Not Found
Purpose or functionality could not be determined from online metadata
3.4.SDK Labeling
While developers provide an attestation of the categorical use-case of their application as a part of submission to the Google Play store, there is no equivalent process for SDKs. Indeed, Maven repositories usually provide only the name of the SDK, a short explanation, and a link back to the originator of the code. This metadata is often incomplete, further obfuscating the use of the SDK. Other datasets are more complete, including the Google Play SDK Index (27), but provide only a small database of SDKs.
Here we borrow techniques from the HCI community and treat the problem as a manual labeling task. Label definitions were collaboratively developed by a team of five expert coders based on the metadata of the SDK, including the SDK’s description in Maven and the content of its developer’s website, from a sample of 100 random SDKs in our dataset. For brevity, an informal description of these categories is in Table 1, and a full explanation (including sub-categories) is in Table 6. Upon reaching saturation, we split the remaining SDKs between reviewers such that each SDK was independently examined twice.
For efficiency, we limited our labeling effort to the 723 SDK families that have been detected in our application dataset, under the assumption that all versions of the same SDK have an equivalent use case and should share the same label. The resultant definitions were robust, with reviewers usually agreeing on SDK labels; the Krippendorff’s alpha inter-rater reliability score of the independent labeling step was 0.804. Finally, all label disagreements were resolved and re-labeled in a meeting of the full group, meaning that any labeling disagreement was addressed by comparing labels from all five coders.
3.5.App–SDK Matching
Inspired by the large body of work in SDK identification for Android apps (65, 67, 66, 6, 41, 38, 58, 38, 30, 64, 61), we created an SDK-identification pipeline using a fine-grained code similarity metric that can be aggregated across code units (e.g., classes, modules) and packaging units (e.g., SDKs, SDK versions). The similarity metric relies on identifiers for system APIs (e.g., operating system calls, standard library calls), opcode frequencies, framework APIs, and string constants.
In our design, we choose parameters for this similarity metric, including the percentage of APK code similar to a known SDK sufficient to declare the SDK as present in the APK, to ensure that our results limited false positives at the risk of some false negatives. In other words, we may miss the presence of an SDK in an APK, and thus the statistical analysis in the rest of the paper provides lower bounds for the prevalence of fingerprinting SDKs. We include a detailed description of our approach in Appendix C.
4.Results
We analyzed the Seed Set, Extended Set, and their prevalence in our application dataset to answer our three research questions (§1).
4.1.RQ1: Self-Identified Fingerprinters
Figure 2.Map of signals collected by known fingerprinting SDKs in the Seed Set, with one dot per API exfiltrated. The top plot shows the percentage of Seed Set SDKs that exfiltrate that API; 80% of APIs are exfiltrated by fewer than 50% of SDKs, and only 2% of APIs are exfiltrated by 75% of SDKs.
Table 2.Seed Set SDKs. List of SDKs that openly admit to fingerprinting, through advertising copy, developer documentation, or other copy. The number of raw number of signals they collect is on the right. Unique Signals is not a total, but the set union of all signals without repetition.
Name Signals
Seon 43
Forter 69
Kaspersky AntiVirus SDK 20
Accertify (InAuth) 213
Castle 31
Microsoft Dynamics 365 128
IP Quality Score 58
Fingerprint.js 30
Shield 148
ThreatMetrix (Lexus Nexus) 94
Ravelin 30
TransUnion TruValidate 55
Socure 43
Incognia 81
Unique Signals 504
What types of behaviors do self-identifying fingerprinting SDKs exhibit?
We find 14 different SDKs that admit to fingerprinting in the wild, listed in Table 2. Our manual analysis found that fingerprinting libraries exfiltrate a minimum of 20 unique signals, and an average of 75.5. Notably, the techniques used by these SDKs were straightforward, with no SDK attempting to collect more than what was available from framework-level API calls. This is a departure from what has been previously measured in the web context (e.g. audio context fingerprinting (19)), as well as the more advanced, hardware-focused techniques discussed by the academic community.
Figure 3.Cosine similarity between Seed Set SDKs, each represented as a one-hot encoding vector of their specific APIs used.
We examine the distribution of signals collected by SDKs in the Seed Set, and present an overview in Figure 2. Though the total number of signals collected is 1043, there are only 504 unique signals, with less than half of all signals collected by at least two SDKs. The set of fingerprinting signals collected are relatively sparse, as the individual signals that SDKs choose to select are somewhat dissimilar—Figure 3 displays the cosine similarity between different fingerprinting SDKs, with only two SDKs scoring above
0.5
(Transunion and Ravelin). Of the
504
unique signals collected, we find that only
21
individual APIs were collected by more than half of the SDKs in our Seed Set, by a maximum of 11 SDKs. We conclude that cosine similarity of individual signals (as used in prior work (22)) is unlikely to be an effective detection mechanism.
Some SDKs stand out with unique API usage patterns. For instance, Forter, Accertify, Microsoft Dynamics, and Shield appear to use a much broader range of APIs compared to others. This could suggest that these SDKs are more complex or serve a wider range of functionalities. Conversely, Kaspersky AntiVirus SDK appears to use a very limited set of APIs, possibly reflecting its more focused security purpose. Under the assumption that all of these SDKs perform fingerprinting equally well, such breadth of API usage patterns implies that some APIs provide more valuable (i.e., more fingerprintable) signals than others.
4.2.RQ2: Purposes of Likely Fingerprinters
What are the stated purposes of SDKs with likely fingerprinting behavior?
The automated exfiltration detection (§3.3) yields 723 SDKs that exhibit behavior similar to the known fingerprinting behavior of the SDKs in the Seed Set. Each SDK may have multiple versions and our SDK dataset identifies 14,178 versions for these 723 SDKs, for an average of 19.60 versions per SDK. With the Extended Set labeled as described in Section 3.3 we consider the prevalence of various purposes across SDKs that exhibit fingerprinting-like behavior and plot the resulting distribution in Figure 4.
Analytics
Security and Authentication
Tools / Other
Unclear / Not found
Ads
77
85
167
173
221
Figure 4.Prevalence of purposes across the Extended Set of 723 likely fingerprinting SDKs.
Several observations are readily available from the plot. First, there is a clear separation between the “Ads” category, the “Tools / Other” category, and the rest of the categories. The “Tools / Other” category being large is expected, as it encompasses a wide variety of functionalities, from cloud storage, to image and video processing, and to consent management. The high number of SDKs with “Ads” as purpose is potentially representative of the complex structure of that particular industry, where ad networks provide their own SDKs and ad platforms act as aggregators and mediators between apps and ad networks, with corresponding mobile SDKs that reflect these relationships. Oftentimes, the ad-mediation systems consists of 5–15 individual SDKs, one for each ad network to which the mediator can connect, easily boosting the number of total SDKs in this vertical. We opted to count these connector SDKs separately — from a security and privacy perspective they are distinct artifacts.
A second observation is there is a large contingent of SDKs whose purpose is not clear from online, public information. The “Unclear / Unfound” category is the second highest source of likely fingerprinting behavior. We know these SDKs are in use in a variety of apps (as we see later in §4.3), so the question is not only what purpose these SDKs serve, but also how to make this purpose information available to security and privacy enforcement mechanisms. One option would be to reverse engineer the code of each SDK and infer from this code its purpose, though this is unlikely to be a scalable long-term solution (and out of scope for this paper).
A further consideration for the use of privacy-preserving alternatives is the long tail of functionalities present in the “Tools / Other” category, which represents 23% of the total number of likely fingerprinting SDKs. While “Ads”, “Security and Authentication”, and “Analytics” are reasonably well understood and studied in terms of privacy, the “Tools / Other” SDKs cover a broad range of algorithms and data types that may not have readily available privacy-preserving alternatives.
SDK Purpose & Behavior.
The challenge of identifying the purpose of likely fingerprinting SDKs of a particular type could be alleviated by focusing on their use of fingerprinting signals, if sufficiently discriminative. For example, if Ads SDKs have distinct fingerprinting behaviors (in terms of signals and API data they exfiltrate) compared to those of Security and Authentication SDKs, one could identify and control fingerprinting appropriately without relying on the SDK’s purpose declaration or on its non-fingerprinting functionality, both of which could be adversarially manipulated. We evaluated this hypothesis by expressing the fingerprinting behavior of each SDK as points in a high-dimensional space defined by a one-hot encoding of the APIs exfiltrated. Each API of interest is an independent dimension in this space and an SDK is placed at position
0
along this dimension if it does not exfiltrate the corresponding API, or
1
if it does. This results in
504
-dimensional space (one for each API observed in the Extended Set of SDKs) in which we locate the
723
SDKs.
In Figure 5, we provide a visual representation of the similarity between different SDK types in the Extended Set using a t-Distributed Stochastic Neighbor Embedding (t-SNE) (56) plot. t-SNE allows us to cluster the SDKs based on their “natural” similarity in fingerprinting behavior by mapping the high-dimensional space (our 504 signals) to a faithful representation in a lower-dimensional space through non-linear transformations while preserving local and global relationships between the data points. Based on the recommendations from Wattenberg, Viégas, and Johnson (59), we set t-SNE perplexity to
25
, learning rate of
10
, and iterations to 5,000, resulting in a final KL divergence of
0.614789
. The resulting t-SNE output is shown in Figure 5, with one dot per SDK, relatively positioned as determined by t-SNE and color coded based on our five SDK purpose labels.
Figure 5.Map of likely fingerprinting behaviors of SDKs in the Extended Set, computed using t-SNE over embeddings constructed by one-hot encoding the exfiltrated APIs. Proximity of SDKs indicates that they exfiltrate data from similar sets of APIs.
The t-SNE plot illustrates the diversity of signal/API usage in fingerprinting behavior, as a large number of small clusters formed. At a minimum this leads us to believe that a corresponding larger number of small, focused permission-based policies may be able to address the fingerprinting problem, with the associate risk of enforcement performance (due to the cost of maintaining and evaluating this many policies) and low usability (due to placing the user in the position of making decision based on seemingly similar but privacy-distinct permissions).
The feasibility of automating the anti-fingerprinting/anti-tracking policies put forth by the industry (advertising: no tracking allowed, anti-fraud: tracking allowed) can be reduced to whether “Ads” SDKs (marked as
in Figure 5) are easily separable from “Security and Authentication” SDKs (
in Figure 5). The right third of t-SNE plot contains most of the “Security and Authentication” SDKs (
), while the “Advertising” SDKs (
) are on the left. Yet there are many “Security and Authentication” SDKs on the left side of the plot, not to mention “Analytics” (
𝑣
) and “Tools / Other” SDKs (
), that appear to have similar fingerprinting-like behavior to the “Ads” SDKs. Thus any automatic enforcement that needs to distinguish between “Ads” SDKs and “Security and Authentication” SDKs will need to rely on non-trivial classifiers that are more expressive than permissions.
Finally we observe that the “Unclear / Unfound” SDKs (
𝑐
), which are declared in APKs and present in Maven repositories but lack any descriptive information, have fingerprinting-like behaviors similar to all other SDK categories. This supports the need for robust categorization and labeling mechanisms for SDKs, and the conclusion that behavioral analysis may be insufficient without additional out-of-band (non-code) information.
Use of Sensitive Signals.
We manually identified 24 APIs which could be used to retrieve location data exactly or approximately and then checked how many of the likely fingerprinting behaviors in SDKs from the Extended Set rely in these APIs. We performed a similar analysis for app-usage signals (based on the three APIs we identified to provide information about the apps the user has installed on the device, the apps that are in use, or the usage statistics for installed apps) and for the account-list signals (based on two APIs to retrieve lists of personal accounts registered on the device). We find that of likely fingerprinting SDKs
72
%
collect coarse-grained location signals,
71.6
%
collect fine-grained location, and
86.29
%
collect at least one or the other. Only only
6.15
%
record account-list signals, and
38.46
%
collect app usage information.
4.3.RQ3: Market Reach
What kinds of apps use SDKs with fingerprinting behavior, and how prevalent are these SDKs in real-world apps?
To answer this question, we consider the SDK categories described in the RQ2 results, and the app categories assigned by the Google Play Store (23).
We are interested in understanding the presence of fingerprinting SDKs in the mobile-app marketplace. For this, we measured how many apps include fingerprinting SDKs in each app category, which categories of fingerprinting SDK are in most use, and which fingerprinting SDKs co-occur most often in apps. The first measurement seeks to determine whether there are app categories with particularly high prevalence of fingerprinting and thus that should be prioritized for any fingerprinting-reduction intervention. The second and third measurements inform any technical efforts to replace fingerprinting-based solutions with privacy-preserving alternatives.
Figure 6shows the prevalence of apps that include fingerprinting SDKs in each app category. 6(a) illustrates that the raw number of apps using fingerprinting SDKs averages at 3.2% across app categories, ranging between 0.8% (for “Events” apps) and 10% (for “Video Players” apps). 6(b) takes into account the number of installs each such app had in the 30-day period, and shows that presence of fingerprinting functionality is heavily skewed towards popular apps.
On average 39.4% of apps in a category contain at least one fingerprinting SDK, and while the minimum prevalence is 5.5% of apps (for the “Libraries and Demo” category), a number of app categories have
50% prevalence: “Maps and Navigation”, “Beauty”, “Shopping”, “Sports”, “House and Home”, “Social”, “Food and Drink”, “Dating”, “Game”, and “Comics”. From a user’s point of view, this indicates that randomly installing a popular app has a 39.4% chance of being fingerprinted and, if they select a dating, game, or comics app, they will be fingerprinted with a 80+% probability.
The “Comics,” “Game,” and “Dating” categories stand out with the highest number of apps incorporating likely fingerprinting SDKs, respectively in this order. This could be attributed to several factors, such as the prevalence of free-to-use functionality (e.g., free-to-play games) that rely on targeted advertising or in-app purchases, or the need for (frictionless) user identification in online settings.
0
50
100
150
200
250
300
350
Education
Game
Business
Tools
Food and Drink
Productivity
Health and Fitness
Entertainment
Shopping
Lifestyle
Music and Audio
Finance
Books and Reference
Personalization
Travel and Local
Communication
Sports
Social
Medical
Maps and Navigation
Auto and Vehicles
News and Magazines
Photography
House and Home
Beauty
Art and Design
Events
Video Players
Dating
Weather
Parenting
Libraries and Demo
Comics
Number of apps (in thousands)
((a))Count of apps (blue) and apps with fingerprinting-like SDKs (red) by app category.
0
20
40
Number of installs (in billions)
((b))Installs of apps (blue) and apps with fingerprinting SDKs (red), by app category.
7.50% 87.76%
1.58% 5.58%
4.98% 38.26%
5.71% 15.78%
4.09% 82.57%
10.45% 23.91%
0.81% 42.98%
2.74% 33.00%
1.11% 52.45%
2.16% 56.51%
5.69% 22.27%
3.85% 45.40%
2.37% 6.14%
3.56% 51.07%
2.33% 31.66%
2.75% 59.25%
2.90% 56.08%
2.17% 10.63%
2.57% 18.28%
2.02% 31.08%
2.65% 31.68%
4.30% 37.48%
2.11% 32.80%
2.22% 33.74%
2.71% 54.44%
3.46% 42.21%
1.77% 31.27%
1.72% 26.71%
1.10% 63.80%
2.98% 13.12%
1.05% 29.18%
7.21% 85.15%
1.21% 48.77%
((c))Percent of apps containing fingerprinting-like SDKs by total count (left) and install volume (right).
Figure 6.The number of apps that come with fingerprinting SDKs is rather small, on average at less than 5% of the total number of apps in a category (6(a)), yet these apps are some of the most installed (6(b)), giving likely fingerprinting SDKs an outsized presence in the market. In 23 app categories (highlighted in (6(c))), apps with fingerprinting SDKs are 10x more popular than other apps.
Fingerprinting Prevalence across App Categories
To further understand the prevalence of likely-fingerprinting SDKs in the mobile-app ecosystem, we use the purpose labels we developed in §4.2 to map app categories to SDK categories. This results in the heatmap shown in 7(a), in which a deeper shade of red indicates that the SDK category of that row dominates the app category of that column. For example, any likely-fingerprinting SDKs used by “Art and Design” apps come primarily from the “Ads” SDK category, while likely-fingerprinting SDKs in “Finance” apps are foremost from the “Analytics” SDK category.
((a))Heatmap of the prevalence of likely fingerprinting SDKs in app categories (X-axis) and SDK categories (Y-axis).
((b))Heatmap of the co-occurrence of likely fingerprinting SDKs across pairs of app categories. Each cell describes the percentage of apps in the corresponding app categories (on the X-axis and Y-axis) that share one or more likely fingerprinting SDKs.
Figure 7.Prevalence of likely fingerprinting SDKs within and across app categories. In (7(a)), the color gradient is computed per app category, allowing the comparison of SDK prevalence between app categories. In (7(b)), the color gradient is computed per pair of app categories given by the X-axis and Y-axis coordinates. For raw data, see Appendix A.
The “Ads” SDKs dominate across almost all app categories as a source of likely fingerprinting behavior, with “Unclear / Unfound” SDKs as the second most common. We note that in many cases the absolute number of “Unclear / Unfound” SDKs is close to that of “Ads” SDKs (e.g., in the “Business” app category, the labels for 16,097 SDKs are unclear, and 17,780 SDKs have the “Ads” label) and thus any shift from “Unclear / Unfound” to “Ads” will only further cement the dominance of “Ads” SDKs a source of likely fingerprinting behavior.
A second observation from this heatmap is that there are several app categories (“Finance”, “Food and Drink”, “Shopping”) where “Analytics” likely-fingerprinting SDKs are more prevalent that other categories of likely-fingerprinting SDKs. We hypothesize that in these app categories, the fingerprinting is less used to track user identities (which are known from the user-account information) and more used to understand user preferences with respect to the items for sale in the app.
Possible Tracking across App Categories via Fingerprint Sharing
A significant privacy risk brought on by fingerprinting is the potential for a third party to track user activity across applications. This can happen when an SDK included in multiple applications fingerprints a user, and thus allows for user activity to be attributed to the same user for both. A service might then learn, say, that a user that engages in particular style of dating app, and also uses a specific medical or finance application. Such cross-app tracking may take place on device or on the server, in both cases powered by the fingerprinting data obtained from the shared SDK.
To estimate a lower bound on the risk of cross-app tracking, we analyze the prevalence of likely fingerprinting SDKs present in distinct categories by computing the probability that two apps randomly selected from each app category share at least one likely fingerprinting SDK. The results are shown in 7(b) as a heatmap (only the lower diagonal presented, as the heatmap is symmetric). A darker shade of blue in the figure indicates a higher prevalence of shared fingerprinting SDKs, as shown, for example, by the (“Game”, “Entertainment”) entry compared with the (“Travel and Local”, “Comics”) entry. We compute these probabilities for the top-1000 apps by total audience size (as defined in §3.1) for each app category, regardless of whether those apps include a likely fingerprinting SDK or not. As a result, the prevalence (and the associated heatmap shown in 7(b)) reflect both the popularity of apps and the distribution of likely fingerprinting SDKs in such popular apps.
Analysis of this heatmap clearly indicates that a few app categories have likely fingerprinting SDKs in common with many other app categories. For example, “Game” apps share likely fingerprinting SDKs with “Art and Design” apps, as well as with apps in the “Beauty”, “Books and Reference”, “Comics”, “Communication”, “Dating”, “Entertainment”, “Health and Fitness”, “Libraries and Demo”, “Lifestyle”, “Maps and Navigation”, “Music and Audio”, “News and Magazines”, “Personalization”, “Photography”, “Productivity”, “Social”, “Tools”, “Video Players”, and “Weather” categories. Similarly, “Personalization” apps share SDKs with 7 out of 33 app categories. From a user’s point of view, this implies that they are at higher risk of cross-app tracking if they install apps from both the “Game” and “Comics” categories.
Alternatively, some categories of apps rarely share likely fingerprinting SDKs with other app categories. We highlight the “Finance” and “Medical” app categories in particular, as such apps often process highly sensitive data. Without further study, one cannot tell why fingerprinting is not more common here, and we note that such apps often require the user to authenticate to access their bank or investment account or their medical record and as such may not need to fingerprint the user through indirect signals.
5.Discussion & Limitations
Our results indicate that the fingerprinting ecosystem is more complex than previously estimated, both in terms of fingerprinting behavior and fingerprinting purpose. We interpret the results below, discuss limitations of our methodology, and present implications for app security mechanisms.
Challenges for Sector-Specific Solutions.
A core observation of this paper is that the current mobile ecosystem has evolved to organically deploy fingerprinting-like behavior in a wide variety of apps through several types of SDKs. Our analysis reveals that while advertising SDKs contribute to fingerprinting, they are not the sole culprits: A significant portion of fingerprinting-like behavior originates from SDKs employed for analytics and anti-fraud purposes, and a large contingent (
23.9
%
) did not have sufficient public information about their purpose or functionality to discern a category. These SDKs are often integrated for reasons directly outside of monetization — understanding user behavior to improve applications or preventing bots and fraud — despite ultimately collecting enough device information to create a fingerprint.
This finding challenges the prevailing notion that fingerprinting is primarily driven by app developer’s need to monetize via advertising, and highlights the opportunity for a broader perspective on privacy-preserving alternatives. Research on how developers can be incentivized to adopt privacy preserving analytics, for example, could prove useful. It is also likely that sandboxing efforts (such as Android’s Privacy Sandbox (1)) could provide additional benefit for both detection and enforcement against SDKs that over-collect — though systems research into lightweight sandboxing to support this setting is necessary, as scaling current process-based techniques to non-ads SDKs requires untenable overhead for constrained mobile environments.
Challenges for API-Specific or Other Behavioral Defenses.
A surprising result of our exploration of the Seed Set (self-identified fingerprinting SDKs) was that the space of APIs used is sparse; the SDKs collected information from dissimilar sets of APIs. This held true regardless of the use-case of the SDK or of its prevalence in our dataset of applications. A potential explanation might be API Proxying (33), which may further complicate any anti-fingerprinting enforcement, though determining the joint entropy or shared entropy of particular API’s is beyond the scope of this work.
In any case, it would appear that targeting specific APIs (akin to Apple’s required reasons (4)) represents a brittle defense against fingerprinting. Developers have a number of signals at their disposal, and could easily move to other sources of entropy. Further, though we (surprisingly) found no evidence of non-API hardware-based fingerprinting in our Seed Set, one might expect developers to shift more advanced methods if comprehensive enforcement at the API level were introduced.
Potential for Sector-Specific Analysis & Targeting.
It is worth noting that certain sensitive application verticals appeared to have an improved privacy stance. Normalized by install volume, only 30% of applications in the medical category used a fingerprinting SDK, and (assuming a normal distribution holds between sample sets) only 19.5% of those did so using an ads SDK — the bulk of identifiable fingerprinting behavior appears to come from analytics. Medical applications also appeared to have a lower potential for cross-application tracking. This heartening result, which is largely repeated in the Finance category, highlights the need for future work focusing on solutions for specific sensitive market verticals.
Need for Multi-Platform Analysis.
It is likely that our results extend to the iOS ecosystem — indeed, all SDKs in our Seed Set appear to have versions readily available for iOS — a finding consistent with prior work on cross-platform tracking (35). However, it is difficult to perform such analysis on iOS, as Apple’s application and operating-system wide DRM restricts third parties’ ability to scalably perform static and dynamic analysis on applications in their App Store. Future work studying the iOS ecosystem would provide invaluable insight into the effectiveness of design choices between the two operating systems.
5.1.Limitations
Any empirical study, including the present paper, is a limited view into real-world conditions and trends and thus it is important to evaluate the factors that threaten its validity. Following the “Campbell Tradition” (11), we consider four types of validity—internal, statistical, construct, and external—and their impact on this study.
Internal validity refers to whether the measured effect (fingerprintable APIs) truly corresponds to the outcome of interest (fingerprinting behavior). A risk is that the use of APIs to retrieve high-entropy data may not be caused by intentional fingerprinting behavior, but instead the result of necessary app functionality. We sidestep this by focusing on documenting the purposes of collection of fingerprintable data. A secondary limitation exists in the selection bias implicit in our Seed Set, which consists of SDKs that self-identify as fingerprinting. It may be that SDKs that fingerprint for hidden reasons use alternative techniques, which would not be caught in our later analyses. We assume that a Seed Set SDK’s self-reporting is honest and make no further inferences about the SDK’s intent.
Statistical validity refers to the risks of underpowered experiments, i.e., without sufficient statistical support. Our large sample size of 228,598 SDKs and 3,025,417 apps mitigates this risk.
Construct validity refers to the choice of metrics to measure the presence of fingerprintable APIs and behaviors. We focused on the number of APIs as an efficient metric of fingerprinting behavior, though we note that not all APIs are equally useful for fingerprinting. For now we make the simplifying assumption that in-the-wild techniques are largely equivalent, and that there is no relationship between signals collected. Using more complex metrics such as collision entropy (10) requires experiments across large sets of devices and users, which we leave for future work.
External validity refers to the generalizability of our results to real-world. Our choice of actual SDKs from popular Maven repositories and mobile apps from the Google Play store ensure minimize this risk. However, we do not attempt to catalog all fingerprinting mobile ecosystem, limiting ourselves to Java-language SDKs that are part of the Android/AOSP framework (excluding non-platform APIs or those from OEMs). It is possible that the Seed Set of fingerprinting SDKs, hand selected through web search, is not representative of all fingerprinting behaviors in the wild, and further study to ensure a comprehensive view is needed.
6.Conclusion
In this paper, we presented the largest-scale analysis of SDK behavior ever conducted, examining over 228,000 SDKs and 178,000 Android applications to understand the prevalence and purpose of fingerprinting-like behavior. Our findings reveal that a significant number of SDKs, beyond those explicitly designed for advertising, collect enough information to potentially track users. This includes SDKs used for analytics and anti-fraud, highlighting the need for privacy-preserving alternatives in these areas. Surprisingly, a large portion of SDKs exhibiting fingerprinting-like behavior lacked clear identification, emphasizing the need for greater transparency in the SDK ecosystem. Moreover, we observed that these SDKs with fingerprinting-like behavior are disproportionately popular and often integrated across diverse application categories. These results underscore the importance of ongoing efforts by Apple and Google to enhance user privacy and emphasize the need for continued research to ensure that such industry efforts are well directed.
References
[1]
- * *.
The Privacy Sandbox: Technology for a More Private Web.
Available online at https://privacysandbox.com/intl/en_us. Last visited: 2024-09-30.
[2]
Irene Amerini, Rudy Becarelli, Roberto Caldelli, Alessio Melani, and Moreno Niccolai.
Smartphone fingerprinting combining features of on-board sensors.
IEEE Transactions on Information Forensics and Security, 12(10), 2017.
[3]
Apple.
Describing data use in privacy manifests.
[4]
Apple.
Describing use of required reason api.
Available online at https://developer.apple.com/documentation/bundleresources/privacy_manifest_files/describing_use_of_required_reason_api. Last visited: 2024-09-30.
[5]
Apple.
User Privacy and Data Use - App Store.
Available online at https://developer.apple.com/app-store/user-privacy-and-data-use/. Last visited: 2024-09-30.
[6]
Michael Backes, Sven Bugiel, and Erik Derr.
Reliable third-party library detection in android and its security applications.
In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, CCS ’16, page 356–367, New York, NY, USA, 2016. Association for Computing Machinery.
[7]
Pouneh Nikkhah Bahrami, Umar Iqbal, and Zubair Shafiq.
Fp-radar: Longitudinal measurement and early detection of browser fingerprinting, 2021.
[8]
Flavio Bertini, Rajesh Sharma, Andrea Iannì, and Danilo Montesi.
Smartphone verification and user profiles linking across social networks by camera fingerprinting.
In Joshua I. James and Frank Breitinger, editors, Digital Forensics and Cyber Crime, pages 176–186, Cham, 2015. Springer International Publishing.
[9]
Hristo Bojinov, Yan Michalevsky, Gabi Nakibly, and Dan Boneh.
Mobile device identification via sensor fingerprinting, 2014.
[10]
Gecia Bravo-Hermsdorff, Róbert Busa-Fekete, Mohammad Ghavamzadeh, Andres Muñoz Medina, and Umar Syed.
Private and communication-efficient algorithms for entropy estimation, 2023.
[11]
Donald T Campbell and Julian C Stanley.
Experimental and quasi-experimental designs for research.
Ravenio books, 2015.
[12]
Yinzhi Cao, Song Li, and Erik Wijmans.
(cross-)browser fingerprinting via os and hardware level features.
In Network and Distributed System Security Symposium, 2017.
[13]
Jing Chen, Yingying Fang, Kun He, and Ruiying Du.
Charge-depleting of the batteries makes smartphones recognizable.
In 2017 IEEE 23rd International Conference on Parallel and Distributed Systems (ICPADS), pages 33–40, 2017.
[14]
Anupam Das, Nikita Borisov, and Matthew Caesar.
Do you hear what I hear? fingerprinting smart devices through embedded acoustic components.
In Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security, CCS ’14, New York, NY, USA, 2014. Association for Computing Machinery.
[15]
Anupam Das, Nikita Borisov, and Matthew Caesar.
Exploring ways to mitigate sensor-based smartphone fingerprinting, 2015.
[16]
Sanorita Dey, Nirupam Roy, Wenyuan Xu, Romit Roy Choudhury, and Srihari Nelakuditi.
Accelprint: Imperfections of accelerometers make smartphones trackable.
In NDSS. The Internet Society, 2014.
[17]
Antonios Dimitriadis, George Drosatos, and Pavlos S. Efraimidis.
How much does a zero-permission android app know about us?
In Proceedings of the Third Central European Cybersecurity Conference, CECC 2019, New York, NY, USA, 2019. Association for Computing Machinery.
[18]
Peter Eckersley.
How unique is your web browser?
In Mikhail J. Atallah and Nicholas J. Hopper, editors, Privacy Enhancing Technologies, pages 1–18, Berlin, Heidelberg, 2010. Springer Berlin Heidelberg.
[19]
Steven Englehardt and Arvind Narayanan.
Online tracking: A 1-million-site measurement and analysis.
In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pages 1388–1401. ACM, 2016.
[20]
Facebook Research.
Faiss.
Online at https://github.com/facebookresearch/faiss. Last accessed June 5, 2023.
[21]
Amin FaizKhademi, Mohammad Zulkernine, and Komminist Weldemariam.
Fpguard: Detection and prevention of browser fingerprinting.
In Pierangela Samarati, editor, Data and Applications Security and Privacy XXIX, 2015.
[22]
Christof Ferreira Torres and Hugo Jonker.
Investigating fingerprinters and fingerprinting-alike behaviour of android applications.
In European Symposium on Research in Computer Security, pages 60–80. Springer, 2018.
[23]
Google.
Play Console Help: Example categories.
Online at https://support.google.com/googleplay/android-developer/answer/9859673.
[24]
Google.
Play console help: View app statistics.
Online at https://support.google.com/googleplay/android-developer/answer/139628?hl=en&co=GENIE.Platform%3DAndroid.
[25]
Google.
Provide information for google play’s data safety section.
[26]
Google.
SDK Runtime overview.
[27]
Google.
Google play sdk index.
Online at https://play.google.com/sdks, 2024.
[28]
Google Research.
ScaNN.
Online at https://github.com/google-research/google-research/tree/master/scann. Last accessed June 5, 2023.
[29]
Catherine Han, Irwin Reyes, Álvaro Feal, Joel Reardon, Primal Wijesekera, Narseo Vallina-Rodriguez, Amit Elazar, Kenneth A. Bamberger, and Serge Egelman.
The price is (not) right: Comparing privacy in free and paid apps.
Proceedings on Privacy Enhancing Technologies, 2020(3), 2020.
[30]
Hongmu Han, Ruixuan Li, and Junwei Tang.
Identify and inspect libraries in android applications.
Wirel. Pers. Commun., 103(1):491–503, nov 2018.
[31]
Kris Heid, Vincent Andrae, and Jens Heider.
Towards detecting device fingerprinting on ios with api function hooking.
EICC ’23, New York, NY, USA, 2023. Association for Computing Machinery.
[32]
Thomas Hupperich, Davide Maiorca, Marc Kührer, Thorsten Holz, and Giorgio Giacinto.
On the robustness of mobile device fingerprinting: Can mobile users escape modern web-tracking mechanisms?
In Proceedings of the 31st Annual Computer Security Applications Conference, ACSAC ’15, 2015.
[33]
Somesh Jha, Mihai Christodorescu, and Anh Pham.
Formal analysis of the api proxy problem, 2023.
[34]
T. Kohno, A. Broido, and K.C. Claffy.
Remote physical device fingerprinting.
IEEE Transactions on Dependable and Secure Computing, 2(2):93–108, 2005.
[35]
Konrad Kollnig, Anastasia Shuba, Reuben Binns, Max Van Kleek, and Nigel Shadbolt.
Are iPhones Really Better for Privacy? A Comparative Study of iOS and Android Apps.
2022(2):6–24.
[36]
Aleksandra Korolova and Vinod Sharma.
Cross-app tracking via nearby bluetooth low energy devices.
In Proceedings of the Eighth ACM Conference on Data and Application Security and Privacy, CODASPY ’18, page 43–52, New York, NY, USA, 2018. Association for Computing Machinery.
[37]
Andreas Kurtz, Hugo Gascon, Tobias Becker, Konrad Rieck, and Felix C Freiling.
Fingerprinting mobile devices using personalized configurations.
Proc. Priv. Enhancing Technol., 2016(1):4–19, 2016.
[38]
Menghao Li, Wei Wang, Pei Wang, Shuai Wang, Dinghao Wu, Jian Liu, Rui Xue, and Wei Huo.
Libd: Scalable and precise third-party library detection in android markets.
In 2017 IEEE/ACM 39th International Conference on Software Engineering (ICSE), pages 335–346, 2017.
[39]
Tianyi Li, Xiaofeng Zheng, Kaiwen Shen, and Xinhui Han.
Fpflow: Detect and prevent browser fingerprinting with dynamic taint analysis.
In Wei Lu, Yuqing Zhang, Weiping Wen, Hanbing Yan, and Chao Li, editors, Cyber Security, pages 51–67, Singapore, 2022. Springer Nature Singapore.
[40]
Xiang-Yang Li, Huiqi Liu, Lan Zhang, Zhenan Wu, Yaochen Xie, Ge Chen, Chunxiao Wan, and Zhongwei Liang.
Finding the stars in the fireworks: Deep understanding of motion sensor fingerprint.
IEEE/ACM Transactions on Networking, 27(5):1945–1958, 2019.
[41]
Ziang Ma, Haoyu Wang, Yao Guo, and Xiangqun Chen.
Libradar: Fast and accurate detection of third-party libraries in android apps.
In Proceedings of the 38th International Conference on Software Engineering Companion, ICSE ’16, page 653–656, New York, NY, USA, 2016. Association for Computing Machinery.
[42]
René Mayrhofer, Jeffrey Vander Stoep, Chad Brubaker, Dianne Hackborn, Bram Bonné, Güliz Seray Tuncay, Roger Piqueras Jover, and Michael A. Specter.
The android platform security model (2023), 2021.
[43]
Microsoft.
SPTAG.
Online at https://github.com/microsoft/SPTAG. Last accessed June 5, 2023.
[44]
Gabi Nakibly, Gilad Shelef, and Shiran Yudilevich.
Hardware fingerprinting using html5, 2015.
[45]
Nick Nikiforakis, Wouter Joosen, and Benjamin Livshits.
Privaricator: Deceiving fingerprinters with little white lies.
In Proceedings of the 24th International Conference on World Wide Web, WWW ’15, 2015.
[46]
Nick Nikiforakis, Alexandros Kapravelos, Wouter Joosen, Christopher Kruegel, Frank Piessens, and Giovanni Vigna.
Cookieless Monster: Exploring the Ecosystem of Web-Based Device Fingerprinting.
In 2013 IEEE Symposium on Security and Privacy, pages 541–555, May 2013.
[47]
Łukasz Olejnik, Gunes Acar, Claude Castelluccia, and Claudia Diaz.
The leaking battery.
In Joaquin Garcia-Alfaro, Guillermo Navarro-Arribas, Alessandro Aldini, Fabio Martinelli, and Neeraj Suri, editors, Data Privacy Management, and Security Assurance, pages 254–263, Cham, 2016. Springer International Publishing.
[48]
Gerald Palfinger and Bernd Prünster.
Androprint: Analysing the fingerprintability of the android api.
In Proceedings of the 15th International Conference on Availability, Reliability and Security, ARES ’20, New York, NY, USA, 2020. Association for Computing Machinery.
[49]
Erwin Quiring, Matthias Kirchner, and Konrad Rieck.
On the security and applicability of fragile camera fingerprints.
In Computer Security – ESORICS 2019: 24th European Symposium on Research in Computer Security, Luxembourg, September 23–27, 2019, Proceedings, Part I, page 450–470, Berlin, Heidelberg, 2019. Springer-Verlag.
[50]
Jingjing Ren, Martina Lindorfer, Daniel J. Dubois, Ashwin Rao, David R. Choffnes, and Narseo Vallina-Rodriguez.
Bug fixes, improvements, … and privacy leaks - a longitudinal study of pii leaks across android app versions.
In Network and Distributed System Security Symposium, 2018.
[51]
Iskander Sanchez-Rola, Igor Santos, and Davide Balzarotti.
Clock around the clock: Time-based device fingerprinting.
In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security, CCS ’18, page 1502–1514, New York, NY, USA, 2018. Association for Computing Machinery.
[52]
Raphael Spreitzer, Felix Kirchengast, Daniel Gruss, and Stefan Mangard.
Procharvester: Fully automated analysis of procfs side-channel leaks on android.
In Proceedings of the 2018 on Asia Conference on Computer and Communications Security, ASIACCS ’18, 2018.
[53]
O. Starov and N. Nikiforakis.
Xhound: Quantifying the fingerprintability of browser extensions.
In 2017 IEEE Symposium on Security and Privacy (SP), pages 941–956, Los Alamitos, CA, USA, may 2017. IEEE Computer Society.
[54]
Junze Tian, Jianyi Zhang, Xiuying Li, Changchun Zhou, Ruilong Wu, Yuchen Wang, and Shengyuan Huang.
Mobile device fingerprint identification using gyroscope resonance.
IEEE Access, 9:160855–160867, 2021.
[55]
Güliz Seray Tuncay, Jingyu Qian, and Carl A. Gunter.
See no evil: Phishing for permissions with false transparency.
In 29th USENIX Security Symposium (USENIX Security 20), pages 415–432. USENIX Association, August 2020.
[56]
Laurens van der Maaten and Geoffrey Hinton.
Visualizing data using t-sne.
Journal of Machine Learning Research, 9(86):2579–2605, 2008.
[57]
Tom Van Goethem, Wout Scheepers, Davy Preuveneers, and Wouter Joosen.
Accelerometer-based device fingerprinting for multi-factor mobile authentication.
In Juan Caballero, Eric Bodden, and Elias Athanasopoulos, editors, Engineering Secure Software and Systems, Cham, 2016.
[58]
Yan Wang, Haowei Wu, Hailong Zhang, and Atanas Rountev.
Orlis: Obfuscation-resilient library detection for android.
In 2018 IEEE/ACM 5th International Conference on Mobile Software Engineering and Systems (MOBILESoft), 2018.
[59]
Martin Wattenberg, Fernanda Viégas, and Ian Johnson.
How to use t-sne effectively.
Distill, 2016.
[60]
Wenjia Wu, Jianan Wu, Yanhao Wang, Zhen Ling, and Ming Yang.
Efficient fingerprinting-based android device identification with zero-permission identifiers.
IEEE Access, 4:8073–8083, 2016.
[61]
Jian Xu and Qianting Yuan.
Libroad: Rapid, online, and accurate detection of tpls on android.
IEEE Transactions on Mobile Computing, 21(1):167–180, 2022.
[62]
Yahoo! JAPAN.
NGT.
Online at https://github.com/yahoojapan/NGT. Last accessed June 5, 2023.
[63]
Yandex.
Hnswlib.
Online at https://github.com/nmslib/hnswlib. Last accessed June 5, 2023.
[64]
Xian Zhan, Lingling Fan, Sen Chen, Feng Wu, Tianming Liu, Xiapu Luo, and Yang Liu.
Atvhunter: Reliable version detection of third-party libraries for vulnerability identification in android applications.
In Proceedings of the 43rd International Conference on Software Engineering, ICSE ’21. IEEE Press, 2021.
[65]
Xian Zhan, Lingling Fan, Tianming Liu, Sen Chen, Li Li, Haoyu Wang, Yifei Xu, Xiapu Luo, and Yang Liu.
Automated third-party library detection for android applications: Are we there yet?
In 2020 35th IEEE/ACM International Conference on Automated Software Engineering (ASE), pages 919–930, 2020.
[66]
Xian Zhan, Tianming Liu, Lingling Fan, Li Li, Sen Chen, Xiapu Luo, and Yang Liu.
Research on third-party libraries in android apps: A taxonomy and systematic literature review.
IEEE Transactions on Software Engineering, 48(10):4181–4213, 2022.
[67]
Xian Zhan, Tianming Liu, Yepang Liu, Yang Liu, Li Li, Haoyu Wang, and Xiapu Luo.
A systematic assessment on android third-party library detection tools.
IEEE Transactions on Software Engineering, 48(11):4249–4273, 2022.
[68]
Jiexin Zhang, Alastair R. Beresford, and Ian Sheret.
Sensorid: Sensor calibration fingerprinting for smartphones.
In 2019 IEEE Symposium on Security and Privacy (SP), pages 638–655, 2019.
[69]
Jiexin Zhang, Alastair R. Beresford, and Ian Sheret.
Factory calibration fingerprinting of sensors.
IEEE Transactions on Information Forensics and Security, 16:1626–1639, 2021.
[70]
Zhe Zhou, Wenrui Diao, Xiangyu Liu, and Kehuan Zhang.
Acoustic fingerprinting revisited: Generate stable device id stealthily with inaudible sound.
In Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security, CCS ’14, page 429–440. Association for Computing Machinery, 2014.
Conference Version
A shorter version of this paper is published at the ACM Conference on Computer and Communications Security 2025 (CCS’25). The present version additionally includes:
• Descriptions of the static analyses performed,
• Description of the SDK-identification algorithm,
• Definition of the codebook used to categorized SDKs, and
• List of the APIs observed in fingerprinting-like behaviors.
Appendix AData Tables for Fingerprinting Prevalence in SDK and App Categories
The following tables provide detailed data on our market measurements.
Table 3presents the prevalence of likely fingerprinting SDKs across various app categories, broken down by SDK type. For instance, in the “Art and Design” app category, 43.3% of apps are likely to contain Ads SDKs that likely engage in fingerprinting. The data shows that “Ads” and “Unclear/Unfound” SDK categories generally have higher prevalence rates across most app categories compared to “Analytics,” “Security and Authentication,” and “Tools/Other” SDKs.
Tables 4 and 5 detail the proportion of apps within each category that contain SDKs also present in apps of other categories. The tables indicate that many apps utilize SDKs that are also prevalent in apps belonging to different categories. For instance, while 0.401 of “Books and Reference” apps share SDKs with “Art and Design” apps, only 0.095 of “Art and Design” apps share SDKs with the “Food and Drink” category, indicating a much lower overlap in SDK usage between these two specific app types.
Table 3.The prevalence of likely fingerprinting SDKs (by SDK category) in app categories.
SDK Category
App Category
Ads
Analytics
Sec. and Authn
Tools / Other
Unclear / Unfound
Art and Design
0.433
0.138
0.132
0.022
0.274
Auto and Vehicles
0.211
0.200
0.183
0.064
0.341
Beauty
0.484
0.146
0.080
0.014
0.277
Books and Reference
0.544
0.093
0.066
0.030
0.267
Business
0.241
0.234
0.138
0.067
0.320
Comics
0.387
0.193
0.099
0.024
0.297
Communication
0.475
0.127
0.096
0.018
0.284
Dating
0.275
0.237
0.108
0.022
0.358
Education
0.516
0.102
0.061
0.036
0.285
Entertainment
0.412
0.128
0.119
0.021
0.320
Events
0.346
0.118
0.154
0.044
0.338
Finance
0.163
0.397
0.111
0.038
0.292
Food and Drink
0.095
0.318
0.225
0.044
0.318
Game
0.339
0.160
0.185
0.010
0.307
Health and Fitness
0.340
0.200
0.117
0.055
0.287
House and Home
0.209
0.252
0.189
0.059
0.291
Libraries and Demo
0.463
0.075
0.100
0.025
0.338
Lifestyle
0.358
0.165
0.129
0.044
0.304
Maps and Navigation
0.325
0.167
0.174
0.051
0.282
Medical
0.195
0.288
0.140
0.052
0.324
Music and Audio
0.452
0.108
0.093
0.063
0.284
News and Magazines
0.386
0.193
0.045
0.052
0.324
Parenting
0.314
0.263
0.082
0.041
0.299
Personalization
0.451
0.094
0.140
0.017
0.298
Photography
0.464
0.140
0.103
0.024
0.268
Productivity
0.463
0.145
0.105
0.028
0.259
Shopping
0.121
0.347
0.177
0.056
0.298
Social
0.317
0.236
0.120
0.051
0.276
Sports
0.375
0.199
0.085
0.038
0.302
Tools
0.454
0.116
0.130
0.029
0.271
Travel and Local
0.179
0.203
0.272
0.059
0.287
Video Players
0.483
0.092
0.121
0.018
0.286
Weather
0.411
0.118
0.148
0.053
0.270
Table 4.The sharing prevalence of likely fingerprinting SDKs across app categories [Part 1]. Part 2 is in Table 5.
App Category
App Category
Art and Design
Auto and Vehicles
Beauty
Books and Reference
Business
Comics
Communication
Dating
Education
Entertainment
Events
Finance
Food and Drink
Game
Health and Fitness
House and Home
Libraries and Demo
Art and Design
0.469
0.202
0.381
0.401
0.210
0.499
0.322
0.365
0.272
0.381
0.272
0.164
0.195
0.794
0.366
0.214
0.304
Auto and Vehicles
0.202
0.182
0.197
0.187
0.174
0.238
0.166
0.253
0.178
0.233
0.193
0.174
0.203
0.378
0.211
0.197
0.143
Beauty
0.381
0.197
0.332
0.327
0.200
0.410
0.269
0.331
0.243
0.330
0.249
0.174
0.203
0.677
0.313
0.211
0.253
Books and Reference
0.401
0.187
0.327
0.347
0.192
0.435
0.278
0.326
0.241
0.335
0.242
0.158
0.186
0.714
0.322
0.198
0.260
Business
0.210
0.174
0.200
0.192
0.171
0.251
0.169
0.254
0.177
0.236
0.187
0.173
0.198
0.403
0.214
0.189
0.145
Comics
0.499
0.238
0.410
0.435
0.251
0.577
0.352
0.432
0.309
0.443
0.296
0.226
0.256
0.847
0.407
0.251
0.322
Communication
0.322
0.166
0.269
0.278
0.169
0.352
0.228
0.278
0.203
0.279
0.206
0.145
0.172
0.598
0.266
0.178
0.210
Dating
0.365
0.253
0.331
0.326
0.254
0.432
0.278
0.419
0.277
0.381
0.274
0.274
0.302
0.669
0.344
0.267
0.237
Education
0.272
0.178
0.243
0.241
0.177
0.309
0.203
0.277
0.199
0.267
0.205
0.167
0.194
0.512
0.247
0.192
0.183
Entertainment
0.381
0.233
0.330
0.335
0.236
0.443
0.279
0.381
0.267
0.381
0.263
0.231
0.261
0.714
0.337
0.249
0.244
Events
0.272
0.193
0.249
0.242
0.187
0.296
0.206
0.274
0.205
0.263
0.227
0.165
0.201
0.483
0.251
0.210
0.192
Finance
0.164
0.174
0.174
0.158
0.173
0.226
0.145
0.274
0.167
0.231
0.165
0.210
0.225
0.338
0.197
0.185
0.108
Food and Drink
0.195
0.203
0.203
0.186
0.198
0.256
0.172
0.302
0.194
0.261
0.201
0.225
0.253
0.381
0.227
0.217
0.134
Game
0.794
0.378
0.677
0.714
0.403
0.847
0.598
0.669
0.512
0.714
0.483
0.338
0.381
0.994
0.659
0.394
0.557
Health and Fitness
0.366
0.211
0.313
0.322
0.214
0.407
0.266
0.344
0.247
0.337
0.251
0.197
0.227
0.659
0.320
0.227
0.243
House and Home
0.214
0.197
0.211
0.198
0.189
0.251
0.178
0.267
0.192
0.249
0.210
0.185
0.217
0.394
0.227
0.217
0.154
Libraries and Demo
0.304
0.143
0.253
0.260
0.145
0.322
0.210
0.237
0.183
0.244
0.192
0.108
0.134
0.557
0.243
0.154
0.212
Lifestyle
0.296
0.203
0.265
0.265
0.203
0.349
0.224
0.318
0.221
0.303
0.227
0.202
0.230
0.568
0.276
0.219
0.196
Maps and Navigation
0.301
0.190
0.261
0.265
0.188
0.338
0.223
0.295
0.212
0.287
0.220
0.174
0.205
0.560
0.268
0.203
0.200
Medical
0.177
0.182
0.183
0.168
0.174
0.217
0.153
0.247
0.172
0.225
0.188
0.178
0.208
0.334
0.200
0.198
0.130
Music and Audio
0.561
0.247
0.457
0.488
0.260
0.605
0.395
0.440
0.335
0.472
0.328
0.201
0.239
0.886
0.445
0.263
0.369
News and Magazines
0.341
0.232
0.309
0.306
0.231
0.381
0.260
0.355
0.257
0.339
0.265
0.222
0.254
0.599
0.318
0.251
0.237
Parenting
0.269
0.203
0.248
0.242
0.200
0.319
0.208
0.308
0.214
0.290
0.222
0.200
0.229
0.516
0.260
0.219
0.181
Personalization
0.610
0.254
0.499
0.530
0.269
0.640
0.430
0.465
0.355
0.494
0.350
0.197
0.237
0.912
0.479
0.268
0.408
Photography
0.511
0.218
0.417
0.442
0.231
0.553
0.356
0.410
0.299
0.423
0.291
0.191
0.220
0.838
0.404
0.228
0.332
Productivity
0.364
0.178
0.299
0.315
0.183
0.398
0.255
0.306
0.224
0.310
0.225
0.155
0.182
0.661
0.298
0.190
0.237
Shopping
0.202
0.190
0.203
0.190
0.190
0.273
0.172
0.318
0.189
0.267
0.182
0.235
0.253
0.416
0.229
0.199
0.128
Social
0.351
0.213
0.305
0.310
0.217
0.411
0.257
0.357
0.245
0.345
0.242
0.217
0.243
0.661
0.312
0.226
0.224
Sports
0.290
0.199
0.264
0.259
0.198
0.328
0.220
0.312
0.219
0.294
0.225
0.197
0.224
0.535
0.272
0.214
0.195
Tools
0.412
0.193
0.334
0.356
0.200
0.457
0.287
0.342
0.248
0.352
0.245
0.170
0.198
0.738
0.333
0.205
0.264
Travel and Local
0.192
0.177
0.190
0.179
0.172
0.237
0.160
0.261
0.174
0.231
0.184
0.185
0.211
0.366
0.208
0.190
0.134
Video Players
0.515
0.222
0.421
0.445
0.234
0.553
0.358
0.398
0.304
0.427
0.299
0.178
0.212
0.841
0.405
0.238
0.337
Weather
0.322
0.169
0.268
0.279
0.172
0.355
0.227
0.276
0.204
0.283
0.209
0.144
0.170
0.603
0.268
0.183
0.210
Table 5.The sharing prevalence of likely fingerprinting SDKs across app categories [Part 2]. Part 1 of this data is in Table 4.
App Category
App Category
Lifestyle
Maps and Navigation
Medical
Music and Audio
News and Magazines
Parenting
Personalization
Photography
Productivity
Shopping
Social
Sports
Tools
Travel and Local
Video Players
Art and Design
0.296
0.301
0.177
0.561
0.341
0.269
0.610
0.511
0.364
0.202
0.351
0.290
0.412
0.192
0.515
Auto and Vehicles
0.203
0.190
0.182
0.247
0.232
0.203
0.254
0.218
0.178
0.190
0.213
0.199
0.193
0.177
0.222
Beauty
0.265
0.261
0.183
0.457
0.309
0.248
0.499
0.417
0.299
0.203
0.305
0.264
0.334
0.190
0.421
Books and Reference
0.265
0.265
0.168
0.488
0.306
0.242
0.530
0.442
0.315
0.190
0.310
0.259
0.356
0.179
0.445
Business
0.203
0.188
0.174
0.260
0.231
0.200
0.269
0.231
0.183
0.190
0.217
0.198
0.200
0.172
0.234
Comics
0.349
0.338
0.217
0.605
0.381
0.319
0.640
0.553
0.398
0.273
0.411
0.328
0.457
0.237
0.553
Communication
0.224
0.223
0.153
0.395
0.260
0.208
0.430
0.356
0.255
0.172
0.257
0.220
0.287
0.160
0.358
Dating
0.318
0.295
0.247
0.440
0.355
0.308
0.465
0.410
0.306
0.318
0.357
0.312
0.342
0.261
0.398
Education
0.221
0.212
0.172
0.335
0.257
0.214
0.355
0.299
0.224
0.189
0.245
0.219
0.248
0.174
0.304
Entertainment
0.303
0.287
0.225
0.472
0.339
0.290
0.494
0.423
0.310
0.267
0.345
0.294
0.352
0.231
0.427
Events
0.227
0.220
0.188
0.328
0.265
0.222
0.350
0.291
0.225
0.182
0.242
0.225
0.245
0.184
0.299
Finance
0.202
0.174
0.178
0.201
0.222
0.200
0.197
0.191
0.155
0.235
0.217
0.197
0.170
0.185
0.178
Food and Drink
0.230
0.205
0.208
0.239
0.254
0.229
0.237
0.220
0.182
0.253
0.243
0.224
0.198
0.211
0.212
Game
0.568
0.560
0.334
0.886
0.599
0.516
0.912
0.838
0.661
0.416
0.661
0.535
0.738
0.366
0.841
Health and Fitness
0.276
0.268
0.200
0.445
0.318
0.260
0.479
0.404
0.298
0.229
0.312
0.272
0.333
0.208
0.405
House and Home
0.219
0.203
0.198
0.263
0.251
0.219
0.268
0.228
0.190
0.199
0.226
0.214
0.205
0.190
0.238
Libraries and Demo
0.196
0.200
0.130
0.369
0.237
0.181
0.408
0.332
0.237
0.128
0.224
0.195
0.264
0.134
0.337
Lifestyle
0.254
0.238
0.199
0.366
0.285
0.244
0.384
0.329
0.247
0.228
0.280
0.246
0.276
0.203
0.330
Maps and Navigation
0.238
0.232
0.182
0.369
0.273
0.227
0.396
0.331
0.247
0.201
0.263
0.233
0.275
0.186
0.333
Medical
0.199
0.182
0.187
0.218
0.225
0.200
0.217
0.191
0.162
0.193
0.204
0.194
0.174
0.179
0.196
Music and Audio
0.366
0.369
0.218
0.669
0.413
0.332
0.711
0.611
0.446
0.247
0.434
0.353
0.505
0.234
0.618
News and Magazines
0.285
0.273
0.225
0.413
0.339
0.276
0.438
0.373
0.287
0.251
0.312
0.288
0.313
0.229
0.378
Parenting
0.244
0.227
0.200
0.332
0.276
0.241
0.344
0.297
0.228
0.225
0.266
0.239
0.252
0.201
0.299
Personalization
0.384
0.396
0.217
0.711
0.438
0.344
0.765
0.659
0.485
0.250
0.457
0.373
0.545
0.239
0.663
Photography
0.329
0.331
0.191
0.611
0.373
0.297
0.659
0.567
0.403
0.239
0.396
0.320
0.458
0.212
0.564
Productivity
0.247
0.247
0.162
0.446
0.287
0.228
0.485
0.403
0.289
0.186
0.287
0.242
0.326
0.172
0.405
Shopping
0.228
0.201
0.193
0.247
0.251
0.225
0.250
0.239
0.186
0.275
0.252
0.224
0.205
0.207
0.217
Social
0.280
0.263
0.204
0.434
0.312
0.266
0.457
0.396
0.287
0.252
0.322
0.271
0.324
0.214
0.393
Sports
0.246
0.233
0.194
0.353
0.288
0.239
0.373
0.320
0.242
0.224
0.271
0.248
0.266
0.199
0.321
Tools
0.276
0.275
0.174
0.505
0.313
0.252
0.545
0.458
0.326
0.205
0.324
0.266
0.372
0.187
0.460
Travel and Local
0.203
0.186
0.179
0.234
0.229
0.201
0.239
0.212
0.172
0.207
0.214
0.199
0.187
0.181
0.209
Video Players
0.330
0.333
0.196
0.618
0.378
0.299
0.663
0.564
0.405
0.217
0.393
0.321
0.460
0.209
0.574
Weather
0.227
0.224
0.158
0.398
0.262
0.212
0.431
0.355
0.256
0.166
0.259
0.220
0.290
0.161
0.362
Appendix BDetails of Our Static Analyses
Dependency Analysis for Android SDKs
It is a common practice for a target software development kit (SDK) to rely on other SDKs. For instance, an advertising SDK may gather user data for fingerprinting and then share the data externally using another SDK (e.g., OkHttp). We refer to the target SDK as the main SDK and the SDKs utilized by the main SDK as its dependency SDKs. It is crucial to accurately infer and integrate the dependencies of each main SDK into the analysis for thorough detection.
To deduce the dependencies for every primary SDK, we analyze the Maven Project Object Model (POM) files of all the SDKs in our repository, extracting the referenced SDKs in each file. From each SDK’s POM file information we construct a dependency graph with directed edges whenever one SDK’s POM file references another’s POM file. In the dependency graph, a distinct SDK version is represented by a trio of factors: a group ID indicating the developer, an artifact ID specifying the SDK name, and a version number, typically written as
𝑋
:
𝑌
:
𝑍
for version
𝑍
of SDK
𝑌
from developer
𝑋
. During the resolution of the dependency graph, version conflict could happen. For instance, an SDK
𝑀
:
𝐴
may reference both SDKs
𝑁
:
𝐵
:
1
and
𝑃
:
𝐶
:
1
. However, SDK
𝑁
:
𝐵
:
1
may require SDK
𝑃
:
𝐶
:
2
. This means that for SDK
𝑃
:
𝐶
, both versions
1
and
2
are listed as dependencies of SDK
𝑀
:
𝐴
. Since importing both versions into the static analysis could lead to one version overriding the other, causing non-deterministic behaviors, we only keep one version for each SDK. Based on Ockham’s razor principle, we prioritize the version with the shortest path from the main SDK in the dependency graph, which is
𝑃
:
𝐶
:
1
in the above example.
During the analysis, we configure the public methods in the main SDK as the analysis’ entry points, meaning that only the execution paths and behaviors triggered by one of these public methods will be reported.
Static Taint Analysis
We employ static taint flow analysis to uncover potential instances of fingerprinting. This process begins by tainting PII data with metadata that allows the system to track its movement throughout the program’s code. To achieve this, the static taint flow analysis encompasses several key phases. First, the system scans the input program to pinpoint likely taint sources (API method call where the program accesses PII) and sinks (API method calls where data exits the device).
Once sources and sinks are established, taint flags are propagated to construct a taint flow graph. This graph utilizes nodes to signify program elements (such as registers or fields), while edges encapsulate the possible transfer of tainted data. The graph expands incrementally until it reaches a fixed point. A detected taint flow path connecting a source and sink flags potential PII exfiltration.
Figure 8.Static Taint Analysis – Identify Sources and Sinks
Figure 9.Static Taint Analysis – Taint Propagation
CoFlow Analysis
CoFlow analysis is a specialized taint flow analysis designed to identify suspicious app behavior patterns based on “crossover flows”. Unlike the traditional taint flow analysis, which focuses on one-to-one relationships between data sources and sinks, CoFlow analysis detects scenarios where multiple sources converge into a single sink. This focus on many-to-one relationships makes it useful for uncovering behaviors like Fingerprinting or ID bridging.
CoFlow analysis builds upon the taint-flow tracking capabilities, utilizing our static taint analysis process to follow the movement of data throughout an application. It applies a set of constraints and rules to pinpoint those flow patterns that match the configured suspicious behavior definitions. CoFlow analysis uses a dedicated configuration file to define its behavior detection rules. This file specifies sets of source APIs and a corresponding sink API and includes fine-grained constraints to minimize false positives.
The analysis process centers on ensuring that, for each configured behavior, at least one source from each defined source group participates in that behavior. The system examines potential sinks, traces the flow of data backward, and compares discovered sources against the rule set. If a match is found for every source group within a rule, the system flags this as an instance of the suspicious behavior. Optimizations exist to streamline this process and improve efficiency. When CoFlow analysis detects a suspicious behavior, it generates an output including a set of sources (with one source representing each configured source group) and the associated sink.
Figure 10.CoFlow Analysis
Fingerprinting Detection
Fingerprinting detection leverages CoFlow analysis to identify fingerprinting behaviors in SDKs. A “fingerprinting behavior” is defined as a set of
𝑁
or more fingerprinting sources that flow to a common interesting sink. The analysis proceeds from the list of fingerprinting APIs observed in our Seed Set of SDKs by checking that at least
𝑁
of the data items from fingerprinting APIs are exfiltrated after possibly being combined into new data objects. The CoFlow source group for fingerprinting is configured with the set of signals collected by self-reported fingerprinting SDKs. The number of sources is quite large, and we expect that most fingerprinting SDKs will include only a subset of these sources in their identifiers. Fingerprinting sinks are configured in two groups: network APIs, which create the potential for fingerprint exfiltration over the network, and encryption functions. Other categories of sinks exist but were not included in this work. For example, fingerprinting sources may be collected in a map or JSON object and simply returned by a publicly visible SDK method. Similarly, an SDK may populate an argument that is passed by reference to a public method with fingerprinting sources.
Fingerprinting diverges slightly from other CoFlow use cases in that we derive higher confidence in the behavior being present as more CoFlows are detected. Therefore, it’s useful to include all detected sources in the output for fingerprinting behaviors, instead of the usual one source per group. Including all detected sources also allows for more thorough analysis of fingerprinting behaviors across the corpus of SDKs, as well as useful information for debugging and reverse engineering to confirm the behavior.
Challenges
We run our analysis on SDKs bundled with their dependencies, and their dependencies’ dependencies, and so on. Our first iteration of fingerprinting detection included taint sources from the entire SDK and dependency bundle. Because these bundles can become quite large, and taint tracking exhibits superlinear behavior with an increasing number of sources, we found that our analysis was prohibitively expensive for very large SDKs and SDKs with a high number of dependencies.
Additionally, tracking taint sources that originate in dependencies leads to identifying fingerprinting behaviors that are fully contained within a dependency, and potentially not utilized by the main SDK at all. This "over-detection" was very noisy, and we saw many SDKs being flagged for including the same popular fingerprinting SDKs as dependencies. Identifying which SDKs include fingerprinting SDKs as dependencies is valuable, but using expensive static analysis techniques to do it is not efficient, and the noise may obscure fingerprinting performed by the main SDK.
To solve for these issues, we reduced the scope of taint sources to only those originating in the main SDK, excluding all sources that originate in dependencies. Notably, taint sinks are still included if they are in a dependency, allowing us to catch cases where an SDK collects fingerprinting sources but uses a dependency to hash or exfiltrate them. We acknowledge that this trade-off means we will miss some cases of fingerprinting, mainly where an SDK uses dependencies that each fetch a number of sources below the threshold for fingerprinting, but that when combined in the main SDK do reach the threshold. The ability to detect boundaries of the SDK that are not interesting and can be excluded from analysis, as well as some dynamic analysis techniques, may be useful in filling this gap.
Figure 11.Fingerprinting detection using CoFlow. There are 2 fingerprinting behaviors in this example. Sources {1,2,3} SinkA, and source {3} SinkB. Sources 4, 5, and 6 are not included in the results because they are in a dependency.
Appendix CDetails of Our SDK Identification Approach
Any approach for SDK identification must not rely on having access to source code (since apps are primarily distributed in binary format), must not assume all SDK code is present in the app (since compilation and linking often minify, optimize, or remove SDK code), must handle code obfuscation (prevalent in Android apps, with more than 75% of them use obfuscation), must handle SDK dependencies (in 30,000 Java SDKs we analyzed, the 80 percentile of SDKs depended on 17 other SDKs), and must operate at scale (the number of SDKs and versions is on the order of
10
4
–
10
7
as the data from modulecounts.com shows, with growth rate of
10
2
–
10
3
new SDKs per day).
Our solution to SDK identification uses a fine-grained similarity metric for code that can be aggregated across code units (e.g., methods, classes, packages) and distribution units (e.g., SDKs, SDK families). The similarity metric for methods is based on a similarity across the following set of features: the type signature of the method, the Java- and Android-framework APIs invoked in the method body, the string constants used by the method body, and the histogram of instruction types present in the method body.
•
The type signature of the method captures the types of its parameters and the type of its return value. Since all types outside of the Java and Android frameworks are defined by the developer, their names are not trustworthy and cannot be relied upon for the purpose of similarity comparison. We eliminate all programmer-chosen type names to obtain an “anonymized” type signature for the method. The resulting type signature is one-hot encoded as a (boolean-valued) feature to be embedded into our vector space.
•
The Java- and Android-framework APIs invoked in the method body each becomes one feature, counting the number of invocations present in the method body. Such a feature does not account for indirect ways to invoke the framework APIs, for example via Java reflection, via code in other languages (e.g., native or Javascript code), or via dynamic code loaded at runtime.
•
The string constants used in the method body each form one (boolean-valued) feature, similar to the anonymized type signature of the method.
•
The instruction types present in the method body each form one feature, measuring the frequency of that specific instruction type in the method body. We consider 34 types of instructions, closely matching the instructions in the Dalvik VM specification, ranging from ASSIGN, to LOAD_INSTANCE, and to THROW.
The feature vector
𝑓𝑣
(
𝑚
)
∈
ℝ
1
×
𝑑
for a method
𝑚
is derived by combining all of the above features and embedding them into a space of
2
64
dimensions by hashing each feature name using a 64-bit hash function. The hash function needs to be collision resistant, to ensure that the adversary cannot easily create code similar to some target SDK code and evade SDK identification, but not necessarily pre-image resistant and thus does not need to necessarily be a cryptographic hash. This embedding step constructs a sparse vector representation for each Java method, since Java methods typically have only hundreds of non-zero features.
A preliminary data analysis showed that these features are non-uniformly distributed across SDKs. The leftmost bar in Figure 12 highlights how across a corpus of about 37,000 SDKs, there were about 11,000,000 features that occurred exactly in one SDK and thus could serve to make those SDKs uniquely reidentifiable in app code. As a result we add a weighing factor to each feature, to account for this non-uniformity:
wfv
(
𝑚
)
=
fv
(
𝑚
)
𝑇
×
weights
.
Figure 12.Features of SDK code are not uniformly distributed across a large set of SDKs. Features that appear in a single SDK are particularly useful for SDK matching and thus need to be weighed more in any similarity metric for SDK identification.
The high-dimensional space into which we represent Java methods naturally allows us to handle code composition. The vector representation of a group of methods (be they organized as a class, a package, an SDK, or other code structure) is the vector-sum of the corresponding vectors. We use cosine similarity as our distance metric for two vectors
fv
1
and
fv
2
:
𝛿
(
fv
1
,
fv
2
)
=
fv
1
fv
2
‖
fv
1
‖
‖
fv
2
‖
.
Each class is assigned a corresponding vector representation as vector-sum of its component methods’ vectors and by abuse of notation we write
fv
(
𝑐
)
for the vector of a class
𝑐
, such that
fv
(
𝑐
)
=
Σ
𝑚
∈
𝑐
fv
(
𝑚
)
. Then for each class in the app we search the set of all SDK classes for the classes that are minimally similar to the app class. Finally, we match SDKs that have sufficient support for being present in the app, measured by the number of classes in the SDK that are similar to app classes. Algorithm 1 presents the pseudo-code version of our technique.
The SDK-matching algorithm is parametrized by a threshold to ensure that app classes and SDK classes satisfy a minimum amount of similarity and by a second threshold to ensure an SDK is a match if a sufficient number of its classes appear in the app. In our experiments we use
0.2
for the similarity threshold and
0.55
for the class-count threshold.
The representation of methods and classes as sparse vectors into a high-dimensional space allows us to borrow solutions to the Approximate Nearest Neighbor (ANN) problem, well studied in the information-retrieval domain. In particular a number of high-performance libraries such as Faiss [20], SPTAG [43], ScaNN [28], Hnswlib [63], and NGT [62] provide capabilities to search for matches of thousands of classes in a typical app into a database of millions of SDK classes.
We evaluated the accuracy of this technique on a set of 302,397 SDKs for which we could find at least one mobile app that declared such an SDK as a dependency in its Gradle build configuration. The Gradle configuration data provided the ground truth against which we measured an average precision for SDK identification of
65.07
%
. Restricting to a smaller set of
9,683
SDKs, the SDK identification reached
99.89
%
precision at
46.16
%
recall. This implies that on this select set of SDKs, the identification technique can accurately detect the absence of an SDK from an app, though it may sometimes fail to detect the SDK’s presence in an app.
In the context of the measurements in the rest of the paper, this application of the SDK-identification technique provides lower bounds on the market reach (cf. RQ3), since it undercounts the presence of SDKs of interest.
Algorithm 1 SDK-Matching Algorithm
inputs :
𝐴
// an app
𝑆
1
,
…
,
𝑆
𝑚
// a set of SDKs
𝜂
// class-dissimilarity upper bound
𝛾
// SDK-similarity lower bound
outputs :
𝒯
// map from app classes to SDKs,
//
𝒯
:
𝐴
↦
2
𝑆
1
∪
⋯
∪
𝑆
𝑚
begin
let
𝐴
=
{
𝑐
1
,
…
,
𝑐
𝑛
}
be the set of app classes
let
𝒞
=
𝑆
1
∪
⋯
∪
𝑆
𝑚
be the set of SDK classes
// compute candidate SDKs for each app class
foreach
𝑐
𝑖
∈
𝐴
,
1
≤
𝑖
≤
𝑛
do
𝒯
(
𝑐
𝑖
)
←
{
𝑐
∈
𝒞
:
𝛿
(
𝑓𝑣
(
𝑐
𝑖
)
,
𝑓𝑣
(
𝑐
𝑦
𝑥
)
)
<
𝜂
}
// filter out SDKs with insufficient presence
for
𝑗
←
1
to
𝑚
do
if
|
𝑆
𝑗
∩
(
𝒯
(
𝑐
1
)
∪
⋯
∪
𝒯
(
𝑐
𝑛
)
)
|
≤
𝛾
then
for
𝑖
←
1
to
𝑛
do
𝒯
(
𝑐
𝑖
)
←
𝒯
(
𝑐
𝑖
)
∖
𝑆
𝑗
end
Appendix DLong-Form Definitions
Category
Definition
Ads
The SDK’s main purpose is to support displaying ads, ads bidding, ads targeting, ads mediation, or user analytics with the express purpose of monetization or conversion.
Examples: Mediation, ad adapters, ad push notifications
Analytics
Analytics (App Health)
The SDK’s main purpose is to track the systems performance of the application.
Examples: crash logging, battery tracking and usage, UI latency
Analytics (User Behavioral Analysis)
The SDK’s main purpose is to track user behavior on-device for business purposes (e.g. growth, retention, engagement, churn) without evidence of direct monetization or feedback into ads using the data collected.
Ex. App engagement measurements
Security & Authentication
The SDK’s main purpose is to protect either the app or the user against malware and fraud.
Security (Anti-Fraud)
SDKs that provide fraud detection capabilities to an app, to handle account fraud, payment fraud, or identity fraud. The fraud detection may be performed on device or on a remote server.
Security (Payments)
SDKs that provide payment functionality, potentially managing all aspects of the payment flow (authorization, clearing, settlement, reversal). All payment methods are covered here.
Security (Authentication)
SDKs that authenticate the user using the app against a local or remote account or identity. All forms of authentication mechanisms (password, biometric, hardware token, knowledge based) are covered here.
Security (Anti-malware)
SDKs that attempt to scan for malware, bugs, or other on-device security issues.
Security (Other)
SDKs that do not clearly fit into one of the above.
Tools / Other
Location
The SDK’s main purpose is to access location data and send it to a remote server.
Location Tracking (Person)
The SDK’s main purpose is to track user location for the purpose of collecting business data (e.g. how often does a person visit a certain location, like Foursquare) or for the purpose of locating family members.
Location Tracking (Object)
Specific object tracking (e.g. SDK that helps track a beacon / tag for finding lost objects)
Location Tracking (Maps)
Any kind of navigational functionality, including the display of maps and map routes, and the location-based retrieval of points of interest.
Location Tracking (Other)
SDKs that do not clearly fit into one of the above
Social
The SDK’s main purpose is to connect the user to a larger social network, or provide a means of discovering other people, or provide a list of contacts for social engagement. Customer service chat SDKs are not included in the social category.
Other
The SDK’s main purpose is to either support or provide functionality unrelated to any other defined categories. The SDK must have a purpose that is clearly defined in the information sources.
Unclear / Not found
The SDK’s main purpose is unclear using data from any of the allowed information sources, or, we cannot find any information on this SDK.
Note that this category is distinct from “Tools / Other”, which captures SDKs with clearly defined purposes.
Table 6.Codebook definitions of categories created by our labeling process.
Table 6 provides a full listing of the definitions developed and used in our labeling process. This table defines a classification system for Software Development Kits (SDKs) based on their primary functions, dividing them into five main categories: Ads, Analytics, Security and Authentication, Tools/Other, and Unclear/Not Found. The “Ads” category covers SDKs related to advertising; “Analytics” includes SDKs for tracking app health and user behavior; “Security and Authentication” encompasses SDKs for anti-fraud, payments, authentication, and anti-malware; “Tools / Other” includes location tracking (person, object, maps), social networking, and miscellaneous SDKs with defined but uncategorized purposes; and “Unclear / Not Found” is used when an SDK’s purpose cannot be determined. Some categories are further broken down into subcategories with specific definitions and examples, such as “Analytics (App Health)” or “Security (Anti-Fraud),” to ensure precise classification.
Our coders used this information to categorize an SDK by first reviewing available information about the SDK’s purpose and functionality. They then compared this information to the definitions provided in the table, starting with the main categories and then moving to the subcategories. By matching the SDK’s features to the descriptions in the table, the rater assigned the most appropriate label. If the SDK’s purpose is clearly defined but does not fit any existing category, they used the “Other” subcategory. If no information can be found, or the available information is insufficient to determine the SDK’s purpose, the rater classified it as “Unclear / Not Found.” This systematic approach ensured consistent and accurate labeling of SDKs based on their primary function.
Appendix EList of Signal APIs
Table 7.List of APIs used in fingerprinting-like behaviors observed in our SDK datasets.
Class Name
Property or Method Name
android.accessibilityservice.AccessibilityServiceInfo
getResolveInfo(), getSettingsActivityName()
android.accounts.Account
name
android.accounts.AccountManager
getAccounts(), getAccountsByType(com.google)
android.app.ActivityManager
getDeviceConfigurationInfo(), getRunningAppProcesses(), getRunningTasks(), isUserAMonkey()
android.app.ActivityManager\$MemoryInfo
availMem, lowMemory, totalMem
android.app.ActivityManager\$RunningTaskInfo
numRunning
android.app.KeyguardManager
isDeviceSecure(), isKeyguardSecure()
android.app.TaskInfo
baseActivity(), numActivities(), taskId()
android.app.UiModeManager
getCurrentModeType()
android.app.WallpaperInfo
getPackageName()
android.app.WallpaperManager
getDrawable(), getWallpaperInfo()
android.app.admin.DevicePolicyManager
getActiveAdmins(), getStorageEncryptionStatus()
android.app.usage.StorageStatsManager
getTotalBytes()
android.app.usage.UsageStats
getPackageName()
android.app.usage.UsageStatsManager
queryUsageStats(INTERVAL\_DAILY)
android.bluetooth.BluetoothAdapter
getAddress(), getBondedDevices(), getDefaultAdapter(), getName(), getScanMode(), getState(), isDiscovering(), isEnabled()
android.content.ClipboardManager
getPrimaryClip(), getPrimaryClipDescription()
android.content.ComponentName
toShortString()
android.content.ContentResolver
query(Uri(content://com.google.android.gsf.gservices)), registerContentObserver(android.provider.MediaStore\$Images\$Media.EXTERNAL\_CONTENT\_URI)
android.content.Context
getPackageName()
android.content.Intent
getIntExtra(health), getIntExtra(plugged), getIntExtra(scale), getIntExtra(status), getIntExtra(temperature), getIntExtra(voltage), getStringExtra(technology)
android.content.pm.ApplicationInfo
flags, loadLabel(), sourceDir
android.content.pm.ConfigurationInfo
reqGlEsVersion
android.content.pm.InstallSourceInfo
getInstallingPackageName()
android.content.pm.PackageInfo
firstInstallTime, lastUpdateTime, packageName, receivers, requestedPermissions, requestedPermissionsFlags, services, signatures, signingInfo, versionCode, versionName
android.content.pm.PackageItemInfo
metaData, name, packageName
android.content.pm.PackageManager
checkPermission(), getApplicationLabel(), getInstallerPackageName(8), hasSystemFeature(android.hardware.fingerprint), hasSystemFeature(android.hardware.location.gps), hasSystemFeature(android.hardware.sensor.accelerometer), hasSystemFeature(android.hardware.sensor.compass), hasSystemFeature(android.hardware.sensor.light), hasSystemFeature(android.hardware.telephony), hasSystemFeature, hasSystemFeature, queryIntentActivities(), resolveActivity(android.content.Intent(‘‘android.intent.action.CALL’’))
android.content.pm.SigningInfo
getApkContentsSigners()
android.content.res.Configuration
getLocales(), locale, orientation, screenLayout, uiMode
android.content.res.TypedArray
getDimensionPixelSize()
android.hardware.Sensor
getName(), getName(), getName(), getPower(), getVendor(), getVersion()
android.hardware.SensorEvent
values
android.hardware.SensorManager
getDefaultSensor(8)
android.hardware.camera2.CameraCharacteristics
CONTROL\_AE\_COMPENSATION\_RANGE, CONTROL\_AE\_LOCK\_AVAILABLE, CONTROL\_AF\_AVAILABLE\_MODES, CONTROL\_MAX\_REGIONS\_AF, FLASH\_INFO\_AVAILABLE, LENS\_FACING, REQUEST\_AVAILABLE\_CAPABILITIES, SCALER\_AVAILABLE\_MAX\_DIGITAL\_ZOOM, SCALER\_STREAM\_CONFIGURATION\_MAP, SENSOR\_INFO\_SENSITIVITY\_RANGE, STATISTICS\_INFO\_MAX\_FACE\_COUNT
android.hardware.camera2.CameraManager
getCameraIdList()
android.hardware.camera2.params.StreamConfigurationMap
getHighSpeedVideoFpsRanges(), getOutputSizes()
android.hardware.usb.UsbManager
getDeviceList()
android.location.Location
getAccuracy(), getAltitude(), getBearing(), getBearingAccuracyDegrees(), getElapsedRealtimeNanos(), getLatitude(), getLongitude(), getProvider(), getSpeed(), getSpeedAccuracyMetersPerSecond(), getTime(), getVerticleAccuracyMeters(), isFromMockProvider()
android.location.LocationManager
getBestProvider(), getLastKnownLocation(), isProviderEnabled(gps), isProviderEnabled(network)
android.media.AudioManager
getDevices(), getRingerMode(), getStreamMaxVolume(), getStreamVolume(), isMusicActive()
android.media.MediaDrm
getPropertyByteArray(deviceUniqueId)
android.media.RingtoneManager
getActualDefaultRingtoneUri(), getDefaultUri(), getRingtone()
android.net.ConnectivityManager
getActiveNetwork(), getActiveNetworkInfo(), getDefaultProxy(), getNetworkCapabilities()
android.net.NetworkCapabilities
hasTransport(4)
android.net.NetworkInfo
getState(), getTypeName(), isConnected(), isRoaming()
android.net.ProxyInfo
getHost()
android.net.TrafficStats
getTotalRxBytes(), getTotalTxBytes()
android.net.sip.SipManager
isVoipSupported()
android.net.wifi.ScanResult
capabilities(), frequency(), level(), SSID()
android.net.wifi.WifiInfo
getBSSID(), getFrequency(), getIpAddress(), getLinkSpeed(), getMacAddress(), getNetworkId(), getRssi(), getSSID()
android.net.wifi.WifiManager
calculateSignalLevel(), getConfiguredNetworks(), getConnectionInfo(), getScanResults(), getSSID(), is5GHzBandSupported(), isDeviceToApRttSupported(), isEnhancedPowerReportingSupported(), isP2pSupported(), isPreferredNetworkOffloadSupported(), isScanAlwaysAvailable(), isTdlsSupported(), isWifiEnabled()
android.os.BaseBundle
getBoolean(present), getLong(install\_begin\_timestamp\_seconds), getLong(referrer\_click\_timestamp\_seconds), getString(technology), getString(install\_referrer)
android.os.Build
BOARD, BOOTLOADER, BRAND, CPU\_ABI, CPU\_ABI2, DEVICE, DISPLAY, FINGERPRINT, getRadioVersion(), getSerial(), HARDWARE, HOST, ID, MANUFACTURER, MODEL, PRODUCT, RADIO, SERIAL, SOC\_MANUFACTURER, SOC\_MODEL, SUPPORTED\_32\_BIT\_ABIS, SUPPORTED\_64\_BIT\_ABIS, SUPPORTED\_ABIS, TAGS, TIME, TYPE, USER
android.os.Build\$VERSION
BASE\_OS, CODENAME, INCREMENTAL, RELEASE, SDK\_INT, SECURITY\_PATCH
android.os.Build\$VERSION\_CODE
FROYO, GINGERBREAD\_MR1, GINGERBREAD, HONEYCOMB\_MR1, HONEYCOMB\_MR2, HONEYCOMB, ICE\_CREAM\_SANDWICH\_MR1, ICE\_CREAM\_SANDWICH, JELLY\_BEAN\_MR1, JELLY\_BEAN\_MR2, JELLY\_BEAN, KITKAT\_WATCH, KITKAT, LOLLIPOP\_MR1, LOLLIPOP
android.os.Debug
isDebuggerConnected()
android.os.Environment
getDataDirectory(), getExternalStorageDirectory(), getExternalStorageState(), getRootDirectory(), isExternalStorageEmulated()
android.os.PowerManager
getCurrentThermalStatus(), getLocationPowerSaveMode(), isDeviceIdleMode(), isInteractive(), isPowerSaveMode()
android.os.StatFs
getAvailableBlocks(), getAvailableBlocksLong(), getBlockCount(), getBlockCountLong(), getBlockSize(), getBlockSizeLong(), getTotalBytes()
android.os.SystemClock
elapsedRealtime(), uptimeMillis()
android.os.SystemProperties
get(gsm.operator.isroaming), get(gsm.operator.numeric), get(gsm.sim.state), get(init.svc.qemu\-props), get(init.svc.qemud), get(qemu.hw.mainkeys), get(qemu.sf.fake\_camera), get(qemu.sf.lcd\_density), get(ro.bootloader), get(ro.bootmode), get(ro.hardware), get(ro.kernel.android.qemud), get(ro.kernel.qemu.gles), get(ro.product.device), get(ro.product.model), get(ro.product.name), get(ro.runtime.firstboot), get(ro.serialno)
android.os.UserManager
getSerialNumberForUser(), getUserProfiles(), isDemoUser(), isSystemUser(), supportsMultipleUsers()
android.os.storage.StorageVolume
isPrimary()
android.provider.Settings\$Global
getString(adb\_enabled), getString(airplane\_mode\_on), getString(airplane\_mode\_radios), getString(always\_finish\_activities), getString(animator\_duration\_scale), getString(auto\_time), getString(auto\_time\_zone), getString(bluetooth\_discoverability), getString(bluetooth\_discoverability\_timeout), getString(bluetooth\_on), getString(boot\_count), getString(data\_roaming), getString(development\_settings\_enabled), getString(device\_provisioned), getString(http\_proxy), getString(mode\_ringer), getString(network\_preference), getString(stay\_on\_while\_plugged\_in), getString(transition\_animation\_scale), getString(usb\_mass\_storage\_enabled), getString(use\_google\_mail), getString(wait\_for\_debugger), getString(wifi\_networks\_available\_notification\_on)
android.provider.Settings\$Secure
getInt(accessibility\_enabled), getInt(mock\_location), getString(accessibility\_display\_inversion\_enabled), getString(allowed\_geolocation\_origins), getString(default\_input\_method), getString(enabled\_input\_methods), getString(input\_method\_selector\_visibility), getString(install\_non\_market\_apps), getString(location\_mode), getString(skip\_first\_use\_hints), getString(tts\_default\_pitch), getString(tts\_default\_rate), getString(tts\_default\_synth), getString(tts\_enabled\_plugins)
android.provider.Settings\$System
getInt(screen\_brightness), getString(accelerometer\_rotation), getString(android\_id), getString(auto\_caps), getString(auto\_punctuate), getString(auto\_replace), getString(dtmf\_tone), getString(dtmf\_tone\_type), getString(end\_button\_behavior), getString(font\_scale), getString(haptic\_feedback\_enabled), getString(mode\_ringer\_streams\_affected), getString(mute\_streams\_affected), getString(notification\_sound), getString(ringtone), getString(screen\_brightness), getString(screen\_brightness\_mode), getString(screen\_off\_timeout), getString(show\_password), getString(sound\_effects\_enabled), getString(time\_12\_24), getString(user\_rotation), getString(vibrate\_on), getString(vibrate\_when\_ringing)
android.provider.Telephony\$Sms
getDefaultSmsPackage()
android.system.StructStat
st\_mtime
android.telecom.TelecomManager
isTtySupported()
android.telephony.CellIdentityGsm
getCid(), getLac(), getMcc(), getMnc()
android.telephony.CellIdentityLte
getCi(), getMcc(), getMnc(), getTac()
android.telephony.CellIdentityWcdma
getCid(), getLac(), getMcc(), getMnc()
android.telephony.CellInfo
isRegistered()
android.telephony.CellInfoCdma
getCellIdentity(), getCellSignalStrength()
android.telephony.CellInfoGsm
getCellIdentity(), getCellSignalStrength()
android.telephony.CellInfoLte
getCellIdentity(), getCellSignalStrength()
android.telephony.CellInfoTdscdma
getCellSignalStrength()
android.telephony.CellInfoWcdma
getCellIdentity(), getCellSignalStrength()
android.telephony.CellSignalStrength
getDbm()
android.telephony.CellSignalStrengthCdma
getDbm()
android.telephony.CellSignalStrengthGsm
getDbm()
android.telephony.CellSignalStrengthLte
getDbm()
android.telephony.CellSignalStrengthTdscdma
getDbm()
android.telephony.CellSignalStrengthWcdma
getDbm()
android.telephony.SignalStrength
getCellSignalStrengths()
android.telephony.SubscriptionInfo
getCarrierName(), getCountryIso(), getDataRoaming(), getDisplayName(), getIccId(), getNumber(), getSimSlotIndex()
android.telephony.TelephonyManager
getActiveModemCount(), getCellLocation(), getDataNetworkType(), getDataState(), getDeviceId(), getDeviceSoftwareVersion(), getGroupIdLevel1(), getImei(), getLine1Number(), getMeid(), getMmsUAProfUrl(), getMmsUserAgent(), getNetworkCountryIso(), getNetworkOperator(), getNetworkOperatorName(), getNetworkType(), getPhoneCount(), getPhoneType(), getSignalStrength(), getSimCountryIso(), getSimOperator(), getSimOperatorName(), getSimSerialNumber(), getSimSpecificCarrierIdName(), getSimState(), getSubscriberId(), getVoiceMailAlphaTag(), getVoiceMailNumber(), hasIccCard(), isHearingAidCompatibilitySupported(), isNetworkRoaming(), isSmsCapable(), isVoiceCapable(), isWorldPhone()
android.telephony.cdma.CdmaCellLocation
getBaseStationId(), getNetworkId(), getSystemId()
android.telephony.gsm.GsmCellLocation
getCid(), getLac()
android.util.DisplayMetrics
density, densityDpi, heightPixels, scaledDensity, widthPixels, xdpi, ydpi
android.view.Display
getMetrics(), getName(), getRealMetrics(), getRefreshRate(), getRotation(), getSize()
android.view.accessibility.AccessibilityManager
getEnabledAccessibilityServiceList(), getInstalledAccessibilityServiceList(), isTouchExplorationEnabled()
android.view.inputmethod.InputMethodInfo
getPackageName()
android.webkit.WebSettings
getDefaultUserAgent(), getUserAgentString()
android.widget.TextView
getTextSize()
androidx.ads.identifier.AdvertisingIdClient
getAdvertisingIdInfo(), isAdvertisingIdProviderAvailable()
androidx.ads.identifier.AdvertisingIdInfo
getId()
androidx.core.hardware.fingerprint.FingerprintManagerCompat
isHardwareDetected()
androidx.core.location.LocationManagerCompat
isLocationEnabled()
com.google.android.gms.ads.identifier.AdvertisingIdClient\$Info
getId()
com.google.android.gms.common.GoogleApiAvailability
isGooglePlayServicesAvailable()
com.google.android.gms.location.FusedLocationProviderClient
getLastLocation()
com.google.android.gms.location.LocationResult
getLastLocation()
com.scottyab.rootbeer.RootBeer
isRooted(), isRootedWithoutBusyBoxCheck()
java.io.File
(/system/fonts/), getFreeSpace(), getTotalSpace()
java.io.FileReader
(/proc/version)
java.io.RandomAccessFile
(/proc/meminfo), (/sys/devices/system/cpu/cpu0/cpufreq/stats/time\_in\_state), (sys/devices/system/cpu/cpu0/cpufreq/cpuinfo\_max\_freq), (sys/devices/system/cpu/cpu0/cpufreq/cpuinfo\_min\_freq)
java.lang.Class
forName(android.os.SystemProperties), forName(com.google.android.gms.loca\-tion.FusedLocationProviderClient), forName(dalvik.system.Taint), forName(androidx.ads.identifier.AdvertisingIdClient), forName(com.google.android.gms.ads.iden\-tifier.AdvertisingIdClient), getDeclaredField(sSystemFontMap), getDeclaredFields(), getDeclaredMethod(getFatVolumeId), getDeclaredMethods(), getMethod(getEnrolledFingerprints), getMethod(getFingerId)
java.lang.Runtime
availableProcessors(), maxMemory()
java.lang.System
currentTimeMillis(), getProperty(http.agent), getProperty(http.proxyHost), getProperty(http.proxyPort), getProperty(os.arch), getProperty(os.name), getProperty(os.version)
java.lang.Thread
getStackTrace()
java.net.InetAddress
getHostAddress()
java.net.NetworkInterface
getHardwareAddress(), getHostAddress(), getMTU(), getName(), isUp()
java.security.KeyStore
getCertificate(), getCreationDate()
java.security.cert.X509Certificate
getIssuerDN(), getKeyUsage()
java.util.Calendar
getCalendarType(), getTime(), getTimeInMillis()
java.util.Locale
getCountry(), getDefault(), getDisplayCountry(), getDisplayLanguage(), getDisplayName(), getDisplayVariant(), getLanguage(), getScript(), toString(), US()
java.util.TimeZone
getDefault(), getDisplayName(), getDSTSavings(), getId(), getOffset(), getRawOffset(), inDaylightTime(), useDaylightTime()
Table 7.(cont’d)
Experimental support, please view the build logs for errors. Generated by L A T E xml .
Instructions for reporting errors
We are continuing to improve HTML versions of papers, and your feedback helps enhance accessibility and mobile support. To report errors in the HTML that will help us improve conversion and rendering, choose any of the methods listed below:
Click the "Report Issue" () button, located in the page header.
Tip: You can select the relevant text first, to include it in your report.
Our team has already identified the following issues. We appreciate your time reviewing and reporting rendering errors we may not have found yet. Your efforts will help us improve the HTML versions for all readers, because disability should not be a barrier to accessing research. Thank you for your continued support in championing open access for all.
Have a free development cycle? Help support accessibility at arXiv! Our collaborators at LaTeXML maintain a list of packages that need conversion, and welcome developer contributions.
We gratefully acknowledge support from our major funders, member institutions, and all contributors.
About
·
Help
·
Contact
·
Subscribe
·
Copyright
·
Privacy
·
Accessibility
·
Operational Status
(opens in new tab)
Major funding support from