云环境中基于时钟偏移的客户端设备识别(AINA 2012) 原文标题:
内容概要总结
本文是 AINA 2012(第 26 届 IEEE 高级信息网络与应用国际会议)论文《Clock Skew Based Client Device Identification in Cloud Environments》(作者 Ding-Jie Huang、Kai-Ting Yang、Chien-Chun Ni、Wei-Chung Teng、Tien-Ruey Hsiang、Yuh-Jye Lee,台湾科技大学与石溪大学)。论文提出一种基于时钟偏移指纹的应用层轻量设备识别方法:用 AJAX 技术在云服务器端周期收集客户端时间戳(每个 AJAX 包带时间戳,sync.js 每 5 秒发一次),再以线性回归估计时钟偏移(散点图斜率)。论文给出时钟行为术语(偏移、频率、偏移率 skew α、漂移),比较四种线性回归机制(累积偏移、滑动窗口偏移、带下界滤波器的滑动窗口偏移、带下界滤波器的累积滑动窗口偏移,后者 20 包内收敛)及快速分段最小值算法,并首次提出跳变点(jump point)检测与消除方案(参数 p=20、k=20),以处理网络切换或时间同步引起的偏移漂移。实验:同一 MacBook 经 LAN/ADSL/3G/Wi-Fi/Tor/VM 接入,偏移在 2.16 ppm 内波动(-21.08 至 -23.24 ppm);100 台设备偏移从 67 ppm 到 -499 ppm。结论:容差阈值取 ±1 ppm 时,最坏情况假阳性率与假阴性率均不超过 8%(误判概率 7.4%)。
翻译内容
原文内容(English)
2012 26th IEEE International Conference on Advanced Information Networking and Applications
云环境中基于时钟偏移的客户端设备识别(Clock Skew Based Client Device Identification in Cloud Environments)
Ding-Jie Huang、Kai-Ting Yang、Chien-Chun Ni、Wei-Chung Teng、Tien-Ruey Hsiang 和 Yuh-Jye Lee
计算机科学与信息工程系
National Taiwan University of Science and Technology, Taipei City, Taiwan 106
Email: {D9515002, D9615007, weichung, trhsiang, yuh-jye}@mail.ntust.edu.tw
计算机科学系
Stony Brook University, Stony Brook, NY 11794, USA
Email: chni@cs.stonybrook.edu
摘要
随着云计算和移动设备的增长,云环境中客户端设备身份关切的重要性日益凸显。为提供一种轻量而可靠的设备识别方法,本文提出一种基于时钟偏移(clock skew)指纹的应用层方法。所开发的实验平台采用 AJAX 技术,在连接期间于云服务器端收集客户端设备的时间戳,然后计算客户端设备的时钟偏移。本文开发了几种基于线性回归和分段最小值(piecewise minimum)算法的方法,以优化精度并缩短时间戳收集过程。本文还提出一种跳变点(jump point)检测方案,以解决通常由切换网络或临时断连引起的偏移漂移问题。最后,进行了两个实验来研究时钟偏移指纹的有效性,结果表明,在适当设置容差阈值时,最坏情况下的假阳性率和假阴性率都不超过 8%。
关键词 —— 时钟偏移、设备身份、云服务、跳变点检测
I. 引言
近年来基于云的服务的发展,显著改变了人们使用计算机和移动设备的方式,也改变了安全攻击可能被发起的方式。云的架构在许多方面不同于经典计算机网络。网络连接成为诸如数据访问等基础操作的基础设施的一部分,而服务器威胁在虚拟化技术下被进一步隔离。因此,用户隐私、数据机密性和监管关切等经典安全问题应在云计算环境下被重新审视。
本文研究云安全的客户端设备识别问题。如今,人们常常通过个人设备(如手机、平板和笔记本电脑)订阅云服务。因此,用户身份可以被关联到专用硬件。云中的设备识别在检测未授权账号访问和定位被盗设备方面很有用。
设备识别通过维护一份与有效用户关联的已注册物理设备列表来实现。一旦用户通过未注册设备登录服务,服务提供者可能要求用户通过额外验证,或向系统管理员发出警报以处理可能的非法用户登录。许多候选者可以被用作物理设备的指纹,例如 IP 地址、MAC 地址、Cookie 或 Web 浏览器的配置 [1]。然而,这些属性存在易于伪造、缺乏唯一性以及随环境变化等弱点,使其不足以充当设备指纹。在本文中,我们提出时钟偏移,一种独特的指纹,来识别不同设备。
时钟偏移是两个时钟之间计时速度的差异。一般来说,现代处理器的时钟呈现两个属性 [2]–[4]:第一,两台设备之间的时钟偏移随时间相对稳定;第二,任意两台物理设备之间存在可区分的时钟偏移。根据这两个属性,时钟偏移可被视为任何具有数字时钟的设备的指纹。时钟偏移已被广泛用作一种攻击方法,来揭示 HoneyPot 或 Tor 网络背后的 Web 主机 [4]–[6]。后来,它也被用于验证传感器节点(sensor motes)的身份 [7], [8]。本工作将时钟偏移应用于识别客户端设备,以服务于云环境乃至经典 Web 系统中的服务器安全。
在本文中,我们从一个新颖的角度看待时钟偏移,把它视为一种对手检测机制。此外,为展示时钟偏移的可用性,构建了一个基于网页的偏移检测系统来收集时间戳,并在其上实现了五种不同的方法来计算客户端设备的时钟偏移。我们提供一种基于 AJAX 的技术,周期性地从云服务订阅者向服务器收集时间信息。设备指纹识别随后通过估计此时钟偏移来执行。通过在一个包含 100 多台设备的测试平台上进行的实验,时钟偏移能有效识别客户端设备,无论底层网络信道是 3G、Wi-Fi 还是 ADSL 等。
此外,在我们的实验中,我们观察到有时包偏移会剧烈漂移,从而造成一个跳变点。此问题主要由全局时间同步或网络连接切换(hand-off)引起。在先前的工作中,由于包收集周期很长,这一现象被忽略且未被解决。为最小化包收集周期并加速时钟偏移计算,本文提供了一个跳变点检测方案。
在本工作中,我们引入时钟偏移作为云服务中一种有效的客户端设备身份,并提供了一个针对跳变点问题的解决方案,据我们所知此前尚未被讨论过。
II. 相关工作
首先,我们回顾几种现代主机识别方案。然后,讨论为无线传感器网络和互联网开发的基于时钟偏移的技术。
A. 主机识别技术
在先前的研究 [9] 中,一旦对手成功伪造其身份以感染云服务器,对手就可能对服务提供者或客户端造成严重损害。因此,一种识别不同客户端的有效方案对任何基于 SaaS 的服务都至关重要。
另一方面,远程主机识别或指纹识别已被广泛研究。可能的应用从蓝牙信号 [10] 到 RFID 标签不等。Gerdes 等人 [11] 提出一种通过分析以太网设备的模拟信号来唯一识别它们的方法。后来,Rasmussen 等人 [12] 把射频指纹识别技术扩展到无线传感器网络,并通过实验证明了其可行性。Eckersley [1] 提出一种基于 Web 浏览器的设备指纹识别方法,使用浏览器的配置和插件信息来刻画设备。Eckersley 还表明 Web 浏览器的信息可以轻易被追踪。
B. 基于时钟偏移的攻击与防御方案
时钟偏移已被广泛用于各种攻击中以检测和揭示隐藏主机。在 [4] 中,Kohno 等人利用时钟偏移,通过暗中记录和分析远程物理设备的 ICMP 或 TCP 时间戳来对其进行指纹识别。然而,对于现实世界应用,使用 ICMP 和 TCP 时间戳有其局限。ICMP 时间戳被许多防火墙阻止,而某些操作系统默认禁用 TCP 时间戳。此外,他们的方法在 Tor 等匿名网络中失败,因为那里无法建立端到端 TCP 连接 [6]。
Murdoch [5] 提出一种基于时钟偏移的攻击来揭示隐藏服务。伪匿名服务器身份因时钟偏移的偏移而暴露,这种偏移源自服务器负载增加以及随之而来的 CPU 温度上升。此攻击在 Tor 网络中有效。然而,它高度依赖隐藏服务器通过 Tor 网络产生的大量流量。此外,它需要在短时间内获取大量时间戳以执行充分的时钟偏移估计。
后来 Zander 和 Murdoch 开发了一种使用同步采样技术的增强攻击 [6],它显著减少了量化误差,从而削减了先前攻击所需的大量网络流量。此外,他们的工作是第一个通过 HTTP 协议进行时钟偏移估计的工作。然而,这种攻击模型不能直接应用于服务器端以用于防御目的。我们稍后在第 III 节讨论其原因。
上述所有基于时钟偏移的技术都是探测互联网上服务提供者的攻击方法。为把时钟偏移部署为一种防御机制,Huang 等人 [7] 提出一种在无线传感器网络中利用时钟偏移进行节点识别的方法。在他们的论文中,时钟偏移指纹被提出作为针对女巫攻击(Sybil attack)的一种对策。Uddin 的工作 [8] 也验证了上述基于偏移的方案在隐蔽信道下是一种有效的指纹识别方法。
III. 基于时钟偏移的主机识别:场景与预备知识
在本节中,介绍基于时钟偏移的主机识别的场景和预备知识。第一部分呈现一个场景,说明基于时钟偏移的方案如何辅助设备识别和恶意对手检测。第二部分讨论计算时钟偏移的预备基础知识和相应术语。然后,给出实验平台的简要描述。在最后部分,进一步分析时间收集过程。
A. 基于时钟偏移的客户端设备识别场景
1) 设备识别系统构建: 客户端设备识别系统的场景如图 1 所示。设想一个用户试图登录一个基于 Web 的云服务器。为确认此次登录的有效性,服务器检查该用户是否已在偏移值数据库中拥有一台或多台已注册设备。
如果没有,服务器对客户端执行二次认证。目前一些流行的方法包括手机验证、电子邮件验证和交互式方法,等等。如果用户无法通过验证,登录被拒绝。通过识别过程的用户可以选择是否注册当前设备。在前一种情况下,一个时间戳收集服务器开始监视它自身与客户端设备之间的时间差,并相应地计算相对时钟偏移。估计出的时钟偏移值成为客户端设备的指纹,并被存储在数据库中供以后登录使用。
相反,如果数据库中存在已注册设备,服务器把当前客户端设备的时钟偏移与已注册的那个进行比较。如果这两个偏移之差在容差阈值以下,服务器接受该设备并认为客户端已通过验证。相反,如果差值超过阈值,此账号可能正遭受账号劫持攻击。在这种情况下,服务器随后要求用户提供进一步信息来验证其身份。接下来的步骤与没有已注册客户端设备的情况相同。最后,当检测到潜在恶意意图时,服务器应发出警报以通知关联的客户端。
2) 云服务的时间戳收集系统: 借助时钟偏移指纹识别技术,云服务获得针对恶意攻击者的额外保护。事实上,作为物理属性的设备指纹,优于 IP 地址、MAC 地址和 Cookie 等网络参数。此外,时钟偏移的两个主要属性——不易伪造和设备独特性——使我们方法成为面向云应用的一个有前景的候选。
如前所述,Zander 和 Murdoch 提供了一种基于 HTTP 请求的时钟偏移估计方法 [6]。然而,他们的方法不能直接应用于上述场景中的服务器端。这是因为他们的方法必须通过 HTTP 请求频繁询问远程主机的时间信息,这对 Web 服务器来说无法以同样方式执行。因此,云服务器需要一种不同的方法来从客户端获取时间信息。
为强制客户端把它自己的时间信息返回给服务器,我们的系统应用了 AJAX 技术,因为 AJAX 生成的每个包都包含相应的时间戳。由于 AJAX 由 JavaScript 实现,而 JavaScript 通常被 Web 浏览器支持,我们假设 JavaScript 的执行是被允许的。
时间戳收集系统的架构如图 2 所示。在客户端侧,客户端通过安全套接字层(SSL)连接到 Web 服务器以提供安全信道。在服务器侧,主 Web 服务器为客户端提供通用服务。同时,设置一个时间戳收集服务器来从客户端收集时间戳。
图 1:基于时钟偏移的主机识别系统流程图(用户输入用户名与密码 → 检查客户端是否有已注册设备 → 若无则检查客户端能否通过其他验证 → 通过则注册设备或拒绝登录 → 计算设备时钟偏移并加入已注册列表 → 数据库 → 登录;若已有注册设备则比较偏移差值是否小于阈值)
图 2:时间戳收集场景(客户端请求 index.php → 主 Web 服务器返回 index.php + sync.js → 客户端通过 AJAX 包(时间戳)定期发往时间戳收集服务器 → 计算的时钟偏移存入数据库服务器)
计算出的时钟偏移被存储在数据库服务器中供将来使用。一旦客户端向主 Web 服务器请求 index 页面,它会返回包含一个脚本 sync.js 的页面。此脚本随后要求客户端设备周期性地把它自己的时间信息发回给时间戳收集服务器。
事实上,我们已经构建了一个原型系统,按照上述场景执行时钟偏移识别实验。主 Web 服务器配备 Ubuntu 服务器操作系统,运行 Apache Web 服务器。由于我们的偏移估计方法基于 JavaScript,时间戳的精度被限制在 1 毫秒。
B. 时钟行为术语
用于表示时钟特性的术语如下。本文改编了来自 [3], [8], [13] 的命名法。把 Cx(t) 视为设备 x 的时钟在真实时间 t 所报告的时间,Cx′(t) ≜ dCx(t)/dt,Cx″(t) ≜ d2Cx(t)/dt2,t ≥ 0。设 Cc 和 Cs 分别为客户端和时间戳收集服务器的时钟:
- 偏移(Offset): Cc 报告的时间与 Cs 报告的时间之差,例如客户端时钟 Cc 相对于服务器时钟 Cs 的偏移是 Cc(t) - Cs(t),t ≥ 0。
- 频率(Frequency): 时钟推进的速率,例如 Cc 在时刻 t 的频率是 Cc′(t)。
- 偏移率/时钟偏移(Skew,α): 两个时钟频率之差,例如 Cc 相对于 Cs 在时刻 t 的偏移率为 α(t) = Cc′(t) - Cs′(t)。
- 漂移(Drift): Cc 相对于 Cs 在时刻 t 的漂移为 Cc(t) - Cs(t)。
根据上述定义,如果服务器累积了足够的客户端时间信息,此客户端的时钟偏移 α 便可由服务器在本地计算出来。
C. 时间戳的使用
借助 AJAX,服务器可以在从客户端返回的每个包中收集时间戳。除非另有说明,所有时钟偏移估计都基于服务器时钟 Cs。假设时间戳收集服务器已从某客户端收到 n 个 AJAX 包。令时间戳 tic 表示客户端发出第 i 个包时的 Cc 时间;类似地,时间戳 tis 表示服务器收到第 i 个包时的 Cs 时间。第 i 个包在服务器与客户端之间的估计偏移记为 oi,其中 oi = tis - tic。此外,根据服务器时钟 Cs,第 i 个包与第 j 个包之间的时段记为 xij,其中 xij = tjs - tis。
为更清楚地说明服务器时间与包偏移之间的关系,把 (tis, oi) 绘制成散点图;此外,服务器与客户端之间的时钟偏移 α 可以被估计为该图的斜率。一个例子如图 3 所示,其中这些数据集的趋势随负斜率递减,这意味着客户端与服务器之间的时钟偏移为负。
图 3:从时间戳收集服务器收集的偏移分布(含被标注的离群值)
IV. 时钟偏移估计
如前一节所示,偏移率可以被估计为散点图的斜率。然而,由于偶发的网络延迟或抖动,存在一些无法用于计算时钟偏移的离群节点。在本节中,说明两种基本技术来估计时间戳收集服务器与客户端之间的时钟偏移:线性回归和分段最小值算法。此外,对于基于时钟偏移的主机识别,我们实现五种方法来估计服务器与客户端之间的时钟偏移,并分析每种估计的性能。
A. 线性回归算法
线性回归是一种使一组数据点逼近一条直线的方法。尽管当数据集中存在显著离群值时此方法不稳健 [3],但它易于实现,且计算开销较小。这里,提出四种基于线性回归的不同类型机制来逼近时钟偏移。在以下段落中,LR(Nij) 表示对数据集 Nij 的线性回归计算,该数据集包含从 (tis, oi) 到 (tjs, oj) 的数据。
1) 累积偏移(Accumulated Skew): 对于累积偏移,当来自客户端的包被服务器收到时,服务器立即计算估计的偏移。在从客户端收到第 i 个请求时,估计的偏移可以表示为 LR(N1i)。
在累积偏移中,每一个数据(甚至离群值)都被累积进数据集。在大量数据和时间下,此方法提供稳定可靠的结果。然而,在短时间内,累积偏移会受离群值剧烈影响。此外,由离群值引起的误差会持续影响之后所有计算的结果。例如,如图 4(a) 所示,由于此时巨大的网络延迟,估计的偏移在红圈部分剧烈波动。
2) 滑动窗口偏移(Skew with Sliding-Windows): 与累积偏移相比,滑动窗口偏移只从最近一小段时间内采样。这防止了大幅波动的数据在长期内毒害偏移估计。
对于大小为 w 的采样窗口,滑动窗口偏移 LR(Nij) 必须满足 j - i = w。
图 4(b) 展示了窗口大小为 200 的滑动窗口偏移的结果;关于窗口大小的建议在第 V 节提供。
如结果所示,滑动窗口减少了滑动窗口之外离群值的影响,但如果离群值存在于窗口内部,此方法仍会受离群值效应影响。因此,为滤除这些离群值,需要一种选择合适内点(inliers)的方法,尤其是在巨大网络延迟随机发生的环境下。
顺带一提,人们可能会注意到图 4(b) 中的偏移估计在开始时全为 0。这是因为窗口大小 w 为 200,200 之前的偏移估计因数据不足被标记为 0。所有基于滑动窗口的方法都有此限制。
3) 带下界滤波器的滑动窗口偏移(Sliding-Windows Skew with Lower-Bound Filter): 为分解离群值造成的影响,最有效的方法是滤除它们,因此下界方法在这里会有帮助。例如,图 3 中红圈部分存在巨大网络延迟(即离群值);这些主要由网络延迟引起的点可能极大地影响偏移估计的正确性。相反,位于较低部分的偏移相对平滑。因此,最接近的偏移估计由下界数据集计算得出,正如先前研究所做的那样 [3]。在 [3] 中,Moon 等人建议把分段最小值算法作为一种简单高效的时钟偏移估计方法。由于分段最小值可用于提取偏移的下界,该算法适合充当低界滤波器。
对于带下界滤波器的滑动窗口偏移,在每个滑动窗口 w 中每 m 个包选取局部最小偏移。因此,用于偏移估计的采样数据量减少到 w/m。使用此下界滤波器,只收集具有最小网络延迟的包,从而最小化网络的影响。此下界偏移估计可以记为 LR(Min(Nij)),其中 Min(Nij) 是第 i 个请求与第 j 个请求之间局部最小偏移的数据集。
如图 4(c) 所示,偏移估计比图 4(b) 中的平滑得多;在此例中窗口大小 w 为 200,m 为 5,因此一次偏移估计的总数据为 40。通过此方法,时钟偏移估计因此稳定,并能够减少巨大网络延迟的影响。
4) 带下界滤波器的累积滑动窗口偏移(Accumulated Sliding-Windows Skew with Lower-Bound Filter): 由于局部最小偏移对找到下界偏移很有用,我们进一步用这些局部最小数据集计算累积偏移,并绘制相应的偏移分布图,如图 4(d) 所示。我们发现此方法既能减少巨大网络延迟的影响,又能在 20 个包内快速收敛。在从客户端收到第 i 个请求时,此偏移估计可以记为 LR(Min(N1i))。
结果,通过比较这四种类型的偏移估计,下界滤波器在我们的实验中对滤除巨大网络延迟的性能最好。
图 4:应用线性回归的各偏移估计比较((a) 累积偏移;(b) 滑动窗口偏移;(c) 带下界滤波器的滑动窗口偏移;(d) 带下界滤波器的累积滑动窗口偏移)
B. 快速分段最小值算法
快速分段最小值算法通过把数据分成 n 段,并分别从第一段和第四段中选取两个最小偏移来实现。通过连接这两个最小偏移,可以得到这条线的斜率。在我们的实验中,n 被设为 4。如图 5 所示,黑色圆圈是偏移,红色方块是对应的段,两个蓝色方块是两个局部最小偏移。因此,估计的偏移是两个蓝色方块之间黑线的斜率。
只要数据集足够稳定,快速分段最小值算法就能以很少的计算量实现高稳定性。图 6 展示了基于与图 3 相同数据用此方法估计的偏移。由于快速分段最小值算法计算开销低,它可以高效地应用于其他云服务。
图 5:快速分段最小值算法的实现
图 6:应用快速分段最小值算法的偏移估计
C. 跳变点检测与消除
在包收集期间,如果客户端正在与时间服务器进行时间同步,或在不同网络提供商之间漫游,就会发生偏移的跳变点。如图 7 所示,一个跳变点示例被红圈标注。如果此跳变点现象未被服务器检测到,偏移估计过程将受大量偏移误差影响,反应为不准确的结果。因此,在实现偏移估计时,需要一种检测跳变点发生的方法。
图 7:服务器与客户端之间含一个跳变点的偏移分布图
此算法的主要思想是把顺序时间戳分成组,其中一组内每一对连续时间戳只产生很小的时间差。假设 diff 表示两个相邻偏移之间的差;diff 记为 dij,其中 dij = oj - oi,i ≠ j,i < j。此外,设置另一个阈值 k 来检查 diff 是否过大,从而检测跳变点。基本的跳变点检测算法如下:
- 每 p 个包选取局部最小偏移。
- 计算每一对连续局部最小偏移之间的 diff。
- 后续检测过程分为两类: a) 如果导出的 diff 全为正或全为负:
- 把导出的 diff 的中位数记为 Med(diff)。
- 如果存在某个 diff 满足 diff > k × Med(diff),则在这 p 个包内存在一个跳变点。 b) 如果只有部分导出的 diff 为正:
- 如果在 x 处一个负 diff 紧跟在正 diff 之后,则此 x 就是跳变点。
- 类似地,正 diff 紧跟在负 diff 之后的情况反过来处理。
图 8:两个段之间的跳变点检测
为防止跳变点造成的误差,偏移估计在跳变点之前和之后分别由快速分段最小值算法执行。如图 8 所示,当跳变点位于第 j 段时,前半部分的估计偏移由第 1 段到第 (j-2) 段计算;后半部分由第 (j+1) 段到最后一段计算。有两段被排除在偏移估计之外,因为跳变点可能位于第 (j-1) 段或第 j 段。通过这种分段方法,所有候选偏移都可以被计算出来。为表示所有偏移的适中值,我们选取这些候选偏移的中位数作为对应客户端的估计偏移。
跳变点的一种同类情况是时间间隙(time gap),即一段空白包的时间段,如图 9 所示。时间间隙通常在用户移动时出现,导致网络连接变化,从而产生一段消失期。为检测这类跳变点,只需检查 xij,其中 j = i + 1。如果 xij 大于所指定 AJAX 包发送周期的两倍,则被视为一个时间间隙,并同样被视为一个跳变点处理。
图 9:服务器与客户端之间含网络断连的偏移分布图
V. 实验与评估
为评估所提出的主机识别方法,本节进行两个不同的实验。我们实验的主要参数设置如下。
为获得高精度的偏移,每个实验至少收集 200 个 AJAX 包用于偏移估计。由于 JavaScript 的时间分辨率为 1 ms,要检测到低至 1 微秒分辨率的偏移至少需要 1000 秒。由于我们的 AJAX 脚本每 5 秒生成一个包,所需的最少包数为 1000/5,即 200。
对于跳变点检测,参数 p 设为 20,阈值 k 也设为 20。p 和 k 可根据不同网络环境调整;使用较小的值,跳变点检测过程会产生更敏感的结果。
A. 同一设备在不同网络下
第一个实验使用一台笔记本电脑通过多种网络接入技术连接到我们的时间戳收集系统。此实验也包括 Tor 网络背后的主机和虚拟机(VM)以供参考。此实验的目的是测量一台设备在不同网络环境下的偏移变化,从而模拟一台设备从各种网络登录云服务的情形。在此实验中,设备是一台 Apple MacBook,配备 2.4 GHz Intel Core 2 Duo 处理器,运行 MAC OS X 10.6.8,估计偏移由跳变点检测结合快速分段最小值算法计算。如表 I 所示,我们首先把此笔记本电脑通过 LAN、ADSL、3G 和 Wi-Fi 等常见网络连接到服务器。
根据前四个结果,它证明了时钟偏移估计与网络接入媒介无关,是相对独立的。值得注意的是,估计的偏移在 2.16 ppm 范围内波动,从 -21.08 ppm 到 -23.24 ppm。由于时钟偏移可能随温度波动(如 [5], [6], [8] 所述),考虑到时间戳收集时间短以及可能与网络环境相关的噪声,这些实验结果似乎是可接受的。
表 I:同一设备在不同网络环境下的估计偏移
| 网络类型 | 偏移估计 | 包数 | IP 地址数 |
|---|---|---|---|
| LAN | -21.91 ppm | 1001 | 1 |
| ADSL | -23.24 ppm | 207 | 1 |
| 3G | -22.74 ppm | 13322 | 1 |
| Wi-Fi | -21.48 ppm | 5837 | 1 |
| Tor | -21.08 ppm | 1400 | 1 |
| -23.24 ppm | 951 | 1 | |
| VM | -23.71 ppm | 1027 | 1 |
| -21.79 ppm | 9810 | 1 | |
| -23.06 ppm | 1470 | 1 | |
| -22.53 ppm | 15007 | 55 | |
| -23.22 ppm | 12922 | 57 | |
| -22.88 ppm | 24120 | 108 | |
| -113.19 ppm | 868 | 1 | |
| -114.22 ppm | 1001 | 1 | |
| -6.40 ppm | 1001 | 1 | |
| -6.83 ppm | 890 | 1 |
由于存在偏移波动,为检测恶意登录设置一个合理的容差阈值是一种权衡。如果阈值太小,假阳性率会高得不可接受,这意味着同一台设备可能经常被当作不同设备。相反,如果阈值太大,假阴性率会升高,导致接受未注册设备的概率很高。因此,在分析表 I 之后,我们认为在当前阶段 ±1 ppm 的阈值是合适的。此阈值产生的误判概率为 7.4%,即 (2.16 - 2)/2.16,假设偏移波动的概率密度函数均匀分布。
此外,Tor 中的实验环境也列在表 I 中。在我们的实验结果中,200 个包不足以计算稳定的估计。然而,只要收集时间足够长,估计的偏移就会趋于一致并接近常见网络的情形。最后,对于运行在 VM 内的客户端主机,偏移稳定但会与真实值产生不可预测的差异。Kohno 等人 [4] 曾指出虚拟机不具有恒定的时钟偏移,我们的实验也显示了相同结果。此外,我们发现 VM 下的估计偏移在不重启它时相对稳定;然而在重启 VM 之后,偏移会随机变化到另一个稳定值。如表 I 所示,两个客户端的估计偏移分别为 -113.19 ppm 和 -114.22 ppm,但在重启系统后它们的偏移变为 -6.4 ppm 和 -6.83 ppm。
B. 不同设备的偏移分布
为研究时钟偏移的可区分性,我们收集了由同一时间戳收集服务器估计的 100 台设备的时钟偏移。结果从 67 ppm 到 -499 ppm 不等,图 10 中只展示了按值排序后最接近的 90 个偏移。此图中每个十字代表一台设备,y 坐标代表其时钟偏移。每个时钟偏移被一个 ±1 ppm 范围的区间界定。如图所示,许多设备的区间与其他设备重叠,这意味着存在不可忽略的概率,用户可能用未注册设备通过识别测试。最坏情况发生在第 20 号设备,其区间与其他 8 台设备重叠。因此,在 1 ppm 的阈值下,目前最大假阴性率为 8%。我们相信,对不同网络环境下偏移估计的进一步分析将有助于降低容差阈值,从而同时降低假阳性率和假阴性率。
图 10:估计设备的排序偏移(横轴为设备编号 1–100,纵轴为估计偏移值 × 10-5)
VI. 结论与未来工作
本文讨论了云环境中的客户端设备识别问题,并提出了一种基于时钟偏移的指纹识别技术和一个实用场景。客户端设备识别增强了账号安全,且所提出的场景在大多数时候建议了一种用户无感知的潜在二次认证方法。本文引入并实现了几种经典方法来估计客户端设备的偏移,并首次讨论了通常由切换网络或临时断连引起的跳变点的处理。为检验时钟偏移的有效性,我们实现了一个基于 Web 的偏移估计平台并进行了两个实验。初步研究包括来自 5 种不同网络媒介的 100 多台客户端设备。实验结果表明,当容差阈值设为 ±1 ppm 时,最坏情况下的假阳性率和假阴性率都不超过 8%。随着时钟偏移指纹被揭示的潜力,进一步的工作将包括利用线性规划方法提高偏移估计的精度,以及积累不同网络媒介中偏移波动的知识。
致谢
本工作在国家科学委员会(National Science Council)资助 100-2218-E-011-008 下完成。
参考文献
[1] P. Eckersley, "How unique is your web browser?" in Privacy Enhancing Technologies. Springer Berlin / Heidelberg, 2010, vol. 6205, pp. 1–18.
[2] V. Paxson, "On calibrating measurements of packet transit times," in Proceedings of the 1998 ACM SIGMETRICS joint international conference on Measurement and modeling of computer systems, ser. SIGMETRICS '98/PERFORMANCE '98. New York, NY, USA: ACM, 1998, pp. 11–21.
[3] S. Moon, P. Skelly, and D. Towsley, "Estimation and removal of clock skew from network delay measurements," in INFOCOM '99. Eighteenth Annual Joint Conference of the IEEE Computer and Communications Societies. Proceedings. IEEE, vol. 1, mar 1999, pp. 227–234 vol.1.
[4] T. Kohno, A. Broido, and K. Claffy, "Remote physical device fingerprinting," in IEEE Transactions on Dependable and Secure Computing, vol. 2, no. 2, April-June 2005, pp. 93–108.
[5] S. J. Murdoch, "Hot or not: revealing hidden services by their clock skew," in CCS '06: Proceedings of the 13th ACM Conference on Computer and Communications Security, New York, NY, USA, 2006, pp. 27–36.
[6] S. Zander and S. J. Murdoch, "An improved clock-skew measurement technique for revealing hidden services," in Proceedings of the 17th conference on Security symposium. Berkeley, CA, USA: USENIX Association, 2008, pp. 211–225.
[7] D.-J. Huang, W.-C. Teng, C.-Y. Wang, H.-Y. Huang, and J. M. Hellerstein, "Clock skew based node identification in wireless sensor networks," in IEEE Global Communications Conference, 2008, pp. 1877–1881.
[8] M. Uddin and C. Castelluccia, "Toward clock skew based wireless sensor node services," in Wireless Internet Conference (WICON), 2010 The 5th Annual ICST, March 2010, pp. 1–9.
[9] Top threats to cloud computing v1.0. [Online]. Available: https://cloudsecurityalliance.org/topthreats/csathreats.v1.0.pdf
[10] J. Hall, M. Barbeau, and E. Kranakis, "Detection of transient in radio frequency fingerprinting using signal phase," in Proceedings of IASTED International Conference on Wireless and Optical Communications (WOC '03), 2003.
[11] R. M. Gerdes, T. E. Daniels, M. Mina, and S. F. Russell, "Device identification via analog signal fingerprinting: A matched filter approach," in Proceedings of the 2006 Network and Distributed System Security Symposium (NDSS '06), February 2006.
[12] K. Bonne Rasmussen and S. Capkun, "Implications of radio fingerprinting on the security of sensor networks," Sept. 2007, pp. 331–340.
[13] S. Jana and S. K. Kasera, "On fast and accurate detection of unauthorized wireless access points using clock skews," in MobiCom '08: Proceedings of the 14th ACM international conference on Mobile computing and networking, 2008, pp. 104–115.
2012 26th IEEE International Conference on Advanced Information Networking and Applications
Clock Skew Based Client Device Identification
in Cloud Environments
Ding-Jie Huang , Kai-Ting Yang , Chien-Chun Ni , Wei-Chung Teng , Tien-Ruey Hsiang , and Yuh-Jye Lee
Department of Computer Science and Information Engineering
National Taiwan University of Science and Technology, Taipei City, Taiwan 106
Email: {D9515002, D9615007, weichung, trhsiang, yuh-jye}@mail.ntust.edu.tw
Department of Computer Science
Stony Brook University, Stony Brook, NY 11794, USA
Email: chni@cs.stonybrook.edu
Abstract--Along with the growth of cloud computing and possible invalid user login. Many candidates can be used as
mobile devices, the importance of client device identity concern the fingerprint of physical devices, such as IP address, MAC
over cloud environment is emerging. To provide a lightweight address, cookie, or web browser's configurations [1]. However,
yet reliable method for device identification, an application layer these attributes suffer the weaknesses of easy to forge, lack of
approach based on clock skew fingerprint is proposed. The uniqueness, and varying with environments, which make them
developed experimental platform adapts AJAX technology to insufficient to serve as device fingerprinting. In this paper,
collect the timestamps of client devices in the cloud server during we propose clock skew, a distinctive fingerprint, to identify
connection time, then calculate the clock skews of client devices. different devices.
Few methods based on linear regression and piecewise minimum
algorithm are developed to optimize the precision and shorten Clock skew is the difference of clocking speed between
timestamp collection process. A jump point detection scheme two clocks. Generally, modern processor's clocks present two
is also proposed to resolve the offset drifting problem, which is properties [2][4]: first, the clock skew between two devices
usually caused by switching network or temporary disconnection. is relatively stable over time; second, there is distinguishable
Finally, two experiments are conducted to study the effectiveness clock skew between any two physical devices. According
of clock skew fingerprint, and the results illustrate that the to these two properties, clock skew can be regarded as a
false positive rate and the false negative rate, in the worst case, fingerprint of any device with digital clock. Clock skew has
are both no more than 8% when the tolerance threshold is set been widely used as an attack method to reveal a web host
appropriately. behind HoneyPot or Tor networks [4][6]. Later, it is also
used to verify the identity of sensor motes [7,8]. This work
Index Terms--clock skew, device identity, cloud service, jump applies clock skew to identify client devices for server security
point detection in cloud environments, or even in classic web systems.
I. INTRODUCTION In this paper, we treat clock skew from a novel aspect
and regard it as a detection mechanism of adversary. Also,
The growth of cloud-based services in recent years has to display the usability of clock skew, a web page based
significantly changed the way how people use computers and skew detection system is constructed to collect the timestamps,
mobile devices, and also the way security attacks may be and five different methods are implemented on it to calculate
launched. The architecture of cloud differs from classic com- the clock skews of client devices. We provide an AJAX-
puter network in many ways. Network connections become based technique to periodically collect time information from
part of infrastructure on basic operations like data accessing, cloud service subscribers to the server. Device fingerprinting
and the server threats are further isolated under virtualization is then performed by estimating the clock skew via this time
technologies. Therefore, classic security issues such as user information. Through experiments performed on a testing plat-
privacy, data confidentiality and regulation concerns should form containing over 100 devices, the clock skew effectively
be reconsidered under the cloud computing environments. identifies client devices regardless of underlying networking
channels such as 3G, Wi-Fi, or ADSL, etc.
This paper studies the client device identification problem
of cloud security. Nowadays, people often subscribe cloud Furthermore, during our experiments, we observed that
services through personal devices such as mobile phones, sometimes packet offsets drift dramatically and thus cause a
tablets, and laptop computers. Therefore, user identity can be jump point. This problem is mainly caused by global time
associated to dedicated hardware. Device identification in the synchronization or network connection hand-off. In previous
cloud is useful in detecting unauthorized account access and work, since the packet collection period is long, this phe-
locating stolen devices. nomenon is ignored and have not been solved. To minimized
the packet collection period and speed up the clock skew
Device identification is realized by maintaining a registered computation, a jump point detection scheme is provided.
list of physical devices which are associated with valid users.
Once a user logs into the service via an unregistered de-
vice, service provider may ask the user to pass additional
verification, or raise an alarm to the system supervisor for
1550-445X/12 $26.00 2012 IEEE 526
DOI 10.1109/AINA.2012.51
In this work, we introduce clock skew as a effective client applied to server side for defense purpose. We discuss more
device identity in cloud services, and provide a solution to about the reason later at section III.
the jump point problem which to our knowledge has not been
discussed before. All of the above clock skew based techniques are attacking
methods that probe service providers on the Internet. To deploy
II. RELATED WORK clock skew as a defense mechanism, Huang et al. [7] proposed
a method to utilize clock skew on node identification in wire-
At first, we review several modern host identification less sensor networks. In their paper, clock skew fingerprinting
schemes. Then, clock skew based techniques developed for is proposed as a countermeasure against Sybil attack. Uddin's
wireless sensor networks and the Internet are discussed. work [8] also verified that the above-mentioned skew-based
scheme an effective fingerprinting approach if under a covert
A. Host Identification Techniques channel.
In previous studies [9], once an adversary successfully forge III. CLOCK SKEW BASED HOST IDENTIFICATION:
its identity to infect the cloud server, the adversary might SCENARIO AND PRELIMINARIES
prompt serious damage either to the service provider or to
the client. Thus, an effective scheme to identifying different In this section, the scenario and preliminaries of clock
clients is essential for any SaaS based service. skew based host identification are introduced. In the first part,
a scenario of how clock skew based scheme assists device
On the other hand, Remote host identification or finger- identification and malicious adversary detection is presented.
printing has been studied broadly. Possible applications vary The second part discusses preliminary basic knowledge of
from Bluetooth signal [10] to RFID tags. Gerdes et al. [11] computing clock skews and corresponding terminology. Then,
proposed a method to uniquely identify Ethernet devices by a brief description of the experimental platform is given. In
analyzing their analog signals. Later, Rasmussen et al. [12] the last part, the time collection process is further analyzed.
extended the radio fingerprinting technology to wireless sensor
networks, and demonstrated its feasibility through experi- A. Scenario of Clock Skew Based Client Device Identification
ments. Eckersley [1] presented a web browser-based device
fingerprinting approach which uses browser's configurations 1) Device identification system construction: The scenario
and plug-in information to characterize the device. Eckersley's of the client device identification system is showed in Fig. 1.
also showed that web browser's information can be easily Consider a user trying to login to a web-based cloud server.
traced. To confirm the validity of this login, the server checks if this
user already has one or more registered devices in the skew
B. Clock Skew based Attack and Defense Schemes value database.
Clock skew has been widely used in various attacks to detect If not, the server performs secondary authentication on the
and to reveal hidden hosts. In [4], Kohno et al. exploited client. Some currently popular methods include cell phone
clock skew to fingerprint a remote physical device by stealthily verification, email verification, and interactive method, to
record and analyze its ICMP or TCP timestamps. However, name but a few. If the user cannot pass the verification, login
for real world applications, using ICMP and TCP timestamps is denied. Users who passed the identification process may
have their limitation. ICMP timestamps are blocked by many choose whether to register the current device or not. In the
firewalls, and some operating systems in default disable TCP former case, a timestamp collection server starts to monitor
timestamps. Furthermore, their approach failed in anonymous the time difference between itself and the client device, and
networks like Tor in which end-to-end TCP connection is not calculates the relative clock skew accordingly. The estimated
possible [6]. clock skew value become the fingerprint of the client device,
and is stored in a database for later login.
Murdoch [5] proposed a clock skew based attack to re-
veal hidden services. The pseudonymous server identity was On the contrary, if there exist registered device in the
revealed due to the shift of clock skew which results from database, the server compares the clock skew of current client
increased server load and accordingly the CPU temperature. device with the registered one. If the difference of these two
This attack works in Tor network. However, it highly relies skews is under a tolerance threshold, the server accepts the
on large amount of traffic from the hidden server through device and consider the client has passed the verification. In
Tor network. Also, it requires large amount of timestamps contrast, if the difference exceeds the threshold, this account
in a short period of time to perform adequate clock-skew might be under an account hijacking attack. In this case, the
estimation. server then requires the user to provide further information to
verify his/her identity. The following steps are the same with
Later Zander and Murdoch developed an enhanced attack the no registered client device case. Finally, the server should
with synchronized sampling technique [6] which significantly raise an alarm to notify the associated client when potential
reduces the quantization error and thus cut the heavy network malicious intention is detected.
traffic necessary to previous attack. Also, their work is the
first one that undertakes the clock skew estimation through 2) Timestamp collection system of a cloud service: With
HTTP protocol. However, this attack model can not be directly the aid of clock skew fingerprinting technique, a cloud service
gains additional protection against malicious attackers. In
527
Clients key in Server Client
username & password
Yes Check if the client has No Request
index.php
a registered device Main Web
Server Return
index.php + sync.js
Collecting time information and Laptop
estimating clock skew
Database AJAX packets
(timestamp)
Firewall
o
Yes | skew difference | No Timestamp o
Collection o
< threshold Yes Check if the client can No Server
pass other verification
Fig. 2. Scenario of timestamp collection.
Pass Register the client Reject to
login
verification Yes device or not No
and login The calculated clock skew are stored in the database server
for future use. Once a client requests the index page from the
Calculate the Login main web server, it returns the page including a script sync.js.
clock skew of This script then asks the client device to periodically send back
the device, then its time information to the timestamp collection server.
add it to In fact, we have built a prototype system to perform
registered list experiments of clock skew identification following the above
scenario. The main web server equips Ubuntu server operat-
DB ing system, running Apache web server. Because our skew
estimating method is based on Javascript, the accuracy of
Login timestamps is limited in one millisecond.
Fig. 1. Flowchart of clock skew based host identification system.
B. Terminology of Clock Behavior
fact, device fingerprint, as a physical attribute, is superior
to network parameters like IP address, MAC address, and The terminology used to represent the clock characteristics
cookie. Moreover, clock skew's two major properties, uneasy is as follows. The nomenclature from [3,8,13] is adapted is
to forge and device distinctness, solid our method as a prospect this paper. Consider Cx(t) as the time reported by the clock
candidate for cloud-based applications. of device x at real time t, Cx (t) dCx(t)/dt and Cx(t)
d2Cx(t)/dt2, t 0. Let Cc and Cs be the clocks of client
As stated earlier, Zander and Murdoch provided a clock and the timestamp collection server respectively:
skew estimating method based on HTTP request [6]. However,
their method cannot be directly applied at the server side in the 1) Offset: The difference between the time reported by Cc
above scenario. This is because their approach must frequently and by Cs, e.g., the offset of the client clock Cc relative
asks for the time information of remote host through HTTP to the server clock Cs is Cc(t) - Cs(t), t 0.
request, which is infeasible for a web server to perform in the
same way. Therefore, a cloud server needs a different approach 2) Frequency: The rate at which the clock progresses, e.g.,
to obtain time information from a client. the frequency at time t of Cc is Cc (t).
To force the client to return its own time informations back 3) Skew (): The difference in the frequencies of two
to server, AJAX technique is applied in our system because clocks, e.g., the skew of Cc relative to Cs at time t
every packet generated by AJAX contains the corresponding is (t) = Cc (t) - Cs (t).
timestamp. Since AJAX is implemented by Javascript and
Javascript is generally supported by web browsers, we assume 4) Drift: The drift of Cc relative to Cs at time t is Cc(t)-
that the execution of Javascript is allowed. Cs(t).
The architecture of the timestamp collection system is According to the above definitions, if the server accumulates
illustrated in Fig. 2. On the client side, a client connects to sufficient client time information, the clock skew of this
web server through Secure Sockets Layer (SSL) to provide client can be computed by the server locally.
a secured channel. On the server side, the main web server
provides general services to clients. Meanwhile, a timestamp C. Usage of Timestamps
collection server is settled to gather timestamps from clients.
With AJAX, timestamps can be collected by server in every
packet returned from client. Unless further specified, all clock
skew estimation is based on server clock Cs. Assume that
the timestamp collection server has received n AJAX packets
528
Outlier can be represented as LR(N1i), while receiving ith request
from the client.
Fig. 3. Distribution of offset collected from the timestamp collection server.
In accumulated skew, every data, even outlier, is accumu-
from a client. Let timestamp tic represents the Cc time when ith lated in the data set. With vast data and time, this method pro-
packet sent out by the client; similarly, timestamp tis denotes vides stable and reliable result. However, in a short period of
the Cs time when the ith packet is received by the server. The time, accumulated skew is dramatically influenced by outliers.
estimated offset between server and client for the ith packet Moreover, errors caused by the outliers continuously affect
is denoted as oi, where oi = tsi - tic. Also, the period between the result for all the following computation. For example, as
the ith packet and the jth packet according to the server's shown in Fig. 4(a), estimated skews fluctuate enormously at
clock Cs is denoted as xij, where xij = tjs - tis. the red circle part due to huge network delay at this time.
To illustrate the relation between server time and packet 2) Skew with Sliding-Windows: Comparing with the accu-
offset more clearly, (tis, oi) is plotted into a scatter diagram; mulated skew, a sliding-window skew takes sampling only
furthermore, the clock skew between server and client can be from the most recently short period of time. This prevent the
estimated as the slope of this diagram. An example is shown largely fluctuated data from poisoning skew estimation in long
in Fig. 3 , where trend of these dataset is decreasing with term.
a negative slope, which means that the clock skew between
client and server is negative. For sampling windows with size w, the sliding-windows
skew LR(Nij) must satisfy j - i = w.
IV. CLOCK SKEW ESTIMATION
Fig. 4(b) shows the result of sliding-windows skew with
As shown in previous section, skew can be estimated as window size 200; the suggestion on the window size is
the slope of the scatter diagram. However, due to occasionally provided in section V.
network delay or jitter, there are some outlier nodes that can
not be used to compute the clock skew. In this section, two As the result indicates, sliding-windows reduced the effect
basic techniques are illustrated to estimate the clock skew of outliers which is outside the sliding window, but if outliers
between the timestamp collection server and the client: lin- exist inside the window, this method still suffer from the
ear regression and piecewise minimum algorithm. Moreover, outlier effect. Therefore, to filter out these outliers, a method
for clock skew based host identification, we implement five to choose the suitable inliers is needed, especially under the
methods to estimate the clock skew between the server and environment that huge network delay happens randomly.
the client, and analyze the performance of each estimation.
Digressively, one may notice that the skew estimation in
A. Linear Regression Algorithm Fig. 4(b) are all 0 in the beginning. This is because the window
Linear regression is a method for a set of data points size w is 200, the skew estimation before 200 are marked as 0
for no enough data. All sliding-windows based methods have
to approach one line. Although this method is not robust this restriction.
while significant outliers exist in the data set [3], it is easy
to implement, and with less computation overhead. Here, 3) Sliding-Windows Skew with Lower-Bound Filter: To
four different types of mechanism based on linear regression disassemble the effect caused by outliers, the most effective
are proposed to approach the clock skew. In the following method is to filter them out, so a lower bound approach would
paragraph, LR(Nij) denotes the linear regression calculation be helpful here. For example, huge network delay exists at
for data set Nij, which contains the data from (tis, oi) to red circle part in Fig. 3 (i.e. the outliers); these points mostly
(tjs, oj ). caused by network delay may affect the correctness of skew
estimation enormously. In contrast, the offsets lie in the lower
- Accumulated Skew: For accumulated skew, while pack- part are relative smooth. Hence, the closest skew estimation is
ets sent from the client are received by the server, the server calculated from the lower-bound dataset, as previous research
computes the estimated skew immediately. The estimated skew did [3]. In [3], Moon et al. suggests piecewise minimum
algorithm as a simple and efficient method for estimating clock
skew. Since piecewise minimum can be used to extract the
lower bound of offset, the algorithm is suitable to serve as the
low-bound filter.
For sliding-windows skew with lower-bound filter, the local
minimum offset is picked for every m packets in each sliding
window w. Thus, the amount of sampling data for skew esti-
mation is reduced to w/m. With this lower-bound filter, only
packets with minimal network delay are collected, thereby
minimizing the influence of the network. This lower-bound
skew estimation can be denoted as LR(M in(Nij)), where
M in(Nij) is the data set of local minimum offset between
ith request and jth.
As shown in Fig. 4(c), the skew estimation is much
smoother than the one in Fig. 4(b); the window size w is
529
200, m is 5 in this example, so the total data for one skew (a) Accumulated Skew
estimation is 40. By this method, the clock skew estimation
hence stable and capable for reducing the effect of huge (b) Skew with Sliding-Windows
network delay.
(c) Sliding-Windows Skew with Lower-Bound Filter
- Accumulated Sliding-Windows Skew with Lower-Bound
Filter: Since the local minimum offset is useful to find the (d) Accumulated Sliding-Windows Skew with Lower-Bound Filter
lower-bound skew, we further calculate the accumulated skews Fig. 4. Comparison of each skew estimation by applying linear regression.
with these local minimum dataset and plot the corresponding
skew distribution diagram in Fig. 4(d). We found that this
method can both reduce the effect of huge network delay
and converge rapidly within 20 packets. This skew estimation
can be denoted as LR(M in(N1i)) while receiving ith request
from the client.
As a result, by comparing these four types of skew esti-
mations, the lower-bound filter has the best performance on
filtering out the huge network delay in our experiments.
B. Quick Piecewise Minimum Algorithm
The quick piecewise minimum algorithm is achieved by
separating data into n segments and picking two minimum
offsets from first segment and fourth segment respectively. By
connecting these two minimum offsets, a slope of this line
can be obtained. In our experiment, n is set to 4. As shown
in Fig. 5, the black circles are offsets, the red squares are the
corresponding segments, and the two blue squares are the two
local minimum offsets. Thus, the estimated skew is the slope
of the black line between two blue squares.
As long as the data set is stable enough, the quick piecewise
minimum algorithm can achieve high stability with little
computation. Fig. 6 illustrates the skew estimated by this
method based on the same data with Fig. 3. Since the quick
piecewise minimum algorithm has low computation overhead,
it can be efficiently apply to other cloud service.
C. Jump Point Detection and Elimination
During packet collection period, a jump point of offset
occurs if the client is performing time synchronization with
a time server or roaming between different network providers.
An example of jump point, as shown in Fig. 7, is marked by
red circle. If this jump point phenomenon does not detected
by server, the skew estimation process would be affected by
the mass offset error, reacting in an inaccurate result. Thus,
an approach to detect jump point occurrence is needed while
implementing the skew estimation.
The main idea of this algorithm is to divide sequential times-
tamps into groups, where every pair of consecutive timestamps
in a group produces only a small time difference. Suppose diff
represents the difference between two adjoining offsets; the
diff is denoted as dij, where dij = oj - oi, i = j, i < j.
Also, another threshold k is set to check if the diff is too
large, thereby detecting the jump point. The basic jump point
detection algorithm is as follows:
- Pick the local minimum offsets for every p packets.
- Compute the diff between every pairs of contiguous
local minimum offsets.
- Following detection process is divided into two classes:
530
Offset Jump Point
(Second)
Offset
(Second)
Server Time (second) Server Time (second)
Fig. 5. Implementation of quick piecewise minimum algorithm. Fig. 8. Detection of jump point between two segments.
Fig. 6. Skew estimation by applying quick piecewise minimum algorithm. a) If derived diffs are all positive or all negative:
Denote the median of derived diffs as Med(diff).
Fig. 7. Offset distribution diagram between server and client with a jump If there exists a diff that diff > k Med(diff), a
point. jump point exists inside these p packets.
b) If only part of derived diffs are positive:
If positive diffs are followed by one negative diff
at x, this x is the jump point.
Similarly, negative diffs followed by one positive
diff is processed vice versa.
To prevent the error caused by the jump point, the skew es-
timation is executed before and after the jump point separately
by quick piecewise minimum algorithm. As shown in Fig. 8,
while the jump point is located at jth segment, the first half
estimated skew is calculated from 1th to (j -2)th segment; the
second half is calculated from (j + 1)th to the last segment.
Two segments are omitted from skew estimation because the
jump point may locate at (j - 1)th or jth segment. By this
segmental method, all candidate skews can be calculated.
To represent the moderate value of all skews, we pick the
median of these candidate skews as the estimated skew for
the corresponding client.
A homogeneous case of jump point is a time gap, which
is a period of time of blank packets, as shown in Fig. 9.
A time gap generally arises when user is moving around,
causing the change of network connection, resulting in a period
of disappearance. To detect this kind of jump point, simply
checking xij, where j = i + 1. If xij is larger than twice of
the assigned AJAX packet sending period, it is considered as
a time gap, and is treated as a jump point as well.
V. EXPERIMENTS AND EVALUATION
To evaluate the proposed host identify method, two different
experiments are held in this section. Principal parameters of
our experiments are set as follows.
To obtain high precision offsets, at least 200 AJAX packets
are collected for skew estimation in every experiment. Since
the time resolution of Javascript is 1 ms, at least 1000 seconds
is required to detect a offset up to 1 microsecond resolution.
Since our AJAX script generates a packet for every 5 seconds,
the minimum number of packets required would be 1000/5,
or 200.
531
Fig. 9. Offset distribution diagram between server and client with network processor, running MAC OS X 10.6.8, and the estimated skew
disconnection. is calculated by jump point detection with quick piecewise
minimum algorithm. As shown in table I, we at first connect
TABLE I this laptop to the server via common networks such as LAN,
THE ESTIMATED SKEW OF THE SAME DEVICE UNDER DIFFERENT ADSL, 3G, and Wi-Fi.
NETWORK ENVIRONMENTS According to the first four results, it justifies that clock
skew estimation is relative independent regardless of network
Network type Skew estimation Packets No. of IP addr. access media. It is worthy to notice that the estimated skew is
LAN fluctuated within 2.16 ppm, from -21.08 ppm to -23.24 ppm.
-21.91 ppm 1001 1 Since the clock skew may fluctuate with temperature men-
ADSL -23.24 ppm 207 1 tioned in [5,6,8], these experimental results seems acceptable
3G -22.74 ppm 13322 1 considering the short timestamp collecting time and possible
Wi-Fi -21.48 ppm 5837 1 noise related to network environments.
Tor -21.08 ppm 1400 1
-23.24 ppm 951 1 With the skew fluctuation, it is a trend-off to set a rea-
VM -23.71 ppm 1027 1 sonable tolerance threshold for detecting malicious login. If
-21.79 ppm 9810 1 the threshold is too small, the false positive rate would be
-23.06 ppm 1470 1 unacceptably high, which means that one device might be
-22.53 ppm 15007 55 treated as different device quite often. On the contrary, if the
-23.22 ppm 12922 57 threshold is too large, the false negative rate would raise, which
-22.88 ppm 24120 108 leads to high probability of accepting unregistered devices.
-113.19 ppm 868 1 Thus, after analyzing the table I, we consider a threshold of
-114.22 ppm 1001 1 1 ppm appropriate at the present stage. This threshold derives
-6.40 ppm 1001 1 a probability of misjudgement 7.4%, or (2.16 - 2)/2.16,
-6.83 ppm 890 1 assuming the probability density function of skew fluctuation
is uniformly distributed.
For the jump point detection, the parameter p is set to 20,
and the threshold k is also set to 20. p and k are adjustable Furthermore, the experimental environment in Tor is also
according to different network environments; with parameters listed in table I. In our experimental results, 200 packets is not
of smaller values, the jump point detecting process leads to a enough to calculate a stable estimation. However, the estimated
much sensitive result. skews diverge and are close to the common network cases,
as long as the collecting time is long enough. Finally, for
A. Same Device under Different Networks client hosts running inside a VM, the skews are stable but
unpredictably different from the real one. Kohno et al. [4]
The first experiment is performed with a laptop connecting had indicated that virtual machines do not have constant clock
to our timestamp collection system via multiple network ac- skews, and our experiments also shows the same results. In
cess techniques. Host behind Tor network and virtual machine addition, we found that the estimated skew under a VM is
(VM) are included in this experiment as well for reference. The relatively stable without rebooting it; however, after restarting
purpose of this experiment is to measure the skew variations the VM, the skew randomly changes to another stable value.
of one device under different network environments, thereby As shown in table I, the estimated skews of two clients are
simulating the situations that a device logins the cloud service -113.19 ppm and -114.22 ppm respectively, but their skews
from various kind of networks. In this experiment, the device changed to -6.4 ppm and -6.83 ppm after rebooting the system.
is an Apple MacBook with a 2.4 GHz Intel Core 2 Duo
B. Skew Distribution of Different Devices
To study the distinguishable property of clock skew, we
collected clock skews of 100 devices estimated by the same
timestamp collection server. The results vary from 67 ppm
to -499 ppm, and only the most close 90 skews, sorted by
their values, are illustrated in Fig. 10. Each cross in this
figure represents one device and the y coordinate stands for
its clock skew. Each clock skew is bound by an interval of 1
ppm range. As shown in this figure, many devices have their
intervals overlapped with others, which means that there exists
not negligible probability a user may pass the identification test
with unregistered devices. The worst case happens at device
number 20, whose interval overlaps with other 8 devices.
Therefore, with the threshold of 1 ppm, the maximum false
negative rate is currently 8%. We believe that further analysis
532
Estimated skew valuex 10-5 [3] S. Moon, P. Skelly, and D. Towsley, "Estimation and removal of clock
skew from network delay measurements," in INFOCOM '99. Eighteenth
0 Annual Joint Conference of the IEEE Computer and Communications
Societies. Proceedings. IEEE, vol. 1, mar 1999, pp. 227 234 vol.1.
-1
[4] T. Kohno, A. Broido, and K. Claffy, "Remote physical device finger-
-2 printing," in IEEE Transactions on Dependable and Secure Computing,
vol. 2, no. 2, April-June 2005, pp. 93108.
-3
[5] S. J. Murdoch, "Hot or not: revealing hidden services by their clock
-4 skew," in CCS '06: Proceedings of the 13th ACM Conference on
Computer and Communications Security, New York, NY, USA, 2006,
-5 pp. 2736.
10 20 30 40 50 60 70 80 90 100
[6] S. Zander and S. J. Murdoch, "An improved clock-skew measurement
Device number technique for revealing hidden services," in Proceedings of the 17th
conference on Security symposium. Berkeley, CA, USA: USENIX
Fig. 10. The sorted skew of estimated devices Association, 2008, pp. 211225.
on skew estimation of different network environments would [7] D.-J. Huang, W.-C. Teng, C.-Y. Wang, H.-Y. Huang, and J. M. Heller-
help reducing the tolerance threshold, thus reduce both false stein, "Clock skew based node identification in wireless sensor net-
positive rate and false negative rate. works," in IEEE Global Communications Conference, 2008, pp. 1877
1881.
VI. CONCLUSIONS AND FUTURE WORK
This paper addressed the client device identification issue [8] M. Uddin and C. Castelluccia, "Toward clock skew based wireless sensor
in cloud environment, and proposed a clock skew based node services," in Wireless Internet Conference (WICON), 2010 The 5th
fingerprinting technique and a practical scenario. Client device Annual ICST, March 2010, pp. 19.
identification strengthen the account safety, and the proposed
scenario suggests a potential secondary authentication ap- [9] Top threats to cloud computing v1.0. [Online]. Available:
proach without users' awareness most of the time. Several https://cloudsecurityalliance.org/topthreats/csathreats.v1.0.pdf
classical methods are introduced and implemented to estimate
the skews of client devices, and the treatment of jump points, [10] J. Hall, M. Barbeau, and E. Kranakis, "Detection of transient in radio
usually caused by switching network or temporary discon- frequency fingerprinting using signal phase," in Proceedings of IASTED
nection, is also discussed at the first time. To examine the International Conference on Wireless and Optical Communications
effectiveness of clock skew, we implemented a web based (WOC '03), 2003.
skew estimation platform and conducted two experiments. The
initial study includes over 100 client devices from 5 different [11] R. M. Gerdes, T. E. Daniels, M. Mina, and S. F. Russell, "Device iden-
network media. The experiment results showed that the false tification via analog signal fingerprinting: A matched filter approach,"
positive rate and the false negative rate, in the worst case, in Proceedings of the 2006 Network and Distributed System Security
are both no more than 8% when the tolerance threshold is Symposium (NDSS '06), February 2006.
set to 1 ppm. With the potential of clock skew fingerprint
be revealed, the further work would include improving the [12] K. Bonne Rasmussen and S. Capkun, "Implications of radio fingerprint-
precision of skew estimation utilizing linear programming ing on the security of sensor networks," Sept. 2007, pp. 331340.
method and accumulating the knowledge of skew fluctuation
in different network media. [13] S. Jana and S. K. Kasera, "On fast and accurate detection of unauthorized
wireless access points using clock skews," in MobiCom '08: Proceedings
ACKNOWLEDGMENT of the 14th ACM international conference on Mobile computing and
This work was supported under the National Science Coun- networking, 2008, pp. 104115.
cil Grants 100-2218-E-011-008.
REFERENCES
[1] P. Eckersley, "How unique is your web browser?" in Privacy Enhancing
Technologies. Springer Berlin / Heidelberg, 2010, vol. 6205, pp. 118.
[2] V. Paxson, "On calibrating measurements of packet transit times,"
in Proceedings of the 1998 ACM SIGMETRICS joint international
conference on Measurement and modeling of computer systems, ser.
SIGMETRICS '98/PERFORMANCE '98. New York, NY, USA: ACM,
1998, pp. 1121.
533