← 资料库索引 ← 个人博客 原始链接 ↗ 🔍
个人博客

p0f 与被动操作系统指纹识别:网络层如何暴露你是谁 原文标题:p0f and passive OS fingerprinting: how the network layer gives you away

发表时间:2026-02-16采集时间:2026-10-09 10:58:20来源:blog.crawlex.net原文语言:en状态:完整

内容概要总结

本文是 crawlex.net 关于 p0f 与被动操作系统指纹识别的深度技术文章(2026 年 2 月)。文章解释了为何 TCP SYN 包(由内核而非应用构建)能在 TLS 和应用数据之前泄露操作系统身份:初始 TTL、通告窗口、TCP 选项集合与顺序。它梳理了 p0f 的历史(Zalewski 2000 年 Bugtraq 发布 v1.0,提出"63 位签名"理念;v2 时代;2012 年 v3 重写并新增 HTTP 模块),并逐字段拆解 p0f v3 的签名语法 ver:ittl:olen:mss:wsize,scale:olayout:quirks:pclass,给出 Linux 3.11+(*:64:0:*:mss*20,10:mss,sok,ts,nop,ws:df,id+:0)、Windows 7/8(*:128:0:*:8192,0:mss,nop,nop,sok:df,id+:0)、macOS 10.x 的真实签名示例,以及 quirks 字段(df、id+、ecn、seq-、ack+、ts2+、exws 等)的完整清单。文章还介绍 TCP 时间戳推算运行时间、HTTP 模块的头部顺序指纹、NAT/代理检测原因码,重点阐述跨层不匹配(TCP 说 Linux TTL 64、User-Agent 说 Windows)这一反代理检测核心信号,以及 Cloudflare 把 p0f 签名编译成 BPF 做线速 SYN 洪泛分流的案例。最后评估该技术在 2026 年的有效性(家族级信号仍有效、中间盒与云出口侵蚀它、负载已加密无法触及)。

翻译内容

原文内容(English)

⚠ 说明:结构对齐(2026-10):按 content_en 校正译文标题层级(节标题 # / 子节 ##),并补回原文「Further reading」下 3 个子标题。

打开一个到服务器的 TCP 连接,你就已经告诉它你运行的是什么操作系统——在你发送任何请求之前,在 TLS 之前,在任何一个字节的应用数据之前。启动握手的 SYN 包是由你的内核而非你的应用程序构建的,而内核之间会以细小而持久的方式存在差异。初始 TTL。通告的窗口。出现哪些 TCP 选项,以及以什么顺序。这些都不是秘密,都未加密,也都不是普通应用程序能够改变的。一个被动的观察者从线路上读取它,把你归档到 Linux、Windows 或 macOS 之下,并对版本有一个合理的猜测。

使这一切成为一门学科的工具是 p0f。它不发送探测。它不触碰你的机器。它坐在一条链路上,观察那些本来就要流动的包,并将它们与一个协议栈怪癖(stack quirks)数据库进行匹配。本文专门讲述 p0f:它从何而来,它从一个 SYN 中读出什么,它的签名语法如何编码一个 TCP/IP 协议栈,以及为什么这门技术在 2026 年仍然有效,即便该工具最后一次真正的发布早于它现在所看到的大部分流量。

路线图:首先是主动与被动的区别,以及为什么被动性正是关键所在。然后是历史,从 2000 年的一篇 Bugtraq 帖子到 v3 的重写。然后是逐字段的机制:一个 SYN 揭示了什么,以及 p0f 的签名格式如何捕捉它。然后是 SYN 之外的部分:用于运行时间的 TCP 时间戳、HTTP 模块、NAT 和代理检测。最后是对该技术当今处境的诚实评估:什么在侵蚀它,什么不会。

被动与主动,以及为什么它重要

有两种从网络行为了解远程主机运行什么操作系统的方法。你可以戳它,或者你可以听。

主动指纹识别会戳。Nmap 是典型例子:它精心构造一批不寻常的包,把它们发往开放和关闭的端口,并读取响应。向关闭端口发一个 SYN,一个带有奇怪标志组合的包,一个故意使用异常窗口的探测。协议栈对畸形或边界情况输入的反应各不相同,而这些差异具有诊断价值。主动探测是精确的,因为探测者精确控制发送什么,并能反复重放实验直到答案毫无歧义。cgsecurity 的相关文章指出,当被动分析无法得出结论时,主动方法给出"更高的精确度"(plus de précision),这就是权衡所在。代价是你产生了流量。不寻常的流量。会被 IDS 记录、可能被防火墙丢弃、并告诉目标你对它感兴趣的流量。

被动指纹识别则只是听。它从不发送任何东西。它分析目标无论如何都要传输的包:它想要发起的连接所打开的那个 SYN,它对你发起的连接所回应的 SYN+ACK。由于它不向线路增加任何东西,它无法被被指纹识别的主机检测到,并且能直接穿透包过滤防火墙和 NAT。p0f 最初的公告正是这样宣称的:"包过滤防火墙、网络地址转换等等对这门技术是透明的,因此你能够获得防火墙背后系统的信息。"你通过阅读一封它已经发出的信件来了解一台主机。

正是这一特性使被动指纹识别对防御者以及任何运行大规模服务的人都很有趣。主机无法知道自己正在被画像,因此无法做出适应。一台 Web 服务器可以连续、免费地指纹识别每一个连接它的客户端,作为它本来就在处理的流量的副产品。Cloudflare 正是这样做的,用来分流 SYN 洪泛,我们稍后会回到这一点。关心自己正在被指纹识别的读者应当明白:网络层会给出一个应用层永远无权置喙的答案。

2000 年的一篇 Bugtraq 帖子与通往 v3 的漫漫长路

p0f 是仍在实际使用中的较老工具之一。Michal Zalewski(署名 lcamtuf)于 2000 年 6 月 10 日将 1.0 版发布到 Bugtraq 邮件列表。公告描述了"p0f - 被动操作系统指纹识别工具",并用一句至今仍然成立的话道出了核心理念:"初始 TTL、窗口大小、最大分段大小、不分片标志、sackOK 选项……nop 选项和窗口缩放选项组合在一起,为每个系统给出唯一的、63 位的签名。"该工具所需的全部只是"至少一个初始化 TCP 连接到你机器或网络的 SYN 包"。二十六年之后,现代签名数据库中的字段名仍然可以辨认出与当年相同。

该工具经历了一个漫长的 2 版时代。到 2004 年 7 月,当前版本是 p0f 2.0.4,它为普通 SYN 之外的更多包类型增加了指纹识别,包括 SYN+ACK 和 RST+ACK,还增加了针对伪装(masquerading)和 IP 共享的启发式规则。那个时期 LWN 的报道将 Zalewski 与 William Stearns 及其他贡献者一并归功,并指出该项目至今仍携带的 LGPL 2.1 许可。v2 在将近十年里都是主力,互联网上流传的许多指纹数据库(包括 CERT 的 NetSA 套件所附带的那些)都源自这一谱系。

3 版是一次彻底的重写,版权 2012 年。它就是今天在用的版本,通过非官方 GitHub 镜像分发,并被各大 Linux 发行版打包。v3 不只是移植。README 描述了一个会推理"IPv4 和 IPv6 头部、TCP 头部、TCP 握手的动态过程,以及应用层负载的内容"的工具。最后那句是重大的新增。v3 长出了一个 HTTP 模块,因此它不仅能指纹识别一条连接底下的内核,还能指纹识别骑在其上的浏览器或服务器软件。签名语法围绕命名字段和显式怪癖被重新设计,这就是我们下面要剖析的格式。

一个 SYN 包实际携带了什么

要理解指纹,你必须理解包。TCP SYN 是三次握手的第一条消息。它不携带应用数据,但它携带了关于连接将如何行为的、令人意外的发送方意图,而这些参数中的大多数由内核的网络协议栈根据编译进去的或 sysctl 调优的默认值设置。

从 IP 头部开始。生存时间(Time To Live,TTL)字段是一个 8 位计数器,每个路由器在包通过时将其减一。它的职责是防环,但它的初始值是一个协议栈默认值,而协议栈只从一个很小的集合中挑选。几乎所有操作系统都把 TTL 初始为 64、128 或 255。Linux 和 BSD 系(包括 macOS)使用 64。Windows 使用 128。某些网络设备使用 255。由于你是在包穿过若干路由器跳之后才观察到它,你看到的是一个被递减的值,但你可以恢复初始值:把观测到的 TTL 向上取整到 {64, 128, 255} 中的下一个成员。cgsecurity 的文章给出了教科书式的例子,观测到的 TTL 为 62 意味着初始值为 64,中间有两跳。观测值与初始值之差也是对网络距离的一个免费估计。

不分片(Don't Fragment,DF)位位于同一个 IP 头部。大多数现代协议栈会设置它,因为它们依赖路径 MTU 发现而非分片,但它是否被设置仍然是一个被记录的签名字节,而 DF 与 IP 标识字段行为的组合会把协议栈进一步区分开。

然后是 TCP 头部,真正的区分力所在。有三件事最重要。

通告的窗口大小。 这是发送方在必须收到确认之前愿意接收多少字节,它深度依赖于协议栈。关键在于,许多协议栈并不通告一个平坦的常量。它们把窗口通告为最大分段大小(MSS)或 MTU 的倍数。指纹在于这个关系,而不仅仅是那个数字。一个通告其 MSS 的十或二十倍的 Linux 内核,会产生一个随路径 MSS 变化而变化的窗口,而乘数保持不变,这个乘数才是识别它的东西。

最大分段大小(MSS),作为一个 TCP 选项携带,告诉对等端发送方希望在一个分段中承载的最大负载。它源自链路 MTU,因此在一条正常的 1500 字节以太网路径上它落在 1460,但在隧道、PPPoE 链路以及任何降低 MTU 的东西上它会移动。p0f 记录 MSS,部分是为了评估窗口关系,部分是因为 MSS 本身,结合它所暗示的 MTU,指向链路类型。

以及 TCP 选项,包括出现哪些选项以及内核把它们排布的顺序。一个协议栈可能发出 MSS、然后一个 SACK-permitted 选项、然后一个时间戳、然后一个用于对齐的 NOP、然后窗口缩放。另一个发出 MSS、NOP、窗口缩放、NOP、NOP、SACK-permitted,且完全没有时间戳。这个集合和顺序无法通过任何正常接口配置;它们被烘焙进内核的 TCP 输出路径。这使得选项布局成为整个包中最强的区分因素之一。

签名语法

p0f v3 把一个 TCP 协议栈编码为一个以冒号分隔的字符串。阅读它是理解该工具究竟以什么为键的最快方式。SYN 的格式为:

ver:ittl:olen:mss:wsize,scale:olayout:quirks:pclass

每个字段对应线路上的某个东西。ver 是 IP 版本:4、6 或 *(两者皆可)。ittl 是推断出的初始 TTL,通过把观测值取整恢复得到,写成规范的 64、128 或 255。olen 是 IPv4 选项或 IPv6 扩展头的长度,通常为零。mss 是来自 TCP 选项的最大分段大小,或 *(当它变化时)。wsize,scale 是通告的窗口和窗口缩放因子,窗口常常相对于 MSS 书写:mss*20 表示 MSS 的二十倍,而不是字面的字节数。olayout 是 TCP 选项按协议栈发出的确切顺序排列的逗号分隔列表,使用短标记:mss、sok 表示 SACK-permitted、ts 表示时间戳、nop 表示单字节填充、ws 表示窗口缩放、eol+n 表示选项结束(end-of-options)后跟 n 个填充字节。quirks 是协议栈怪癖的逗号分隔列表。pclass 把负载分类为 0(空)、+(非空)或 *(任意);一个正常的 SYN 没有负载。

一个 label(标签)把一个签名绑定到一个人类可读的身份。取自随附的数据库:

label = s:unix:Linux:3.11 and newer

sig = *:64:0:*:mss*20,10:mss,sok,ts,nop,ws:df,id+:0

从左到右读它。任一 IP 版本。初始 TTL 64,因此是 Unix 系协议栈。没有 IP 选项。任意 MSS。窗口是 MSS 的二十倍,缩放因子为 10。选项按 MSS、SACK-permitted、时间戳、NOP、窗口缩放的顺序。怪癖:DF 被设置,且尽管 DF 被设置,IP ID 字段仍为非零。空负载。那个字符串就是 Linux 3.11+ 内核的 TCP 人格被写下来。

label = s:win:Windows:7 or 8

sig = *:128:0:*:8192,0:mss,nop,nop,sok:df,id+:0

初始 TTL 128,Windows 的特征。一个平坦的 8192 窗口,没有缩放。选项按 MSS、NOP、NOP、SACK-permitted 的顺序,且明显没有时间戳。然后是 macOS:

label = s:unix:Mac OS X:10.x

sig = *:64:0:*:65535,1:mss,nop,ws,nop,nop,ts,sok,eol+1:df,id+:0

像其 BSD 血统一样 TTL 为 64,一个平坦的 65535 窗口、缩放因子 1,以及一个长长的、独特的选项串,以一个显式的选项结束标记加一字节填充收尾。三个操作系统,三个签名,全都从一个不携带应用数据、不加密的包中读出。

OS 标签本身有结构。一个 v3 标签携带一个类型、一个类、一个名称和一个变体(flavor)。类是宽泛的家族:unix、win、cisco。名称是具体的操作系统,Linux 或 Windows。变体是限定词,即版本范围。这就是为什么 p0f 能以不同的分辨率作答:当只有类匹配时是"某个 Windows",当完整签名对得上时是"Windows 7 或 8"。

quirks 字段

quirks 列表是 p0f 捕捉协议栈表现出的那些轻微违规行为的地方,那些与其说是参数不如说是特征的东西。README 列举了它们。在 IP 侧:df 表示不分片标志,id+ 表示 DF 被设置而 IP ID 仍为非零,id- 表示 DF 被清除而 ID 为零,ecn 表示显式拥塞通知支持,0+ 表示一个规范要求必须为零的字段中出现了非零值,flow 表示非零的 IPv6 流标签。在 TCP 侧:seq- 表示序列号为零,ack+ 表示 ACK 标志未被设置时确认号却非零,ack- 是相反情况,uptr+ 表示没有 URG 标志却出现非零紧急指针,此外还有 push 和 urgent 标志的异常。在时间戳上:ts1- 表示自身时间戳为零,ts2+ 表示一个 SYN 上出现了非零的对等端时间戳(这本不该发生)。以及兜底项:opt+ 表示选项之后出现尾随的非零数据,exws 表示超过 14 的过大窗口缩放值,bad 表示解析器无法理解的畸形选项。

这些怪癖大多描述的是任何应用程序都无法产生、任何正常用户也永远不会注意到的行为。它们存在,是因为 TCP 协议栈是由不同的人在不同的时间针对一份带有边角的规范编写的,而这些边角被以不同方式处理。这正是它们成为持久标识符的原因。一个"必须为零"却不为零的字段,告诉你关于发出该包的代码路径的某些具体信息。

SYN 之外:时间戳、运行时间与 HTTP 模块

p0f 读取的不止第一个包。它两个更有趣的技巧来自随时间观察连接,以及观察 TCP 之上的层。

TCP 时间戳,当存在时,让 p0f 能够估计远程主机的运行时间。时间戳选项携带一个由时钟驱动的值,该时钟以协议栈特定的频率滴答。通过观察时间戳在多个包之间的推进并知道滴答频率,p0f 可以向后外推计数器何时为零,那大致就是协议栈启动的时间。README 指出,该工具需要观察到"至少约 25 毫秒的合格流量"才能锁定其推进过程。结果是对任何启用了时间戳的协议栈主机(默认情况下大多数 Linux 和 BSD 系统)的一个免费运行时间读数。这生动地展示了被动观察者能从一个出于完全无关原因(往返时间测量与防止序列号回绕的 PAWS 保护)而存在的参数中推断出多少东西。

HTTP 模块是 v3 的头号新增。它把同样的哲学应用上一层。它不是解析请求说了什么,而是看请求是如何构造的,基于结构比内容更难伪造的理论。一个 HTTP 签名具有以下形式:

ver 是 HTTP 版本,1.0 为 0,1.1 为 1。horder 是有序的头部列表,可选的 name=[value] 对特定头部值做子串匹配。habsent 列出不得出现的头部。expsw 是 User-Agent 或 Server 头部中预期的子串,用于捕捉那些谎报自身的软件。其洞见在于:浏览器以特征的顺序发送其头部,并包含或省略一个特征性的集合,而这个顺序是 HTTP 客户端实现的一个属性,而非所获取页面的属性。一个声称自己是某浏览器、却按另一种浏览器的顺序排列头部的客户端,已经暴露了自己。这与驱动现代头部顺序和大小写指纹以及 Accept 头部三元组签名的逻辑相同,而 p0f 在 2012 年就在 HTTP 层做这件事了。

抓住一个代理:操作系统不匹配信号

被动指纹识别在反滥用语境中最有用的单一功能,是抓住一个各层彼此不一致的网络。p0f 为此有明确的机制。

当 p0f 看到一台主机的签名以一种看起来系统化而非随机的方式发生变化时,它会标记它。README 列出了它附加的原因代码:os_sig 表示操作系统签名本身变化,sig_diff 表示协议层变化,tstamp 表示时间戳不一致,ttl 表示 TTL 变化,port 表示本不该发生的源端口递减,mtu 表示 MTU 偏移。从一个表观主机连续出现这些,就是 NAT、代理或地址共享的指纹——多台真实机器藏在一个 IP 背后,各有自己的协议栈人格。

这更锐利的版本是跨层不匹配,而它正是网络指纹识别在 2026 年对机器人检测仍然重要的原因。设想一个流经代理的请求。代理的内核终结你的 TCP 连接,并向目标打开一条新的连接。目标因此看到的是代理的 SYN,带有代理的 TTL、窗口和选项布局,而不是你的。所以 TCP 协议栈所暗示的操作系统是代理的操作系统。现在假设骑在其中的 HTTP 请求通过其 User-Agent 声称自己是一个 Windows 浏览器,而发出该 SYN 的代理运行的是 Linux。TCP 指纹说是 TTL 64,Linux。User-Agent 说是 Windows。对于一台机器,这两者不可能同时为真。这个矛盾就是检测。正如 pydoll 关于网络指纹识别的文章直白地说的那样,"User-Agent 说是 Windows(TTL 128),但 TCP 指纹显示是 Linux(TTL 64)"就是暴露代理或伪造代理字符串的特征。

这就是为什么 TCP/IP 指纹识别位于现代反机器人技术栈的底层,而不是被它所取代。一个抓取器可以完美伪造其 User-Agent。它可以用 uTLS 模仿浏览器的 TLS ClientHello。它可以伪造头部顺序。但打开连接的那个 SYN 出自实际发送它的任何内核,而伪造它需要控制网络协议栈本身,而不仅仅是应用程序。这种不匹配检测远远超越了 p0f 自己的数据库,并直接连接到通过操作系统不匹配来检测代理这个更广泛的问题。

一次指纹匹配实际是如何发生的

匹配的机制比数据库让它们看起来的要简单。p0f 从观测到的包中提取字段,构建签名字符串,并针对其已加载的指纹寻找最佳匹配。一次典型的安装会从 p0f.fp 文件加载大约 320 个 SYN 签名。匹配不是纯粹的相等;数据库中的通配符(* 用于 IP 版本、MSS、负载类别)意味着单个签名可以覆盖一系列真实包,而窗口可以相对于 MSS 表达,因此它能跨具有不同 MTU 的路径匹配。

当多个签名都可能匹配时,p0f 在它能确信的最粗粒度上做出判断。一个 TTL 为 64、但其选项布局不匹配任何精确条目的包,可能仍会凭借 TTL 和宽泛的选项形态被判定为"泛 Linux"。带有类/名称/变体字段的标签结构,正是为了让该工具能给出一个有用的部分答案,而不是一个无用的空值。这也解释了为什么这个数据库在一个方向上优雅地老去,在另一个方向上则很差:一个未知的新 Linux 在类一级看起来仍像 Linux,即使没有变体匹配;但一个具有不寻常选项布局的真正新颖的协议栈可能完全落入"未知"。

数据库是柔软的软肋。p0f 的 v3 指纹集是 2012 年时代的,加上一些社区补丁。此后发布的操作系统其协议栈是原始数据库从未见过的。p0f 读取的字段没有改变,TTL 仍是 TTL,窗口仍是窗口,但现代内核发出的具体值组合可能没有带标签的条目。在实践中,这意味着一个全新的 p0f 面对 2026 年的流量,正确识别操作系统家族的频率远高于它精确锁定具体版本,因为家族级的特征(TTL 64 对 128、时间戳存在与否、粗略的选项顺序)在许多内核发布之间是稳定的,而精细细节则会漂移。

Cloudflare 的 BPF 编译器,或者说线速指纹识别

p0f 的签名比 p0f 程序本身活得更久的一个绝佳例证,是 Cloudflare 用它们做的事。在 2016 年 8 月的一篇工程博客中,Cloudflare 描述了把 p0f 的签名格式编译成 Berkeley Packet Filter(BPF)字节码。其动机是 SYN 洪泛防御:在攻击期间,他们想"对攻击包进行速率限制,并实际上优先处理其他(希望是合法的)包"。真实操作系统产生的 SYN 匹配已知的 p0f 签名;许多洪泛工具产生的 SYN 则不匹配,或者匹配某个特定攻击工具的签名。

他们没有在包路径上运行 p0f 守护进程,而是采用人类可读的签名语法,构建了一个编译器来生成 BPF,使分类能在内核的包过滤器中以线速运行,位于 iptables 内部。该博客用 p0f 自己的格式给出了实际示例:一个 Linux SYN 为 4:64:0:*:mss*10,6:mss,sok,ts,nop,ws:df,id+:0,一个 Windows 7 SYN 为 4:128:0:*:8192,8:mss,nop,ws,nop,nop,sok:df,id+:0,以及一个 hping3 攻击包,它用一个稀疏的签名和一个 ack+ 怪癖暴露了其合成来源。这种签名语言比定义它的工具活得更久,这是对那个抽象有多好的一个公平衡量。Cloudflare 编译成 BPF 的同一套字段,正是为 DataDome 的 HTTP/2 和网络指纹识别等厂商系统的网络层一侧提供输入的东西。

这门技术在 2026 年的处境

从 SYN 出发的被动操作系统指纹识别比当今运行的大多数生产 TCP 协议栈都要老,而且它不会消失,但它有值得直说的局限。

仍然有效的: 家族级信号。TTL 为 64 对 128 仍然干净地把 Unix 血统的世界与 Windows 分开,而且没有正常应用程序能改变它。选项顺序和时间戳选项的存在与否仍然区分主要协议栈。窗口对 MSS 的关系仍然成立。最重要的是,跨层不匹配检查——网络指纹与 User-Agent 不一致——现在比 2012 年更有用,因为现代互联网充斥着代理、VPN 和 CGNAT,而检测某人是否身处其中之一本身就很有价值。

侵蚀它的: 中间盒与规范化。NAT、VPN 集中器、负载均衡器和流量规范化器会重写 TTL 或重建 TCP 选项,这要么改变观测到的签名,要么把许多真实主机涂抹成一个。一台身处企业 VPN 背后的主机可能被指纹识别为 VPN 设备。云出口的兴起意味着现在有巨大份额的流量源自 Linux 虚拟机管理程序上屈指可数的几种协议栈类型,这压缩了这门技术赖以生存的多样性。而 v3 数据库的年迈意味着即使在家族级识别仍然成立的地方,版本级精度也已经衰减。

它从未触及的: 负载。被动性的全部要点在于 p0f 读取的是线上已有的东西,而线上越来越多的东西是加密的。p0f 的 HTTP 模块看不到 TLS 连接内部的任何东西。这就是为什么应用指纹识别的重心转移到了仅存的明文握手——TLS ClientHello——以及 HTTP/2 帧的结构上,这两者在正文加密时仍以明文泄露。p0f 所读取的网络层位于这一切之下的一层,而它之所以保持可读,恰恰是因为 TTL、窗口大小和选项顺序是出于必然而非选择地以明文传输。

p0f 持久的教训是架构性的,而非关于某个工具。身份在应用程序无法控制的那一层泄露。一个程序可以对它写下的所有东西撒谎——它的 User-Agent、它的头部、它声称的操作系统。它无法轻易对 SYN 撒谎,因为 SYN 属于内核。在一篇 Bugtraq 帖子宣称每个系统都有一个 63 位签名二十六年之后,网络上最廉价、最不可检测的信号仍然是那个应用程序从未有机会触碰的信号。

来源与延伸阅读

  • Zalewski, M. (2000),p0f - passive os fingerprinting tool(Bugtraq 公告)——最初的 v1.0 发布帖,阐述了 63 位 SYN 签名理念与防火墙透明性主张。
  • Zalewski, M. (2012),p0f v3 项目页——作者本人对 v3 重写、其范围以及它从 IP/TCP/HTTP 读取什么的描述。
  • p0f 项目 (2012),p0f v3 README——关于签名语法、完整怪癖列表、运行时间估计和 NAT 原因代码的权威参考。
  • p0f 项目,p0f.fp 指纹数据库——随附的签名文件,含分节结构和 Linux、Windows、Mac OS X 的真实带标签签名。
  • p0f 项目,p0f 非官方 git 仓库——受维护的源码镜像,LGPL 2.1,含文档和 BPF 辅助材料。
  • LWN.net (2004),p0f, the Passive OS Fingerprinter——关于 v2.0.4 时代、作者身份以及主动与被动的区别的同时期报道。
  • Cloudflare (2016),Introducing the p0f BPF compiler——Cloudflare 如何把 p0f 签名编译成 BPF 用于 SYN 洪泛分流,含实际示例签名。
  • CGSecurity,OS fingerprinting——关于主动与被动方法,以及如何通过把观测值取整到 64/128/255 来恢复初始 TTL 的背景。
  • Pydoll docs (2025),Network fingerprinting deep dive——现代 Linux/Windows/macOS TCP 签名值,以及代理操作系统不匹配的检测逻辑。
  • CERT NetSA,p0f fingerprint database snapshots——归档的社区指纹文件,展示数据库如何被版本化和分发。

延伸阅读

TCP/IP 协议栈指纹识别:TTL、窗口大小和 MSS 作为操作系统身份

追溯一个 SYN 包中的初始 TTL、TCP 窗口大小、MSS 以及 TCP 选项顺序如何识别发送方操作系统,以及为什么这个身份是由内核而非浏览器设置的。
2026 年 2 月 17 日(周二)· 22 分钟阅读

跨操作系统的 TCP 时间戳与窗口缩放指纹

追溯一个 SYN 包中 TCP 选项的顺序、窗口缩放移位计数、SACK-permitted 标志、NOP 填充和时间戳时钟如何识别操作系统,以及每连接随机化如何改变了时间戳所泄露的东西。
2026 年 2 月 15 日(周日)· 22 分钟阅读

MTU、路径 MTU 发现,以及隐藏在包大小中的指纹

隧道和 VPN 如何移动 MTU 和 MSS,为什么一个 SYN 包中的非标准 MSS 暴露了一条被封装的路径,以及路径 MTU 发现行为如何把一个包大小值变成一个信号。
2026 年 2 月 13 日(周五)· 20 分钟阅读

Open a TCP connection to a server and you have already told it which operating system you are running, before you send a request, before TLS, before a single byte of application data. The SYN packet that starts the handshake is built by your kernel, not your application, and kernels disagree with each other in small, durable ways. The initial TTL. The advertised window. Which TCP options appear, and in what order. None of it is secret, none of it is encrypted, and none of it is something a normal application can change. A passive observer reads it off the wire and files you under Linux, or Windows, or macOS, with a fair guess at the version.

The tool that made this a discipline is p0f. It does not send a probe. It does not touch your machine. It sits on a link, watches packets that were going to flow anyway, and matches them against a database of stack quirks. This piece is about p0f specifically: where it came from, what it reads out of a SYN, how its signature grammar encodes a TCP/IP stack, and why the technique still works in 2026 even though the tool’s last real release predates most of the traffic it now sees.

The road map: first the distinction between passive and active fingerprinting, and why passivity is the whole point. Then the history, from a 2000 Bugtraq post to the v3 rewrite. Then the mechanics, field by field, of what a SYN reveals and how p0f’s signature format captures it. Then the parts beyond the SYN, TCP timestamps for uptime, the HTTP module, NAT and proxy detection. And finally the honest assessment of where the technique stands today, what erodes it, and what does not.

Passive versus active, and why it matters

There are two ways to learn what operating system a remote host runs from its network behavior. You can poke it, or you can listen.

Active fingerprinting pokes. Nmap is the canonical example: it crafts a battery of unusual packets, sends them at open and closed ports, and reads the responses. A SYN to a closed port, a packet with strange flag combinations, a probe with a deliberately odd window. Stacks respond to malformed or edge-case input differently, and those differences are diagnostic. Active probing is precise, because the prober controls exactly what gets sent and can replay the experiment until the answer is unambiguous. The cgsecurity write-up on the subject notes that active methods give “plus de précision” when passive analysis is inconclusive, and that is the trade. The cost is that you generate traffic. Unusual traffic. Traffic that an IDS logs, that a firewall may drop, and that tells the target you are interested in it.

Passive fingerprinting listens. It never sends anything. It analyzes packets the target was going to transmit regardless: the SYN that opens a connection it wanted to make, the SYN+ACK with which it answers a connection you made. Because it adds nothing to the wire, it cannot be detected by the host being fingerprinted, and it works straight through packet-filtering firewalls and NAT. The original p0f announcement made exactly this claim: “packet filtering firewalls, network address translation and so on are transparent to this technique, so you’re able to obtain information about systems behind the firewall.” You learn about a host by reading mail it already sent.

That property is what makes passive fingerprinting interesting to defenders and to anyone running a service at scale. The host cannot tell it is being profiled, so it cannot adapt. A web server can fingerprint every client that connects to it, continuously, for free, as a byproduct of traffic it was already handling. Cloudflare does precisely this to triage SYN floods, which we will come back to. The reader who cares about being fingerprinted should understand that the network layer gives an answer the application layer never gets a vote on.

A 2000 Bugtraq post and the long road to v3

p0f is one of the older tools still in working use. Michal Zalewski, who signs his work lcamtuf, posted version 1.0 to the Bugtraq mailing list on 10 June 2000. The announcement described “p0f - passive OS fingerprinting utility” and laid out the core idea in a sentence that has aged well: the combination of “initial TTL, window size, maximum segment size, don’t fragment flag, sackOK option … nop option and window scaling option combined together gives unique, 63-bit signature for every system.” All the tool needed was “at least one SYN packet initializing TCP connection to your machine or network.” Twenty-six years later the field names in the modern signature database are recognizably the same.

The tool went through a long version 2 era. By July 2004 the current release was p0f 2.0.4, which added fingerprinting for more packet types beyond the plain SYN, including SYN+ACK and RST+ACK, plus heuristics for masquerading and IP sharing. The LWN write-up from that period credits Zalewski along with William Stearns and other contributors, and notes the LGPL 2.1 license that the project still carries. v2 was the workhorse for the better part of a decade, and a lot of the fingerprint databases floating around the internet, including ones shipped by CERT’s NetSA suite, descend from that lineage.

Version 3 is a clean rewrite, copyright 2012. It is the version in use today, distributed through the unofficial GitHub mirror and packaged by every major Linux distribution. v3 is more than a port. The README describes a tool that reasons about “IPv4 and IPv6 headers, TCP headers, the dynamics of the TCP handshake, and the contents of application-level payloads.” That last clause is the big addition. v3 grew an HTTP module, so it can fingerprint not just the kernel underneath a connection but the browser or server software riding on top of it. The signature grammar was redesigned around named fields and explicit quirks, which is the format we will dissect below.

What a SYN packet actually carries

To understand the fingerprint you have to understand the packet. A TCP SYN is the first message of the three-way handshake. It carries no application data, but it carries a surprising amount of the sender’s intent about how the connection will behave, and most of those parameters are set by the kernel’s network stack from compiled-in or sysctl-tuned defaults.

Start with the IP header. The Time To Live field is an eight-bit counter that each router decrements by one as the packet passes. Its job is loop prevention, but its initial value is a stack default, and stacks pick from a tiny set. Almost every operating system starts TTL at 64, 128, or 255. Linux and the BSDs (including macOS) use 64. Windows uses 128. Some network gear uses 255. Because you observe the packet after it has crossed some number of router hops, you see a decremented value, but you can recover the initial: round the observed TTL up to the next member of {64, 128, 255}. The cgsecurity article gives the textbook example, an observed TTL of 62 implies an initial of 64, with two hops in between. The difference between observed and initial is also a free estimate of network distance.

The Don’t Fragment bit lives in the same IP header. Most modern stacks set it, because they rely on path MTU discovery rather than fragmentation, but whether it is set is still a recorded signature bit, and combinations of DF with IP identification field behavior split stacks further.

Then the TCP header, where the real discrimination lives. Three things matter most.

The advertised window size. This is how many bytes the sender is willing to receive before it must get an acknowledgment, and it is deeply stack-dependent. Crucially, many stacks do not advertise a flat constant. They advertise the window as a multiple of the maximum segment size, or of the MTU. The relationship is the fingerprint, not just the number. A Linux kernel that advertises ten or twenty times its MSS produces a window that varies with the path’s MSS while the multiplier stays constant, and the multiplier is what identifies it.

The maximum segment size, carried as a TCP option, tells the peer the largest payload the sender wants in one segment. It is derived from the link MTU, so on a normal 1500-byte Ethernet path it lands at 1460, but it shifts on tunnels, PPPoE links, and anything that lowers the MTU. p0f records MSS partly to evaluate the window relationship and partly because the MSS itself, combined with the MTU it implies, points at the link type.

And the TCP options, both which ones appear and the order in which the kernel lays them out. A stack might emit MSS, then a SACK-permitted option, then a timestamp, then a NOP for alignment, then window scale. Another emits MSS, NOP, window scale, NOP, NOP, SACK-permitted, and no timestamp at all. The set and the ordering are not configurable through any normal interface; they are baked into the kernel’s TCP output path. That makes option layout one of the strongest discriminators in the whole packet.

The signature grammar

p0f v3 encodes a TCP stack as a colon-separated string. Reading it is the fastest way to understand what the tool actually keys on. The format for a SYN is:

ver:ittl:olen:mss:wsize,scale:olayout:quirks:pclass

Each field maps to something on the wire. ver is the IP version: 4, 6, or * for both. ittl is the inferred initial TTL, recovered by rounding the observed value, written as the canonical 64, 128, or 255. olen is the length of IPv4 options or IPv6 extension headers, usually zero. mss is the maximum segment size from the TCP option, or * when it varies. wsize,scale is the advertised window and the window scaling factor, and the window is frequently written relative to MSS: mss*20 means twenty times the MSS, not a literal byte count. olayout is the comma-delimited list of TCP options in the exact order the stack emits them, using short tokens: mss, sok for SACK-permitted, ts for timestamp, nop for the one-byte padding, ws for window scale, eol+n for end-of-options followed by n padding bytes. quirks is a comma-separated list of stack oddities. pclass classifies the payload as 0 for empty, + for non-empty, or * for any; a normal SYN has no payload.

A label binds a signature to a human-readable identity. From the shipped database:

label = s:unix:Linux:3.11 and newer

sig = :64:0::mss*20,10:mss,sok,ts,nop,ws:df,id+:0

Read that left to right. Either IP version. Initial TTL 64, so a Unix-family stack. No IP options. Any MSS. Window is twenty times MSS with a scale factor of 10. Options in the order MSS, SACK-permitted, timestamp, NOP, window scale. Quirks: DF set, and the IP ID field is non-zero despite DF being set. Empty payload. That string is a Linux 3.11+ kernel’s TCP personality written down.

label = s:win:Windows:7 or 8

sig = :128:0::8192,0:mss,nop,nop,sok:df,id+:0

Initial TTL 128, the Windows tell. A flat window of 8192 with no scaling. Options in the order MSS, NOP, NOP, SACK-permitted, and notably no timestamp. And macOS:

label = s:unix:Mac OS X:10.x

sig = :64:0::65535,1:mss,nop,ws,nop,nop,ts,sok,eol+1:df,id+:0

TTL 64 like its BSD ancestry, a flat 65535 window with scale factor 1, and a long, distinctive option string ending in an explicit end-of-options marker with one byte of padding. Three operating systems, three signatures, all read from a packet that carries no application data and no encryption.

The OS label itself has structure. A v3 label carries a type, a class, a name, and a flavor. The class is the broad family: unix, win, cisco. The name is the specific OS, Linux or Windows. The flavor is the qualifier, the version range. This is why p0f can answer at different resolutions, “some Windows” when only the class matches, “Windows 7 or 8” when the full signature lines up.

The quirks field

The quirks list is where p0f captures the small illegal-ish behaviors that stacks exhibit, the things that are not parameters so much as tells. The README enumerates them. On the IP side: df for the don’t-fragment flag, id+ for DF set while the IP ID is still non-zero, id- for DF clear while the ID is zero, ecn for explicit congestion notification support, 0+ for a non-zero value in a field that the spec says must be zero, and flow for a non-zero IPv6 flow label. On the TCP side: seq- for a zero sequence number, ack+ for a non-zero acknowledgment number when the ACK flag is not set, ack- for the inverse, uptr+ for a non-zero urgent pointer with no URG flag, plus push and urgent flag oddities. On timestamps: ts1- for an own-timestamp of zero, ts2+ for a non-zero peer timestamp on a SYN, which should not happen. And the catch-alls: opt+ for trailing non-zero data after the options, exws for an excessive window-scale value above 14, and bad for malformed options the parser could not make sense of.

Most of these quirks describe behavior no application can produce and no normal user would ever notice. They exist because TCP stacks were written by different people at different times against a spec with corners, and the corners got handled differently. That is precisely what makes them durable identifiers. A field that “must be zero” but is not tells you something specific about the code path that emitted the packet.

Beyond the SYN: timestamps, uptime, and the HTTP module

p0f reads more than the first packet. Two of its more interesting tricks come from looking at the connection over time and at the layers above TCP.

TCP timestamps, when present, let p0f estimate the remote host’s uptime. The timestamp option carries a value driven by a clock that ticks at a stack-specific frequency. By watching the timestamp advance across packets and knowing the tick rate, p0f can extrapolate backward to when the counter was zero, which is roughly when the stack started. The README notes the tool needs to observe “at least about 25 milliseconds worth of qualifying traffic” before it can lock onto the progression. The result is a free uptime readout for any host whose stack enables timestamps, which is most Linux and BSD systems by default. It is a striking demonstration of how much a passive observer can infer from a parameter that exists for an entirely unrelated reason, round-trip-time measurement and PAWS protection against wrapped sequence numbers.

The HTTP module is v3’s headline addition. It applies the same philosophy a layer up. Rather than parse what a request says, it looks at how the request is structured, on the theory that structure is harder to fake than content. An HTTP signature has the form:

ver is the HTTP version, 0 for 1.0 or 1 for 1.1. horder is the ordered list of headers, with optional name=[value] substring matching on specific header values. habsent lists headers that must not appear. expsw is an expected substring in the User-Agent or Server header, used to catch software that lies about itself. The insight is that a browser sends its headers in a characteristic order and includes or omits a characteristic set, and that ordering is a property of the HTTP client implementation, not of the page being fetched. A client claiming to be one browser while ordering its headers like another has given itself away. This is the same logic that drives modern header order and casing fingerprints and the Accept-header triad signature, and p0f was doing it at the HTTP layer in 2012.

Catching a proxy: the OS-mismatch signal

The single most useful thing passive fingerprinting does in an anti-abuse context is catch a network whose layers disagree with each other. p0f has explicit machinery for this.

When p0f sees a host’s signature change in a way that looks systematic rather than random, it flags it. The README lists the reason codes it attaches: os_sig when the OS signature itself changes, sig_diff for protocol-level changes, tstamp for inconsistent timestamps, ttl for a TTL change, port for a source-port decrease that should not happen, and mtu for an MTU shift. A run of these from one apparent host is the fingerprint of NAT, a proxy, or address sharing, multiple real machines hiding behind one IP, each with its own stack personality.

The sharper version of this is the cross-layer mismatch, and it is the reason network fingerprinting still matters for bot detection in 2026. Consider a request that flows through a proxy. The proxy’s kernel terminates your TCP connection and opens a fresh one to the target. The target therefore sees the proxy’s SYN, with the proxy’s TTL, window, and option layout, not yours. So the OS that the TCP stack implies is the proxy’s OS. Now suppose the HTTP request riding inside claims, via its User-Agent, to be a Windows browser, while the proxy that emitted the SYN runs Linux. The TCP fingerprint says TTL 64, Linux. The User-Agent says Windows. Those cannot both be true for one machine. The contradiction is the detection. As the pydoll write-up on network fingerprinting puts it plainly, “the User-Agent says Windows (TTL 128) but the TCP fingerprint shows Linux (TTL 64)” is the tell that exposes a proxy or a spoofed agent string.

This is why TCP/IP fingerprinting sits underneath the modern anti-bot stack rather than being replaced by it. A scraper can spoof its User-Agent perfectly. It can mimic a browser’s TLS ClientHello with uTLS. It can fake header order. But the SYN that opened the connection came out of whatever kernel actually sent it, and faking that requires control of the network stack itself, not just the application. The mismatch detection generalizes far past p0f’s own database, and it connects directly to the broader problem of detecting a proxy by OS mismatch.

How a fingerprint match actually happens

The mechanics of matching are simpler than the database makes them look. p0f extracts the fields from an observed packet, builds the signature string, and looks for the best match against its loaded fingerprints. A typical install loads on the order of 320 SYN signatures from the p0f.fp file. Matching is not pure equality; wildcards in the database (* for IP version, MSS, payload class) mean a single signature can cover a range of real packets, and the window can be expressed relative to MSS so it matches across paths with different MTUs.

When several signatures could match, p0f resolves at the coarsest level it can be confident about. A packet whose TTL is 64 and whose option layout matches no exact entry might still resolve to “generic Linux” on the strength of the TTL and the broad option shape. The label structure, with its class/name/flavor fields, exists exactly so the tool can give a useful partial answer instead of a useless null. This is also why the database ages gracefully in one direction and badly in another: an unknown new Linux still looks like Linux at the class level even when no flavor matches, but a genuinely novel stack with an unusual option layout can fall through to “unknown” entirely.

The database is the soft underbelly. p0f’s v3 fingerprint set is the 2012-era one, give or take community patches. Operating systems shipped since then have stacks that the original database never saw. The fields p0f reads have not changed, TTL is still TTL, window is still window, but the specific value combinations that modern kernels emit may not have a labeled entry. In practice this means a fresh p0f against 2026 traffic correctly identifies the OS family far more often than it nails the exact version, because the family-level tells (TTL 64 versus 128, timestamp present versus absent, the gross option ordering) are stable across many kernel releases while the fine details drift.

Cloudflare’s BPF compiler, or fingerprinting at line rate

A good illustration of p0f’s signatures outliving the p0f program is what Cloudflare did with them. In an August 2016 engineering post, Cloudflare described compiling p0f’s signature format into Berkeley Packet Filter bytecode. The motivation was SYN-flood defense: during an attack they want to “rate limit attack packets, and in effect prioritize processing of other, hopefully legitimate, ones.” Real operating systems produce SYNs that match known p0f signatures; many flood tools produce SYNs that do not, or that match the signature of a specific attack tool.

Rather than run the p0f daemon in the packet path, they took the human-readable signature grammar and built a compiler that emits BPF, so the classification runs in the kernel’s packet filter at line rate, inside iptables. The blog gives worked examples in p0f’s own format, a Linux SYN as 4:64:0:*:mss*10,6:mss,sok,ts,nop,ws:df,id+:0, a Windows 7 SYN as 4:128:0:*:8192,8:mss,nop,ws,nop,nop,sok:df,id+:0, and a hping3 attack packet that betrays its synthetic origin with a sparse signature and an ack+ quirk. The signature language outlived the tool that defined it, which is a fair measure of how good the abstraction was. The same field set that Cloudflare compiles to BPF is what feeds the network-layer side of vendor systems like DataDome’s HTTP/2 and network fingerprinting.

Where the technique stands in 2026

Passive OS fingerprinting from the SYN is older than most production TCP stacks running today, and it is not going away, but it has limits worth stating plainly.

What still works: the family-level signal. TTL of 64 versus 128 still cleanly separates the Unix-descended world from Windows, and no normal application can change it. Option ordering and the presence or absence of the timestamp option still split the major stacks. The window-to-MSS relationship still holds. Most importantly, the cross-layer mismatch check, network fingerprint disagreeing with the User-Agent, is more useful now than it was in 2012, because the modern internet is saturated with proxies, VPNs, and CGNAT, and detecting that someone is behind one is valuable in itself.

What erodes it: middleboxes and normalization. NATs, VPN concentrators, load balancers, and traffic normalizers rewrite TTL or rebuild TCP options, which either changes the observed signature or smears many real hosts into one. A host behind a corporate VPN may fingerprint as the VPN appliance. The rise of cloud egress means a huge share of traffic now originates from a handful of stack types on Linux hypervisors, which compresses the diversity the technique feeds on. And the v3 database’s age means version-level precision has decayed even where family-level identification holds.

What it never reached: the payload. The whole point of passivity is that p0f reads what is already on the wire, and increasingly what is on the wire is encrypted. p0f’s HTTP module sees nothing inside a TLS connection. That is why the center of gravity for application fingerprinting moved to the one plaintext handshake that remains, the TLS ClientHello, and to the structure of HTTP/2 framing, both of which leak in the clear even when the body does not. The network layer that p0f reads is one floor below all of that, and it stays readable precisely because TTL and window size and option order travel in the clear by necessity, not by choice.

The durable lesson of p0f is architectural, not about any one tool. Identity leaks at the layer that the application does not control. A program can lie about everything it writes, its User-Agent, its headers, its claimed OS. It cannot easily lie about the SYN, because the SYN belongs to the kernel. Twenty-six years after a Bugtraq post claimed a 63-bit signature for every system, the cheapest, most undetectable signal on the network is still the one the application never gets to touch.

Sources & further reading

  • Zalewski, M. (2012), p0f v3 project page — the author’s own description of the v3 rewrite, its scope, and what it reads from IP/TCP/HTTP.
  • p0f project (2012), p0f v3 README — the authoritative reference for the signature grammar, the full quirks list, uptime estimation, and NAT reason codes.
  • p0f project, p0f.fp fingerprint database — the shipped signature file with the section structure and real labeled signatures for Linux, Windows, and Mac OS X.
  • CGSecurity, OS fingerprinting — background on passive versus active methods and how initial TTL is recovered from the observed value by rounding to 64/128/255.

Further reading

TCP/IP stack fingerprinting: TTL, window size, and MSS as OS identity

Traces how the initial TTL, TCP window size, MSS, and the order of TCP options in a single SYN packet identify the sending operating system, and why that identity is set by the kernel rather than the browser.

Tue, February 17, 2026 ·22 min read

TCP timestamp and window-scaling fingerprints across operating systems

Traces how the order of TCP options in a SYN packet, the window-scale shift count, the SACK-permitted flag, NOP padding, and the timestamp clock identify an operating system, and how per-connection randomization changed what the timestamp leaks.

Sun, February 15, 2026 ·22 min read

MTU, path MTU discovery, and the fingerprint hiding in packet sizes

How tunnels and VPNs shift MTU and MSS, why a non-standard MSS in a SYN packet betrays an encapsulated path, and how path-MTU discovery behavior turns a packet-size value into a signal.

Fri, February 13, 2026 ·20 min read