Replicating a Mobile App Protocol from the Client Side: Black-Box Differential Debugging, Coordinate Systems, and Deterministic Simulation
声明 / Disclaimer:本文基于已获授权的个人账号与教学研究目的撰写。签名算法与
密钥、载荷字段全集、接口路径清单、验证码绕过操作步骤等内容已从文中省略;文中
给出的实验方法、几何与测试工程细节可完整复现,防御部分建议先行通报厂商后公开。This work is based on a personally authorized account and written for educational
research. The signature algorithm and salts, the complete field inventory, API path
lists, and CAPTCHA-bypass procedures are deliberately omitted. The methodology,
geometry, and testing details are fully reproducible; the defensive findings should
be reported to the vendor before public release.
摘要
校园跑步类 App 将”成绩”作为核心业务资产,其客户端—服务端协议天然偏向
“客户端声明、服务端记账”的信任模型。本文以一个可运行的纯 Python 客户端为对象,
系统梳理了在不修改 APK、不注入运行时、不可得密钥的前提下,仅通过黑盒观测复刻
该协议栈的工程方法。全文围绕四个研究问题展开:(RQ1)当服务端把所有失败折叠成
同一句文案时,如何建立可观测的差分调试框架;(RQ2)如何从有限样本反推字段的
类型、单位、空值与布尔规范化语义;(RQ3)如何在 WGS84/GCJ-02 双坐标系与多边形
地理围栏约束下生成几何可信的轨迹,并使随机性不破坏任何恒等式;(RQ4)如何通过
种子化双阶段生成与属性测试,使”预览即上传”成为可验证不变量,并把服务端自身的
规则判定回显当作测试 Oracle。文章给出四个方法论命题、一套含 24 项单元测试的
验证体系、三条端到端闭环证据,并从服务端视角提出六条可落地的防御建议及其检测
数学形式。最后讨论本工作的局限性、伦理边界与负责任披露路径。
关键词:移动应用安全;协议逆向;黑盒差分调试;地理围栏;GCJ-02;确定性仿真;
属性测试;验证码自动化;API 安全
Abstract
Campus running applications treat “performance records” as their core business asset,
and the resulting client–server protocol is inherently biased towards a
“client-declares, server-books” trust model. This paper presents a working pure-Python
client and systematically describes how the complete protocol surface can be replicated
through black-box observation alone—without modifying the APK, without runtime
instrumentation, and without access to any key material. Four research questions are
addressed. (RQ1) How can one build an observable differential-debugging framework when
the server collapses every failure into a single canned message? (RQ2) How can field
types, units, null handling, and boolean canonicalization be inferred from a finite
number of samples? (RQ3) How can geometrically credible trajectories be generated under
dual coordinate systems (WGS84/GCJ-02) and polygonal geofences, while ensuring that
randomization never violates any conservation invariant? (RQ4) How can “preview equals
upload” be turned into a verifiable invariant through seeded two-stage generation and
property-based testing, and how can the server’s own rule echo be used as a test oracle?
The paper contributes four methodological propositions, a validation regime comprising
24 unit tests, and three pieces of end-to-end closure evidence. Six concrete defensive
recommendations are derived from the observed weaknesses, each accompanied by an
executable detection formulation. Limitations, ethical boundaries, and a responsible
disclosure path are discussed at the end.
Keywords: mobile application security; protocol reverse engineering; black-box
differential debugging; geofencing; GCJ-02; deterministic simulation; property-based
testing; CAPTCHA automation; API security
第一部分 · 中文版
1 引言
1.1 研究背景
定位类校园跑步 App 的业务闭环可以概括为:手机采集轨迹 → 客户端组装成绩报文 →
服务端验签、套规则、写库、聚合。这条链路上,成绩的所有要素都由客户端声明:
里程、时长、步数、配速、点位通过情况,乃至”记录是否完整”这一布尔值。服务端的
职责被压缩为三件事:验证签名、套用学校配置规则、维护聚合口径。
从安全工程的角度看,这是一个典型的”不可信边界放错了位置”的系统。本文不讨论
该系统应然的形态(那是第 11 节的主题),而是回答一个实证问题:如果攻击者只
拥有一个普通账号与网络抓包能力,复刻这个协议面需要哪些方法、会遇到哪些可观测
性障碍、这些障碍如何被工程手段化解。
1.2 问题定义:失衡的信任模型
| 要素 | 声明方 | 校验方 | 观察到的现状 |
|---|---|---|---|
| 轨迹坐标序列 | 客户端 | 无重算 | 直接入库 / 对象存储 |
| 总里程、总时长、总步数 | 客户端 | 与学校阈值比较 | 不与轨迹重算比对 |
| 配速 | 客户端 | 与配置区间比较 | 持久化时缩放 10⁻³ |
| 点位通过情况 | 客户端 | 与学校点位配置比较 | 按客户端布尔值判定 |
记录完整性 complete |
客户端 | 无 | 直接信任 |
| 签名 | 客户端 | 服务端重算比对 | 唯一的强校验 |
表 1 的结论是不言自明的:唯一的强校验保护的是”报文没有被中间人篡改”,而不是
“报文内容为真”。 签名回答的是”话有没有被改”,不回答”话是不是真的”。
1.3 研究问题
本文将上述问题域拆解为四个可独立验证的研究问题:
- RQ1(可观测性):当服务端把解析失败、校验失败与系统异常折叠为同一句
文案时,能否仅凭返回类别建立可归因的差分调试框架? - RQ2(语义辨识):在不可得源码的条件下,字段的类型、量纲与规范化规则
能否被系统性地反推? - RQ3(几何可信性):如何在双坐标系与多边形围栏的双重约束下生成轨迹,
并使随机性在引入多样性的同时不破坏任何守恒式? - RQ4(可验证性):如何把”预览”与”上传”收敛为同一条可重放路径,并以
服务端回显的规则判定作为验收 Oracle?
1.4 本文贡献
- 错误码三分类驱动的差分调试流程(对应 RQ1):将”统一错误文案”这一
可观测性缺陷转化为稳定的布尔探针,给出实验矩阵、推论链与三条方法论命题; - 字段语义黑盒辨识方法(对应 RQ2):类型探测、单位反推、空值与布尔
规范化,附本产品上的实测案例与完整算式; - 种子化双阶段生成架构与随机性治理原则(对应 RQ3、RQ4):预览与上传
共享随机源状态,”预览即上传、预览零副作用”成为可测试性质;任何抖动都必须
在归一化后保持总量恒等; - 以服务端规则回显为 Oracle 的验证体系:24 项单元测试 + 三条端到端
闭环证据(status = 0记录、有效里程由 0 增长至 3.0 km); - 六条防御建议(对应第 11 节),每条给出触发证据与可执行的检测形式。
1.5 文章结构
第 2 节回顾相关工作;第 3 节给出系统概述、威胁模型与实验纪律;第 4 节回答 RQ1;
第 5 节回答 RQ2;第 6 节回答 RQ3;第 7 节回答 RQ4 的仿真与测试部分;第 8、9 节
分别讨论会话管理与有效性验证;第 10 节描述控制台实现;第 11 节给出防御建议;
第 12 节讨论伦理与披露;第 13 节陈述局限性;第 14 节总结。
2 相关工作
2.1 早期公开分析(2016 年前后)
该产品的协议面最早由若干安全爱好者在 2016 年前后公开分析 [2][3]。当时的研究
以抓包为主,结论集中在三点:传输层为明文 JSON;鉴权采用可解码的 Basic 凭据;
服务端对成绩字段几乎不做交叉校验,客户端可构造任意里程与时长。这些工作确立了
“客户端声明式成绩”这一基本判断,但受限于当时的版本,其字段结构、签名机制与
现代版本已相去甚远。
2.2 社区工具与运行时模块
后续社区工作可分两类。一类是宿主在系统层的运行时修改(如将有序点位改回无序
的 Xposed 模块 [9]),依赖 root 与框架环境;另一类是独立客户端实现(如开源的
纯 Python 复刻项目 [1] 及各类脚本 [3]),不依赖宿主环境,但通常只覆盖部分运动
类型,且普遍记录”计分跑需要账号已有跑步计划”这一前置条件而未给出解释。本文与
[1] 最接近,差异在于:本文将方法学(差分调试的归因纪律、语义辨识的量纲闭环、
确定性测试)作为主要贡献对象,并补充了会话生命周期、副作用分级与防御视角。
2.3 移动端反滥用的一般方法
反作弊与反自动化领域已有成熟共识:服务端权威计算、统计特征异常检测、分层错误
可观测性、设备与会话风控 [7]。然而在本类产品中,这些共识的落地程度普遍偏低——
成绩字段直接信任客户端是最典型的偏差。本文第 11 节将把观察到的每个弱点映射到
上述共识条目,并给出具体检测形式。
2.4 坐标系与地理围栏
WGS84 与 GCJ-02 的系统性偏差及其互转是中文地理信息工程中的常识问题 [8],但
在协议复刻场景下,它常以”静默偏差”的形式出现:本地校验通过、服务端判定越界,
且两侧都不报错。本文第 6 节把坐标系边界纪律作为一等工程约束处理,并给出一个
真实反例。
3 系统概述、威胁模型与实验纪律
3.1 分层架构
┌───────────────────────────────────────────────────────────┐
│ Web Console stdlib HTTP + 单页前端(仅绑定 127.0.0.1) │
│ 接口按副作用分级:plan / submit │
├───────────────────────────────────────────────────────────┤
│ Run Planner 风格 → 距离/配速/步频分布 │
│ 有效时段内的回填时刻 → 可序列化"计划"对象 │
├───────────────────────────────────────────────────────────┤
│ Protocol Core 报文建模 · 字段规范化 · 加盐摘要签名 │
│ 轨迹几何生成 · 十秒分桶 · OBS 载荷组装 │
├───────────────────────────────────────────────────────────┤
│ Transport 会话与设备头 · 预签名对象存储上传 │
├───────────────────────────────────────────────────────────┤
│ Challenge 极验 v4:渲染 → 边缘定位 → 轨迹生成 → 回传 │
└───────────────────────────────────────────────────────────┘
图 1:五层架构。层间依赖自上而下单向,Protocol Core 不持有任何会话状态,
使其可被离线单元测试覆盖。
3.2 一次提交的数据流
plan(本地) ──► build(seed) ──► [几何 + 报文 + 签名]
│
挑战校验 ◄── ① ──► 滑块求解 ──► ② 验证回传
│
③ OBS 预签名上传(gzip + base64 × 11 段)
│
④ 保存记录(含 recordUrl / isUpload 后置字段)
│
⑤ 详情回读(规则逐条判定 + 成绩聚合)
图 2:端到端共 5 次往返。注意第 ④ 步中的两个字段是在签名之后才注入的——
它们被有意排除在签名集合之外。这提示我们:签名集合是一份显式的、需要精确复刻的
清单,而不是”所有字段”。
3.3 威胁模型
| 维度 | 设定 |
|---|---|
| 攻击者能力 | 持有合法账号;可抓包;可离线分析 APK 字符串与 DEX;可运行任意本地代码 |
| 攻击者不具备 | 服务端源码、日志、数据库、密钥材料、运行时插桩环境 |
| 环境假设 | 账号所有者授权;网络可达;学校规则配置可经接口读取 |
| 研究目标 | 完整刻画协议面并构造一致性客户端,评估该信任模型的稳健性 |
3.4 实验纪律
黑盒研究最大的风险不是失败,而是污染:一次未登出的会话、一次意外的状态写入,
都会让后续几十组实验失去可比性。本文的纪律是:
- 单批次单登录:每个批次只登录一次,批量执行探针后立即登出,释放设备绑定;
- 副作用分级:只读探针与写探针严格分脚本,写探针文件名显式标注;
- 状态变更前置授权:涉及账号业务状态(选课、目标、计划)的写操作一律先
征得账号所有者同意; - 每批留存原始响应:所有实验结果落盘为 JSON,便于事后复算与回溯;
- 先做对照组:任何”新变量”实验都必须伴随一个已知行为的对照,否则类别
跳变无法归因。
4 方法(RQ1):错误码三分类与单变量差分调试
4.1 形式化:一个三值输出的服务端
该保存接口的全部可观测输出只有三类(表 2):
| 记号 | 码值 | 语义 | 实验价值 |
|---|---|---|---|
E_sig |
签名校验失败 | 本地规范化与服务端重算不一致 | 指示类型/规范化错误 |
E_biz |
统一文案 | 签名通过后服务端抛异常并被吞掉 | 指示内容/上下文错误 |
E_ok |
成功 | 完整走通并入库 | 当前变量组合正确 |
表 2:三分类。E_biz 的文案固定为”服务器开小差”,不携带堆栈、不区分反序列化
失败与业务校验失败、不区分空指针与类型转换异常。
定义 1(观测函数)。令 X 为请求空间,定义 f : X → {E_sig, E_biz, E_ok}。
对运维而言,f 的坍缩是可观测性缺陷;对差分实验而言,f 是一个稳定的布尔
探针——只要改动引起类别跳变,就说明该改动落在了”签名之后的第一个处理分支”上。
由此,实验问题形式化为:寻找使 f(x) = E_ok 的输入 x,且 E_biz → E_ok 的
跳变点即为进入正确处理分支的充要条件。
4.2 实验设计三原则
原则一:严格单变量。 同一批次内,除目标自变量外的一切字段(时间、坐标、
签名、会话头)全部冻结。混杂变量会使跳变无法归因。
原则二:必须有对照组。 每个取值至少执行一次已知对照(例如对照组换成另一种
运动类型,期望 E_ok),以排除”环境漂移”——时段变化、会话过期、限流都会让
结果改变,但它们与目标变量无关。
原则三:跳变必须回退复验。 出现类别跳变后立即把变量改回原值,确认类别回到
原状态。没有回退复验的跳变不计入结论,因为多数伪跳变(环境漂移、限流窗口)
不满足”可逆”这一性质。
4.3 实验矩阵
表 3 是一个批次的真实矩阵,共 20 余个自变量,节选 10 行。除最后一行外,全部
落在 E_biz:
| # | 自变量 | 变化范围 | 结果 | 推论 |
|---|---|---|---|---|
| 1 | 时间窗口 | 有效时段内 / 窗外 5h / 未来时刻 | E_biz |
时间校验不在异常点之前 |
| 2 | 里程 | 50 m / 1000 m / 50 km | E_biz |
距离阈值校验不在异常点之前 |
| 3 | 时长、步数 | 极小 / 常规 / 极大 | E_biz |
同上 |
| 4 | 十秒分桶 | 全零 / 均匀 / 抖动 | E_biz |
分桶解析不抛异常 |
| 5 | 地理位置 | 围栏内 / 围栏外 600 m | E_biz |
围栏判定不在异常点之前 |
| 6 | 会话标识 | 头部 / 查询串 / 缺省 | E_biz |
该字段非必需入参 |
| 7 | 目标与主题 | 有值 / JSON null / 字符串 “null” | E_biz |
空值形态不是异常源 |
| 8 | 数值类型 | 浮点 / 整数 | E_sig 或 E_biz |
反序列化目标类型可被探测 |
| 9 | 序列化形态 | 裸集合 / 外包配置对象 | 裸集合 E_biz;包装 E_ok |
异常源定位 |
| 10 | 对照组:另一运动类型 | 同上任意 | E_ok |
环境正常,跳变可归因 |
表 3:实验矩阵(节选)。第 9 行是结论行。
4.4 推论链与异常穿透效应
第 9 行的解释是:服务端在按运动类型分派后的入口处直接把某个嵌套字段
反序列化为包装类型,裸集合触发空指针,异常逃逸到统一捕获层,于是所有下游校验
(时间、距离、围栏、点位)都被跳过——这正是第 1~5 行全部失效的原因:
它们根本没被执行到。
由此得到一个可推广的观察:
命题 1(异常穿透效应)。若反序列化入口缺少类型校验,且异常被统一错误
中间件捕获,则任何位于该入口之后的业务校验都构成旁路:改变这些校验对应的
输入不会改变观测结果。因此,”改什么都无效”本身就是入口处异常的强证据。
4.5 方法论小结
命题 2(类别即信号)。当失败被折叠为单一文案时,返回类别仍构成
二值探针;观测信息量的下界不为零。
命题 3(可解释性准则)。一个异常源假设只有在能解释全部历史矩阵时才
被采纳。本文的假设同时解释了 20 余行为何全部 E_biz,这是它比”试出来的
正确姿势”更强的地方。
5 方法(RQ2):字段语义的黑盒辨识
签名通过只说明”格式一致”,不说明”语义正确”。本节讨论四类语义的辨识方法。
5.1 类型探测
向疑似整型字段注入浮点值:若被签名层以 E_sig 拒绝,而等值整型通过,即可判定
反序列化目标为整型——因为服务端在反序列化之后、签名之前已把值规范化为整型,
本地按浮点字符串参与摘要必然不一致。这一手段的价值在于:它不需要任何源码,
仅凭错误类别即可推断出字段的 Java 类型。
推论。类型探测可以批量执行:对每个候选字段做两次请求(浮点、整型),记录
类别矩阵,即可得到一张”字段 → 类型”的黑盒表。
5.2 单位反推:一个 10³ 因子的定位过程
观测链如下:
- 本地上送字段值
2381; - 详情回显同一记录该字段为
2.381; - 假设持久化存在缩放
k,则k = 10⁻³; - 取同批次的真实记录做交叉验证:某记录总里程 3450 m、总时长 1257 s,
则1257 / 3.45 = 364.3 s/km = 6.07 min/km,恰为该记录回显值 6.07。
结论:该字段语义为配速(分钟/公里),线路上行值为其 1000 倍。
单凭第 1、2 步只能得到”k = 10⁻³”,第 4 步才把量纲钉死为 min/km——这就是
对照组在语义辨识中的作用:缩放因子可以猜,量纲必须用独立样本闭环。
由此还解释了一个现象:早期客户端按 距离/时长 × 1000 构造该字段,恰好是
m/s×1000 而非 min/km×1000,签名可通过、语义完全错误,导致配速规则必然判负。
这是”格式对、语义错”的典型样本——它不会被任何签名校验捕获,只能被规则回显
(第 9 节的 Oracle)暴露。
5.3 空值、布尔与字符串的规范化
Java 侧的 String.valueOf(null) 是签名差异的高发区。本地必须区分三种形态
(表 4):
| 形态 | 参与签名的字符串 | 适用条件 |
|---|---|---|
| 字段缺省 | 不参与 | 服务端模型中无此属性 |
JSON null |
"null" |
服务端为对象类型且值为 null |
字符串 "null" |
"null" |
等价于上一行 |
表 4:空值的三种形态。布尔值同样需要规范化(True/true → true),整数与
字符串数字(1 vs "1")不能混用。本文把这些规则集中到一个纯函数中,使其可被
单元测试穷举。
5.4 时间格式的双重角色
对同一时刻注入毫秒时间戳、ISO 字符串、日期字符串三种形态,观察哪一种进入签名
集合且最终入库一致。毫秒时间戳同时承担两个角色:业务上的起止时刻与
几何上的打点基准——因此它既参与签名,又必须与轨迹点的时间轴严格对齐
(stop − start = duration,轨迹点时间在 [start, stop] 内单调)。
任何一处不自洽,都会在详情回读或地图渲染阶段暴露。
5.5 小结
本节的四类辨识共同支撑一个结论:签名一致性是必要条件而非充分条件。
类型、量纲、规范化规则必须分别独立验证,且量纲验证必须引入独立样本闭环。
6 方法(RQ3):双坐标系与围栏约束下的轨迹生成
6.1 坐标系选择与边界纪律
地图瓦片与地理常识使用 WGS84,而国内地图供应商与移动端定位 SDK 输出 GCJ-02,
二者在中国大陆存在数百米的系统性偏移。工程原则只有一条:
选定唯一建模坐标系,只在系统边界处转换。
本文的约定:
- 围栏数据(服务端下发,GCJ-02)在读入时立即逆变换为 WGS84;
- 轨迹在 WGS84 中生成、在 WGS84 中做围栏预检;
- 仅在构造上传报文的那一刻正变换为 GCJ-02;
- 地图预览使用 WGS84,与瓦片天然对齐。
6.2 逆变换的不动点迭代
GCJ-02 → WGS84 没有闭式逆,采用不动点迭代:
w₀ = g
wₙ₊₁ = wₙ − ( F(wₙ) − g ) # F 为 WGS84 → GCJ-02
本文取 6 次迭代、收敛阈值 10⁻⁷ 度(约 1 cm)。实践表明三次以内已进入阈值,
取 6 次是防御性冗余。由于 F 在局部近似为仿射变换且偏移量远小于坐标本身,
该迭代的收敛性在实际区域内是良态的。
反例(静默偏差):本项目早期版本把”GCJ-02 围栏顶点”与”WGS84 生成点”直接
比较,产生约 600 m 的静默偏差,表现为”围栏内预检通过、实际入库越界”。这类错误
的可怕之处在于它不抛异常——只有把两个坐标系的量纲写进同一张校验表,才会暴露。
6.3 多边形围栏的本地预检
围栏为 6~19 顶点多边形。预检包含两项:
- 点在多边形内:射线法(crossing number),对每个轨迹点执行;
- 点到边的最小距离:逐边点到线段距离取最小,与阈值(本文 3 m)比较,
防止”擦边”点在地图放大后越界。
预检不通过的参数组合直接判为不可用,避免”上传后才知道越界”的高成本试错——
这正是把服务端校验前置到本地的典型收益。形式化地,令 P 为围栏多边形、
T = {p₁,…,p_N} 为轨迹点集,则接受条件为
∀pᵢ ∈ T : pᵢ ∈ interior(P) ∧ dist(pᵢ, ∂P) ≥ δ,本文 δ = 3 m。
6.4 跑道形状的弧长参数化
轨迹不是”圆”,而是跑道形状:两段直线 + 两端半圆。生成算法:
loop = 直线段点列 + 半圆弧点列 # 局部米制坐标,闭合
cum = 累积弧长表
lap = cum[-1] # 单圈周长(约 301 m)
θ = 跑道长轴方位角
point_at(s): # s 为弧长
在 cum 上二分查找 → 段内线性插值 → 乘 0.95 内缩
for i in 0..N-1:
s = (phase + v·t_i) mod lap # phase 为起跑相位
p = point_at(s)
if i > 0: p += 三角分布噪声(±2.5 m)
旋转 θ → 加到围栏中心 → WGS84 → 转 GCJ-02
三个设计点,每个都对应一个可识别的关联风险:
- 相位参数
phase:使每次运行的起点不同。若起点固定,同一账号多次记录的
首坐标在小数点后六位完全相同,这是一个零成本即可检出的关联特征; - 内缩因子 0.95:让路径整体位于中心线内侧约一个跑道,为噪声留出余量,
保证”中心线贴边”时噪声仍不越界; - 噪声分布选三角分布而非高斯:高斯无界,理论上必然越界;三角分布有界、
单峰、中心化,形态接近真实 GPS 误差且可证明不越界。
6.5 速度与分桶的随机性,以及它不能破坏的恒等式
- 瞬时上报速度:
v·(1 + 0.04·sin(6πt/T) + U(−0.025, 0.025)),均值等于目标速度; - 十秒分桶:距离份额与步数份额各乘
1 + U(−0.08, +0.08)的权重; - 归一化:份额必须精确还原声明总量,否则”随机性”本身会成为异常特征。
步数为整数,采用最大余数法:
raw = [total_steps * w_i / Σw for w in weights] # 浮点份额
alloc = [int(x) for x in raw] # 向下取整
left = total_steps - sum(alloc) # 必然 0 ≤ left < n
order = 按小数部分降序
alloc[order[:left]] += 1 # 严格守恒
距离份额为浮点,按权重归一后求和即为声明总里程。两条守恒式
Σ stepsBucket = totalSteps、Σ distBucket = totalDis 各有一条单元测试看守。
命题 4(守恒优先)。随机性只允许改变分布形态,不允许改变总量。
任何引入多样性而破坏守恒的实现,都会把随机性转化为新的异常特征。
6.6 点位(检查点)的生成
五个检查点按行进顺序取样:frac_k = k/6 + U(−0.04, +0.04)(k = 1..5),
排序后落到实际轨迹点上,从而天然满足”每个检查点坐标都能在完整轨迹中找到”——
这是另一条被单元测试看守的性质。顺序跑模式下位置编号即行进顺序,保证”按顺序
通过所有点位”这一规则可被判定为通过。
7 方法(RQ4):确定性仿真、随机性治理与测试体系
7.1 计划对象:把一次跑步变成可序列化数据
规划器输出一个可序列化的”计划”对象:
{
"style": "slow",
"distance_m": 2663, "duration_s": 1382, "steps": 3600,
"stopMs": 1790345740000, "zoneIdx": 1, "seed": 271828182,
"phaseFrac": 0.4137, "checkpoints": [0.15, 0.31, 0.52, 0.66, 0.83],
"goalId": 406727, "selDistance": 1000, "avgPower": 151
}
字段的含义:seed 决定整条轨迹的随机序列;phaseFrac 决定起跑相位;
checkpoints 决定点位分布;stopMs 决定记录落在有效时段的哪个位置。
全部显式写进计划,是为了让同一份计划在任何时候重放结果都一样。
7.2 随机源边界与两条性质
预览与上传是两条代码路径,但它们在生成几何之前都执行 random.seed(plan.seed),
之后调用同一生成函数并传入同样的 phaseFrac / checkpoints。只要两次调用之间
没有其他随机源消耗,产出必然逐点相同。
由此得到本文最重要的两条可测试性质:
- P1(确定性):同一计划连续两次
build()的坐标序列全等; - P2(无副作用):预览接口不触网、不写库,可无限次调用。
P1 的实现陷阱在于:uuid4、时间戳等若误用同一随机源,会破坏两次调用的序列对齐。
本文的对策是明确随机源边界——几何随机只用 random 模块并在入口统一播种,
标识符随机走 os.urandom。随机源的边界必须显式声明,否则确定性无法成立。
7.3 随机性治理:多样性必须合法
“每次不同”与”每次都合规”是一对矛盾,治理办法是把随机变量全部约束在规则的
内点上,并对边界留出安全余量(表 5):
| 随机维度 | 分布 | 约束 | 余量来源 |
|---|---|---|---|
| 风格 | fast/steady/slow 均匀 | 距离 ∈ [1000, 3000] m | 日上限 3000 m |
| 配速 | 按风格区间均匀 | ∈ (5.0, 10.0) min/km | 上下各留 0.3–0.4 |
| 步频 | 按风格区间均匀 | ∈ [60, 300] spm | 实际取 152–190 |
| 回填时刻 | 有效时段内均匀 | 整次跑步落在时段内 | 两端各留 4 分钟 |
| 起跑相位 | U(0,1) | 起点在围栏内 | 内缩 + 噪声有界 |
| 分桶抖动 | U(−8%, +8%) | 归一化守恒 | — |
表 5:随机性治理表。
其中配速余量是针对分段分析的防御:若学校开启分段校验,单桶配速会因 ±8%
抖动而偏离整体配速,5.4 × 0.92 ≈ 4.97 说明按 5.4 起步仍可能贴边,因此
速跑风格的下界被设为 5.45。先算最坏情况,再定参数。
7.4 测试体系:24 项单元测试的分组
| 测试文件 | 数量 | 看守的性质 |
|---|---|---|
test_coordinate_consistency.py |
8 | 锚点与首轨迹点一致、检查点必须属于轨迹、坐标边界与非有限值拒绝、敏感字段脱敏 |
test_outdoor_record.py |
6 | 点位载荷的包装形态、配速字段的 1000 倍语义、状态与主题默认值、围栏内包含性、重复运行起点/点位必须不同、分桶求和守恒 |
test_random_run.py |
5 | 300 次采样全部落在学校规则内、风格间距离区间互不重叠、越界覆盖值被钳制、回填时刻必在时段内、时段外回退语义 |
test_web_server.py |
5 | 错误与规则文案全英文、学期名国际化、计划与几何确定性、计划形状不重复 |
表 6:测试分组。三类测试对应三种回归风险:协议语义回归(签名与载荷)、
几何回归(越界、重复)、界面语言回归(错误码翻译被删)。最后一类测试
直接断言输出字符串不含 CJK 码位——把”文案规范”从 code review 变成机器约束。
7.5 属性测试
对规划器这种”随机输入 → 必须合法输出”的组件,逐例断言没有意义,应写属性测试:
for _ in range(300):
plan = plan_run("random", rng)
assert MIN_DISTANCE <= plan.distance_m <= MAX_DISTANCE
assert PACE_TOP <= plan.pace <= PACE_BOTTOM
assert plan.duration_s >= MIN_TIME
assert 60 <= plan.steps * 60 / plan.duration_s <= 300
300 次采样覆盖三档风格与全部区间端点,任何一次越界即失败。这套断言同时是
文档:它精确写出了”什么叫一次合法的跑步计划”。
8 会话、设备绑定与副作用管理
8.1 登录的三段式结构
登录由三段构成:状态检查 → 挑战求解(可跳过)→ 凭证交换。凭证交换使用
HTTP Basic 携带账号口令,换取 token 并同时完成设备绑定。工程要点:
- 挑战求解平均 15 s,必须作为异步进度暴露给界面(否则用户会以为卡死);
- 登录成功后立即把
uid/token/unid写入会话对象,后续所有请求共用同一设备头; - 失败必须区分”凭证错误”“设备保护”“挑战失败”三种,否则用户无法自救。
8.2 会话状态机与保护态代价
logged_out ──login──► pending ──ok──► active ──logout──► released
▲ │ │ │
└───── 冷却期后 ◄────┴── 保护态 ─────┘ │
(同账号他处在线) │
└────────────────── released 即可重新 login ────────────┘
图 3:会话四态。实测语义:同一账号的新登录会把旧会话打入保护态,服务端要求
“原设备登出或等待一小时”。这一语义对自动化是硬约束:
- 每次登录成功后必须登记待登出标记,即使后续实验抛异常也要在
finally
中释放; - 探针脚本把登出视为一等结果(打印其码值),而非可忽略的善后;
- 网页控制台把”登出失败”直接反馈给用户,因为它意味着账号在一段时间内不可用。
本项目在开发过程中真实遭遇过一次保护态,代价是整批实验推迟一小时——这构成了
“失败必须可区分”的反向证据:忘记登出的代价不是本次失败,而是后续全部实验被锁。
8.3 副作用分级
| 接口 | 方法 | 副作用 | 触发条件 |
|---|---|---|---|
/api/state |
GET | 无(若已登录则回读成绩与历史) | 页面加载 |
/api/plan |
POST | 无(纯本地几何与报文构造) | 预览按钮 |
/api/submit |
POST | 写库 + 对象存储 | 提交按钮,需 active 会话 |
/api/login |
POST | 建立会话、设备绑定 | 显式登录 |
/api/logout |
POST | 释放设备绑定 | 显式登出 |
分级的收益是把”误触”的代价降为零。实现上的一个教训是:辅助函数若以”有无
body”隐式决定 GET/POST,就会出现”退出按钮发 GET 被 404 静默吞掉”这类缺陷——
显式优于隐式,副作用分级必须体现在函数签名上,而不是调用约定上。该缺陷在本项目
中真实出现过,并表现为”用户点击退出毫无反应”。
9 有效性验证:把服务端判定当作 Oracle
9.1 规则回显是最廉价的测试 Oracle
黑盒测试最难的是”不知道什么算对”。本系统的幸运之处在于:记录详情会逐条回显
规则判定结果(每条含类型编号、规则文案、通过与否)。这相当于服务端把判定函数的
中间结果暴露了出来,可以被直接当作断言:
reasonList = [
{type: 9, ok: true, "按顺序通过所有点位"},
{type: 1, ok: true, "最少要跑1.00公里"},
{type: 7, ok: true, "跑7分钟以上"},
{type: 11, ok: true, "步频控制在60-300步/分钟之间"},
{type: 12, ok: true, "配速控制在5′00″-10′00″"},
]
status = 0 # 0 = 计入成绩;1 = 异常无效
一次提交的验收标准因此是:五条规则全 ok 且 status = 0。
9.2 Oracle 在归因中的作用
在修复过程中,这个 Oracle 精确地把问题收敛到第 12 条(配速):配速字段单位错误
导致其数值落在区间之外,而其余四条一直是通过的。若没有逐条回显,”成绩无效”将
是一个无法归因的黑箱结论——Oracle 的价值不在于判定成败,而在于定位失败原因。
9.3 记录状态机与聚合口径
- 记录级:
status ∈ {0, 1},1伴随”异常成绩”标签与申诉入口; - 聚合级:有效里程与总里程是两个字段,只有
status = 0的记录贡献
有效里程; - 实测中曾观察到”逐条相加与聚合回显不一致”(3.53 vs 3.00 km),怀疑与异步
刷新或取整口径有关。本文的处理原则是:以服务端聚合为准,不自行推算,
并把该差异记为待观察项而非缺陷。
9.4 端到端验收序列与闭环证据
修复的最终验收不是”返回成功”,而是一条完整序列:
plan → build(seed) → OBS 上传 → 保存 → 详情回读(五条规则全过)
→ 成绩聚合回读(有效里程增长) → 登出释放
本项目以三条 status = 0 的记录、有效里程从 0 增长到 3.0 km 作为闭环证据,
其中一条来自 Web 控制台的完整链路(含回填时刻、随机风格、地图预览),
一条来自命令行链路,一条用于验证围栏与回填时刻的组合。
10 实现:Web 控制台
10.1 零依赖选型
控制台选择 Python 标准库 http.server(多线程版)+ 单文件前端,而非引入 Web
框架,理由有三:(1)工具的攻击面应尽可能小,本地服务只需绑定回环地址;
(2)依赖越少,越容易在他人机器上复现;(3)接口数量在十以内,框架的收益接近于
零。服务端状态(会话、规则缓存)保存在进程内字典中,不落盘、不持久化凭据。
10.2 错误码国际化与顺序陷阱
服务端错误与规则文案原生为中文,界面目标语言为英文。处理方式是在边界处
维护映射而非在业务逻辑里散落翻译:
精确匹配表:完整状态文案 → 英文
关键短语规则:含"开小差" → "Server rejected the payload"
含"已在其他设备登录" → "Device limit reached: ..."
规则类型表:type 9/1/7/11/12 → 对应英文规则名
学期名规则:第一学期 → Fall,第二学期 → Spring
顺序很重要:先精确匹配再包含匹配,否则”成功”二字会把
“太棒了,恭喜您完成了跑步,成功挑战了自己!”误判为 “OK”。这类缺陷在本项目中
真实出现过,并由一条单元测试(输出不得含 CJK)固定下来——文案规范应当是
机器约束,而非 review 纪律。
10.3 Clean view:为截图而设计的呈现模式
需求侧有一个具体场景:用户需要截取干净的数据视图。实现上做了一个 body.clean
类切换:隐藏侧栏与标题,卡片放大,地图增高,切换按钮以悬浮胶囊形式保留。
两个工程细节值得记录:
- 退出通道必须永存:首次实现把标题栏整体隐藏,导致退出按钮一并消失,
用户被困在呈现模式里(刷新可解)。修复方式是把按钮从被隐藏的容器中解耦,
并追加Esc键退出——任何全屏/沉浸模式都必须保留至少两个退出路径; - 回归验证用坐标命中测试而非肉眼:用
document.elementFromPoint在原侧栏
位置与标题位置取样,断言命中地图容器,即可机器化地证明”确实隐藏了”。
截图可能因缓存而失真,坐标命中不会。
10.4 视觉风格的取舍
初版采用暗色玻璃拟态 + 渐变 + 发光 + emoji,被评价为”模板感”。终版改为浅色、
系统字体、单一强调色、1px 描边、无装饰性效果——本质上是把”设计感”从视觉装饰
转移到信息结构(字重层次、分隔线、表格化数据)。对工具类界面而言,这条路径的
截图效果反而更接近”真实产品”。
11 防御性建议(含检测形式)
本研究观察到的弱点几乎全部指向同一根因:服务端信任了客户端的声明。
以下每条给出触发证据与可执行形式。
11.1 服务端权威计算成绩字段
- 证据:
totalDis / totalTime / totalSteps / speed / complete全部由客户端
声明,服务端仅与学校阈值比较,不与轨迹重算比对; - 做法:以轨迹重算里程与时长(Haversine 累加、时间轴单调性),与声明值比较,
偏差超阈值即拒绝或标记;speed直接由totalTime / totalDis计算,不接受上行值。
11.2 轨迹统计特征检测
| 特征 | 形式 | 直觉 |
|---|---|---|
| 分桶距离变异系数 | CV = σ(d_i)/μ(d_i) |
真实跑步 CV 通常 > 0.15,均匀分桶 CV ≈ 0 |
| 曲率方差 | 相邻三点夹角序列的方差 | 完美椭圆方差极小 |
| 起止点跨记录重复 | 不同记录首坐标距离 < 1 m 且同账号 | 固定起点是强关联信号 |
| 瞬时速度熵 | 速度直方图香农熵 | 过窄分布可疑 |
| 时序自洽 | stop − start == duration 且轨迹点时间覆盖全区间 |
伪造常在此处露出 |
| GPS 精度字段分布 | accuracy 序列是否恒定 | 恒定 3.0 m 不是真实分布 |
表 7:低成本、高区分度的检测特征。
11.3 错误可观测性分级
4xx/10xxx 客户端错误:格式、类型、必填缺失 ← 可对外
5xx/20xxx 服务端内部异常 ← 仅日志
业务码 规则不通过(配速/里程/时段) ← 可对外,带原因
这不是”给攻击者递刀”:差分调试的核心信号是类别跳变,而细分后的类别恰恰
要求攻击者用更多次实验去区分;真正的收益在于运维能定位问题。
11.4 关联配置的前置校验
- 证据:计分目标、跑步计划、日历等关联数据缺失时,保存入口处的空引用会逃逸
到统一捕获层,表现为”服务器开小差”且跳过全部下游校验; - 做法:写路径前置
null检查并返回业务错误码;启动期/定时任务扫描”配置为
空的活跃用户”,把被动 500 变成主动巡检。
11.5 嵌套载荷的防御性反序列化
多层嵌套结构(列表外包配置对象)应显式校验结构类型;类型不符时返回 400 而非
抛异常逃逸。反序列化入口的异常必须被捕获并归类,不允许穿透到统一错误中间件——
异常穿透的代价是”所有下游校验被跳过”(命题 1),这是本研究中最具教学价值的
一条。
11.6 会话、设备与风控
- 并发登录、高频登出重登、设备指纹复用进入风控;
- 为人工申诉保留通道(该产品已具备成绩申诉机制,方向正确);
- 对同一账号短时间内多条轨迹的几何相似度做横向比对(椭圆重合度、长度归一化后
的 Fréchet 距离)。
12 伦理、法律与负责任披露
- 授权与最小权限:本研究仅使用作者本人授权账号,所有写操作经账号所有者
明确同意;实验遵循”单批次单登录、用后登出”的最小干扰原则; - 不得用于规避评价:该类工具不应被用于绕过学校体育考核或其他第三方评价——
那是对评价体系的欺诈,与技术能力无关,本文不提供任何可直接操作的载荷、密钥
或接口清单; - 合规义务:逆向、复刻与公开发布须遵守服务条款、当地法律及软件保护的
相关规定;商业验证码产品的绕过细节不宜公开扩散; - 差异化披露:公开方法论、测试工程与防御建议;隐去签名算法、字段全集、
接口路径与载荷形态。判断标准是:读者能否据此在半小时内构造一次成功提交; - 披露顺序:第 11.3、11.4、11.5 条涉及的具体缺陷建议先通报厂商修复,
再公开防御部分;披露文档应包含可复现的最小触发条件与建议修复位点。
13 局限性
任何诚实的技术报告都必须交代边界。本工作的局限包括:
- 样本规模:单学校、单账号、单版本(v7.x)。结论对其他学校配置、其他
版本的可迁移性未经验证; - 缺少对照客户端:未能在受控条件下抓取官方 App 自身的成功提交报文做逐字段
比对(需要第二台设备与中间人环境),因此部分字段语义依赖间接推断; - 聚合口径未定论:”逐条相加 3.53 km 与聚合回显 3.00 km”的差异仅被记录为
待观察项,未定位到根因; - 反作弊深度有限:本文只验证了”能提交且被判有效”,未对服务端是否存在
异步风控、离线复核做压力测试,也不打算做; - 厂商未确认:防御建议基于观察推导,尚未获得厂商的反馈或修复确认。
14 结论
黑盒协议复刻的难点从来不在于”加密多强”,而在于可观测性工程。本文的做法
可以压成五条:
- 当服务端只会说”失败”时,把失败折叠成的类别做成稳定探针,用单变量二分与
回退复验控制归因(RQ1); - 当语义不可见时,用类型探测定格式、用独立样本定量纲、用对照组闭环(RQ2);
- 当坐标系有系统性偏移时,选定唯一建模系、只在边界转换,并把服务端校验前置为
本地预检(RQ3); - 当随机性是必需品时,让它服从守恒式与规则内点,否则它本身就是异常特征
(命题 4); - 当结果难以判定时,把服务端自己回显的规则结果当作 Oracle,把”返回成功”升级为
“五条规则全过且聚合口径增长”(RQ4)。
对防御方,结论同样清晰:不要信任客户端的声明,不要折叠异常,不要让关键配置的
缺失变成空引用。 这三条比任何签名算法都更接近安全的本质——因为签名只保证
“话没被改”,不保证”话是真的”。
参考文献
[1] Inverx. com.zjwh.android_wh_physicalfitness-V7.3.40 [EB/OL]. GitHub.
[2] Desgard. 玩转运动世界校园 [EB/OL]. 一片瓜田, 2016.
[3] faldict. 运动世界校园刷跑步记录脚本 [EB/OL]. 2016.
[4] 网易易盾. NetSecKit 移动应用安全组件 [EB/OL].
[5] GeeTest. GeeTest v4 滑块验证方案 [EB/OL].
[6] OpenStreetMap Foundation. Tiles Usage Policy [EB/OL].
[7] OWASP. API Security Top 10 [EB/OL].
[8] WGS-84 与 GCJ-02 坐标系差异及双向转换实践 [EB/OL].
[9] LiuYiGL. RunWorldSchoolMod — 运动世界校园辅助模块 [EB/OL]. GitHub.
[10] FengLi666. sports — 跑步数据生成脚本 [EB/OL]. GitHub.
Part II · English Version
1 Introduction
1.1 Background
The business loop of location-based campus running applications can be summarized as:
the phone collects a trajectory; the client assembles a performance record; the server
verifies a signature, applies institutional rules, persists the row, and aggregates
statistics. On this loop, every element of a performance record is declared by the
client: distance, duration, step count, pace, checkpoint outcomes, and even the
boolean “this record is complete”. The server’s responsibility is compressed into three
tasks: verify the signature, apply the school’s configuration rules, and maintain the
aggregation.
From a security-engineering perspective, this is a textbook case of a misplaced
trust boundary. This paper does not discuss how the system ought to be designed
(that is Section 11), but instead answers an empirical question: given only an ordinary
account and packet capture, what methods are required to replicate the protocol
surface, what observability obstacles arise, and how are those obstacles resolved by
engineering means?
1.2 Problem Statement: A Deficient Trust Model
| Element | Declared by | Verified by | Observed reality |
|---|---|---|---|
| Trajectory coordinates | client | no recomputation | persisted / uploaded as-is |
| Distance, duration, steps | client | compared with thresholds | never cross-checked against the trajectory |
| Pace | client | compared with configured range | scaled by 10⁻³ at persistence |
| Checkpoint outcomes | client | compared with point configuration | decided by client booleans |
| Record completeness flag | client | none | trusted directly |
| Signature | client | recomputed server-side | the only strong check |
The conclusion of Table 1 is self-evident: the only strong check protects against
transit tampering, not against false content. A signature answers “was the message
altered in transit”, never “is the message true”.
1.3 Research Questions
- RQ1 (Observability): When the server collapses parse failures, validation
failures, and internal exceptions into one canned message, can an attributable
differential-debugging framework be built from the return category alone? - RQ2 (Semantic identification): Without source access, can field types, physical
units, and canonicalization rules be systematically inferred? - RQ3 (Geometric credibility): How can trajectories be generated under dual
coordinate systems and polygonal geofences, while keeping randomization free of any
conservation violation? - RQ4 (Verifiability): How can “preview equals upload” become a replayable
invariant, and how can the server’s own rule echo serve as an acceptance oracle?
1.4 Contributions
- A differential-debugging procedure driven by a three-way error taxonomy
(RQ1): the canned message is converted into a stable boolean probe, with an
experiment matrix, a chain of inferences, and three methodological propositions; - A black-box method for field semantics (RQ2): type probing, unit back-derivation,
and null/boolean canonicalization, with worked measurements; - A seeded two-stage generation architecture and a randomization governance
principle (RQ3, RQ4): preview and upload share random-source state, making
“preview equals upload with zero side effects” a testable property; randomization
may alter distributions but never totals; - A validation regime using the server’s rule echo as an oracle: 24 unit tests
plus three end-to-end closure records (valid distance growing from 0 to 3.0 km); - Six defensive recommendations, each with the evidence that motivated it and an
executable detection formulation.
1.5 Organization
Section 2 reviews related work. Section 3 presents the system overview, threat model,
and experimental discipline. Sections 4–7 answer RQ1–RQ4 respectively. Sections 8 and
9 cover session management and validation. Section 10 describes the console
implementation. Section 11 lists defensive recommendations; Section 12 discusses ethics
and disclosure; Section 13 states limitations; Section 14 concludes.
2 Related Work
2.1 Early Public Analyses (circa 2016)
The protocol surface of this product was first analyzed publicly around 2016 [2][3].
Those efforts were capture-driven and reached three conclusions: transport was
plaintext JSON; authentication used a decodable Basic credential; the server performed
almost no cross-validation of performance fields, so a client could declare arbitrary
distance and duration. They established the “client-declared performance” premise, but
the field structure and signature mechanism of that era differ substantially from the
modern version.
2.2 Community Tooling and Runtime Modules
Subsequent community work falls into two categories: runtime modification hosted in
the operating system (for example, an Xposed module that reverts ordered checkpoints
to unordered ones [9]), which requires root and a framework environment; and standalone
client implementations such as an open-source pure-Python replica [1] and assorted
scripts [3], which avoid the host environment but typically cover only some sport
types, and generally note—without explanation—that scored runs require an existing
running plan on the account. This work is closest to [1]; the difference is that
methodology (attribution discipline in differential debugging, dimensional closure
in semantic identification, deterministic testing) is treated as the primary object of
contribution, complemented by session lifecycle, side-effect tiering, and a defensive
perspective.
2.3 General Approaches to Mobile Anti-Abuse
Mature consensus exists in anti-cheat and anti-automation: server-authoritative
computation, statistical anomaly detection, layered error observability, and device /
session risk control [7]. In this class of products, however, landing those practices
is weak—the direct trust of client-declared performance fields is the canonical
deviation. Section 11 maps each observed weakness onto that consensus with concrete
detection forms.
2.4 Coordinate Systems and Geofencing
The systematic offset between WGS84 and GCJ-02, and their conversion, are common
knowledge in Chinese geospatial engineering [8]. In protocol replication, however, it
manifests as a silent offset: local checks pass while the server judges the record
out of bounds, and neither side raises an error. Section 6 treats coordinate discipline
as a first-class constraint and reports a real counterexample.
3 System Overview, Threat Model, and Experimental Discipline
3.1 Layered Architecture
┌───────────────────────────────────────────────────────────┐
│ Web Console stdlib HTTP + single-page front end │
│ side-effect tiering: plan / submit │
├───────────────────────────────────────────────────────────┤
│ Run Planner style → distance / pace / cadence │
│ distribution; backdated stop inside the │
│ valid window → serializable plan object │
├───────────────────────────────────────────────────────────┤
│ Protocol Core payload modeling · normalization · salted │
│ digest · geometry · 10-second buckets │
├───────────────────────────────────────────────────────────┤
│ Transport session & device headers · pre-signed OBS │
├───────────────────────────────────────────────────────────┤
│ Challenge GeeTest v4: render → locate → drag → send │
└───────────────────────────────────────────────────────────┘
Figure 1: five layers with strictly downward dependencies. Protocol Core holds no
session state, which is what allows offline unit testing.
3.2 Submission Data Flow
Five round trips: challenge check and solving (①②), pre-signed object upload of eleven
gzip+base64 segments (③), record save with two fields injected after signing (④),
and a detail read-back returning per-rule verdicts plus aggregation (⑤). The two
post-signature fields are deliberately excluded from the signed set, which reveals that
the signature set is an explicit inventory that must be replicated exactly—not “every
field”.
3.3 Threat Model
| Dimension | Setting |
|---|---|
| Attacker capability | valid account; packet capture; offline analysis of APK strings and DEX; arbitrary local code |
| Attacker lacks | server source, logs, database, key material, runtime instrumentation |
| Assumptions | account owner authorization; network reachability; readable institutional rule configuration |
| Goal | fully characterize the protocol surface and build a conformant client to evaluate the trust model |
3.4 Experimental Discipline
The dominant risk in black-box research is contamination, not failure. The rules
enforced here: one login per batch with immediate logout; strict separation of read-only
and side-effecting probes; explicit owner consent before any account-state mutation;
persisting raw responses per batch; and a known-behavior control for every new
variable, without which category jumps cannot be attributed.
4 Method (RQ1): Three-Way Error Taxonomy and Single-Variable Debugging
4.1 Formalization
Define the observation function f : X → {E_sig, E_biz, E_ok} where E_sig denotes
signature mismatch (normalization/type error), E_biz denotes a server exception
swallowed into the canned message (content/context error), and E_ok denotes success.
For operations the collapse of f is an observability defect; for differential
experiments it is a stable boolean probe: any edit that changes the category must
land on the first branch executed after signature verification.
The experimental task is therefore: find x such that f(x) = E_ok, where the
E_biz → E_ok jump is a necessary and sufficient indicator of entering the correct
branch.
4.2 Three Design Principles
P-I Strict single variable. All other fields are frozen within a batch.
P-II Mandatory control. Every setting is paired with a known control (for example
another sport type that is expected to succeed) to exclude environmental drift such
as window changes, session expiry, or rate limiting.
P-III Reversal check. Every category jump must be reproduced in reverse by
restoring the variable; jumps that do not reverse are discarded, because most
spurious jumps (environmental drift, throttling windows) are not reversible.
4.3 Experiment Matrix
Table 3 reports a real batch of more than twenty independent variables, ten of which
are shown. All but the last row land in E_biz: time window, distance extremes,
duration and step extremes, bucket content, geographic position inside/outside the
fence, session identifier placement, null-valued goal/theme fields, numeric types,
serialization shape of one nested payload (bare collection versus wrapper object),
and a control sport type that returns E_ok.
The final row is the conclusion row: the same field with two serialization shapes
produces the category jump, which means the server deserializes that field into a
wrapper type at the entry point right after dispatching by sport type. The bare
collection triggers a null pointer, the exception escapes to the unified catch, and
every downstream check—time, distance, fence, checkpoints—is skipped. This is exactly
why rows 1–5 were inert: they were never executed.
4.4 Propositions
Proposition 1 (Exception bypass). If a deserialization entry point lacks type
validation and exceptions are caught by a unified error middleware, then every
business rule located after that entry point forms a bypass: mutating the inputs
of those rules does not change the observation. Consequently, “nothing I change has
any effect” is itself strong evidence of an exception at the entry point.
Proposition 2 (Category as signal). When failures are collapsed into a single
message, the return category still constitutes a boolean probe; the information
content of the observation has a non-zero lower bound.
Proposition 3 (Explainability criterion). An exception-source hypothesis is adopted
only if it explains all rows of the historical matrix. The hypothesis adopted here
explains why more than twenty rows all returned E_biz, which is why it is stronger
than “a working recipe found by trial”.
5 Method (RQ2): Black-Box Identification of Field Semantics
Passing the signature only proves format agreement, never semantic correctness.
5.1 Type Probing
Inject a floating-point value into a suspected integer field. If the signature layer
rejects it with E_sig while the equivalent integer passes, the deserialization target
is an integer: the server normalizes the value after deserialization and before signing,
so a locally formatted float cannot match. Type probing can be batched to produce a
black-box “field → type” table with two requests per field.
5.2 Unit Back-Derivation
Observation chain: (1) the client sends 2381; (2) the detail view echoes 2.381;
(3) the persistence scale is therefore k = 10⁻³; (4) an independent record is used
for dimensional closure—3450 m in 1257 s gives 1257 / 3.45 = 364.3 s/km =
6.07 min/km, exactly the echoed value. The field is thus pace in minutes per
kilometer, transmitted multiplied by 1000.
Steps (1)–(2) alone only establish the scale; step (4) fixes the dimension. This is
the role of a control sample in semantic identification: a scale factor can be
guessed, a dimension must be closed with an independent sample. It also explains an
earlier defect: constructing the field as distance / duration × 1000 yields m/s×1000
rather than min/km×1000—signature-valid, semantically wrong, and guaranteed to fail
the pace rule. This is the canonical “format correct, semantics wrong” sample, which
no signature check can catch and only the rule echo (Section 9) exposes.
5.3 Null, Boolean, and String Canonicalization
String.valueOf(null) on the Java side is the primary source of signature divergence.
Three forms must be distinguished: an omitted field (not signed), a JSON null
(signed as "null"), and the string "null" (equivalent to the previous). Booleans
are lower-cased, and integer 1 never matches string "1". These rules are
centralized in a single pure function so that unit tests can enumerate them.
5.4 The Dual Role of Timestamps
Millisecond timestamps must be probed across representations (epoch, ISO, date) to
determine which enters the signed set. They carry two roles at once—the business
interval and the geometric sampling base—so they must satisfy stop − start = duration
with monotone point timestamps inside the interval; any inconsistency surfaces at
detail read-back or map rendering.
5.5 Summary
Signature consistency is necessary but not sufficient: type, dimension, and
canonicalization must each be verified independently, and dimensional verification
requires closure with an independent sample.
6 Method (RQ3): Trajectory Generation under Dual Coordinates and Geofences
6.1 Coordinate Discipline
Map tiles use WGS84 while domestic providers and mobile SDKs emit GCJ-02, differing by
hundreds of meters. The single rule adopted here: choose one modeling coordinate
system and convert only at system boundaries. Fence data is inverse-transformed on
read; trajectories are generated and pre-checked in WGS84; conversion back to GCJ-02
happens only when the upload payload is assembled; the preview map stays in WGS84 to
align with tiles.
6.2 Inverse by Fixed-Point Iteration
There is no closed-form inverse, so w₀ = g, wₙ₊₁ = wₙ − (F(wₙ) − g) is iterated
six times to a threshold of 10⁻⁷ degrees (about 1 cm); three iterations already
converge in practice, six are defensive redundancy. A counterexample from an earlier
revision compared GCJ-02 fence vertices against WGS84-generated points, producing a
silent 600 m offset: local pre-checks passed while the stored record was out of
bounds, on both sides without an exception. Such errors surface only when both
coordinate spaces are written into the same validation table.
6.3 Local Geofence Pre-Check
For a 6- to 19-vertex polygon, every point must satisfy
pᵢ ∈ interior(P) ∧ dist(pᵢ, ∂P) ≥ δ with δ = 3 m, using a ray-casting test plus a
per-edge distance computation. Parameter combinations failing the check are rejected
locally, moving a server-side check to the client and eliminating the expensive
“upload first, learn later” loop.
6.4 Arc-Length Parameterization of a Track Shape
The trajectory is a stadium (two straights plus two semicircles), not a circle: a
closed loop with an accumulated arc-length table, a lap of about 301 m, and a bearing
θ for the long axis. point_at(s) binary-searches the table, interpolates, and applies
a 0.95 inset; points sample s = (phase + v·t) mod lap and receive bounded
triangular jitter (±2.5 m) for indices greater than zero. Three design points address
recognition risks: the phase randomizes the start (a fixed start repeats to six
decimals across records), the inset reserves room for noise so edge-hugging centerlines
stay inside, and the triangular distribution is bounded and unimodal—Gaussian noise is
unbounded and would eventually leave the fence.
6.5 Bucket Randomization and Conservation
Instantaneous speed breathes around the target mean; ten-second distance and step
shares are weighted by 1 + U(−8%, +8%) and then renormalized so that
Σ stepsBucket = totalSteps and Σ distBucket = totalDis hold exactly. Integer step
shares use largest-remainder rounding. Proposition 4 (Conservation first):
randomization may only change the distribution shape, never the total; any diversity
that breaks conservation converts randomization itself into a new anomaly signature.
6.6 Checkpoints
Five checkpoints sample arc fractions k/6 + U(−0.04, +0.04) (k = 1..5), sorted and
mapped onto actual trajectory points, which guarantees that every checkpoint coordinate
appears in the full track—a property guarded by a unit test. In sequence mode the
position index equals travel order, so “passed all checkpoints in order” evaluates to
true.
7 Method (RQ4): Deterministic Simulation, Randomization Governance, and Testing
7.1 The Plan Object
A run is serialized as a plan containing style, distance, duration, steps, stop time,
zone index, random seed, phase fraction, checkpoint fractions, goal identifier, target
distance, and average power. Every source of randomness is explicit so that the same
plan replays identically at any later time.
7.2 Random-Source Boundaries and Two Properties
Preview and upload are separate code paths that both execute random.seed(plan.seed)
before invoking the same generator with the same phase and checkpoints; absent any
interleaved consumption, outputs are point-wise identical. This yields:
- P1 (Determinism): two consecutive
build()calls on one plan produce equal
coordinate sequences; - P2 (No side effects): the preview endpoint performs no network I/O and no
persistence, and may be invoked without limit.
The trap behind P1 is that uuid4 and timestamps must not share the geometric random
source, or the two sequences fall out of alignment; identifiers therefore use
os.urandom. Random-source boundaries must be declared explicitly, otherwise
determinism cannot hold.
7.3 Randomization Governance
Every random variable is confined to the interior of the rule envelope with explicit
margins: distance in [1000, 3000] m against a daily cap of 3000 m; pace in
(5.0, 10.0) min/km with 0.3–0.4 reserved at both ends; cadence in [60, 300] spm with
samples drawn from 152–190; stop time inside the valid window with four minutes
reserved at both ends; phase in U(0,1); bucket jitter at ±8% under renormalization.
The pace margin is derived from worst-case segment analysis: with ±8% bucket jitter,
5.4 × 0.92 ≈ 4.97 would touch the lower bound, so the fast style starts at 5.45.
Compute the worst case first, then choose the parameter.
7.4 Test Regime: 24 Unit Tests
Four groups guard three regression classes: protocol semantics (signature and
payload), geometry (containment, repetition), and interface language (deleted
translations). The language test asserts directly that outputs contain no CJK
codepoints, converting a copy rule from review discipline into a machine constraint.
7.5 Property-Based Testing
For a component whose inputs are random and whose outputs must be legal, example-based
assertions are meaningless; a property test samples the planner 300 times and asserts
distance, pace, duration, and cadence bounds. The assertion block doubles as
documentation: it states precisely what “a legal running plan” means.
8 Sessions, Device Binding, and Side-Effect Management
8.1 Three-Stage Sign-In
Status check → challenge solving (skippable) → credential exchange via HTTP Basic,
which returns a token and simultaneously performs device binding. The challenge
averages 15 s and must be surfaced as asynchronous progress; credentials are stored in
a session object shared by all subsequent requests; failures must distinguish wrong
credentials, device protection, and challenge failure.
8.2 Session State Machine
Four states: logged_out → pending → active → released, with a protection state in
which a second sign-in from elsewhere locks the session until the original device
signs out or one hour elapses. For automation this is a hard constraint: every
successful login registers a pending-logout marker released in finally; probe scripts
print the logout code as a first-class result; the console surfaces logout failure
because it means the account is unusable for a period. One missed logout in this
project cost a full hour of experiments, which is the counter-evidence for
“distinguishable failures”.
8.3 Side-Effect Tiering
/api/plan performs pure local construction with zero side effects; /api/submit
writes to the database and object storage and requires an active session; login and
logout manage the session explicitly. A real defect is worth recording: a helper that
implicitly chose GET versus POST based on the presence of a body caused the sign-out
button to emit GET, receive 404, and fail silently. Side effects belong in the
function signature, not in calling conventions.
9 Validation: Using the Server’s Verdict as an Oracle
9.1 The Rule Echo
Record detail returns per-rule verdicts—type id, human-readable rule, pass/fail—so a
submission is accepted only if all five rules are ok and status = 0.
9.2 The Oracle in Attribution
During the repair cycle, the oracle narrowed the failure to rule 12 (pace): the unit
error put the value outside the interval while the other four rules passed. Without
the per-rule echo, “record invalid” would have been an unattributable black box.
The oracle’s value lies not in judging success, but in localizing failure.
9.3 Status Machine and Aggregation Semantics
Record level: status ∈ {0, 1}; only status = 0 contributes to valid distance,
while total distance counts everything. A discrepancy between per-record summation
(3.53 km) and aggregated echo (3.00 km) was observed and recorded as an open item,
with the principle: defer to the server’s aggregation, never recompute locally.
9.4 End-to-End Acceptance and Closure Evidence
Acceptance is a full sequence—plan → build(seed) → object upload → save → detail
read-back with five green rules → aggregation growth → logout. Evidence consists of
three status = 0 records and valid distance growing from 0 to 3.0 km, one of them
through the complete console path (backdated stop, random style, map preview), one
through the command-line path, and one exercising the fence-plus-backdate combination.
10 Implementation: The Web Console
10.1 Zero-Dependency Choice
Python’s stdlib multi-threaded http.server plus a single-file front end. Rationale:
minimize the attack surface of a local tool bound to loopback, maximize reproducibility
on other machines, and avoid framework overhead for fewer than ten endpoints. Session
and rule caches live in process memory; no credential is ever written to disk.
10.2 Internationalization and the Ordering Trap
Translations are maintained at the boundary in an explicit table: exact status phrases
first, then key-phrase rules, then rule-type mappings, then semester-name rewriting.
Order matters: exact before substring, otherwise the word “success” inside a longer
congratulation message collapses to “OK”. This defect occurred in this project and is
now pinned by a unit test asserting no CJK output. Copy rules should be machine
constraints, not review discipline.
10.3 Clean View: A Presentation Mode Designed for Screenshots
A body.clean class hides the side panel and title, enlarges cards, and raises the
map. Two lessons: (1) the exit control must survive—hiding the whole header trapped
users in presentation mode, so the button was decoupled from the hidden container and
an Esc binding was added; every immersive mode needs at least two exit paths;
(2) hiding must be verified by coordinate hit-testing (document.elementFromPoint
against the former panel and title positions must hit the map), not by screenshots,
which can be served stale.
10.4 Visual Style
The first version used dark glassmorphism with gradients, glow, and emoji and was
judged “template-like”. The final version uses a light background, system typography,
a single accent color, 1px borders, and no decorative effects, moving perceived
quality from ornament to information structure.
11 Defensive Recommendations
Every observed weakness traces to one root cause: the server trusts client
declarations.
11.1 Server-authoritative computation. Recompute distance and duration from the
trajectory (Haversine accumulation, monotone time axis) and compare against declared
values; derive pace from duration / distance instead of accepting the upload.
11.2 Statistical trajectory detection. Table 7 lists six low-cost features:
coefficient of variation of bucket distances (real runs typically > 0.15, uniform
buckets ≈ 0), curvature variance, cross-record start-point recurrence (< 1 m for the
same account), entropy of the instantaneous speed histogram, time-axis self-consistency
(stop − start == duration with points covering the interval), and whether the GPS
accuracy field is suspiciously constant.
11.3 Layered error observability. Separate client errors (4xx), internal faults
(5xx, log only), and rule outcomes (business codes with reasons). This is not handing
attackers a map: the core signal for differential debugging is the category jump, and
finer categories cost the attacker more experiments while giving operations real
diagnosis.
11.4 Precondition checks for relational configuration. Missing goal/plan/calendar
rows currently produce null dereferences at the save entry that escape to the unified
catch and bypass all downstream validation. Return explicit business codes and run
scheduled scans for active users with empty configuration.
11.5 Defensive deserialization of nested payloads. Validate the structural type of
wrapper objects explicitly and return 400 on mismatch; never let exceptions escape to
the unified middleware. Exception escape = validation bypass (Proposition 1).
11.6 Session, device, and risk control. Monitor concurrent sign-ins, rapid
sign-out/sign-in cycling, and device-fingerprint reuse; keep a human appeal channel
(the product already has one); compare geometric similarity across records of one
account (Fréchet distance after length normalization).
12 Ethics, Legal Boundaries, and Responsible Disclosure
- Authorization and least privilege: only personally authorized accounts, explicit
owner consent for writes, one login per batch with immediate logout; - Not for evading evaluation: the tool must not be used to bypass institutional
assessment—this is fraud against the evaluation system regardless of technical
merit, and this paper ships no directly usable payloads, keys, or endpoint lists; - Compliance with terms of service, applicable law, and software-protection rules;
CAPTCHA-bypass details of commercial products should not be spread publicly; - Differential disclosure: publish methodology, testing engineering, and
defensive findings; withhold signature algorithm, field inventory, endpoint paths,
and payload shapes. The test is: could a reader construct one successful
submission within half an hour? - Disclosure order: report the specific defects behind Sections 11.3–11.5 to the
vendor first, with reproducible minimal triggers and suggested fix locations, and
publish the defensive section afterwards.
13 Limitations
- Sample size: one school, one account, one version (v7.x); transferability to
other institutional configurations and versions is unverified; - No reference capture: a controlled, field-by-field comparison against a genuine
official-client submission was not obtained (it requires a second device and a
man-in-the-middle setup), so some field semantics rest on indirect inference; - Aggregation semantics unresolved: the 3.53 km versus 3.00 km discrepancy is
logged as an open item, not root-caused; - Limited anti-cheat depth: this work establishes “submits and is judged valid”
and deliberately does not stress asynchronous risk control or offline re-auditing; - No vendor confirmation: defensive recommendations are derived from observation
and await vendor feedback.
14 Conclusion
The difficulty of black-box protocol replication is never “how strong is the
encryption” but observability engineering. The method compresses into five rules:
turn the collapsed failure category into a stable probe with single-variable bisection
and reversal checks (RQ1); infer types by probing, dimensions by independent samples,
closure by controls (RQ2); pick one coordinate system and convert only at boundaries
while front-loading server checks as local pre-checks (RQ3); confine randomization to
conservation identities and rule interiors, or it becomes the anomaly itself
(Proposition 4); and promote “returned success” to “all five rules green plus
aggregation growth” using the server’s own echo (RQ4).
For defenders the conclusion is equally short: do not trust client declarations, do
not collapse exceptions, do not let missing configuration become null references.
None of these require a stronger signature—because a signature only guarantees the
words were not changed, not that the words are true.
References
[1] Inverx. com.zjwh.android_wh_physicalfitness-V7.3.40 [EB/OL]. GitHub.
[2] Desgard. 玩转运动世界校园 [EB/OL]. 一片瓜田, 2016.
[3] faldict. 运动世界校园刷跑步记录脚本 [EB/OL]. 2016.
[4] 网易易盾. NetSecKit 移动应用安全组件 [EB/OL].
[5] GeeTest. GeeTest v4 滑块验证方案 [EB/OL].
[6] OpenStreetMap Foundation. Tiles Usage Policy [EB/OL].
[7] OWASP. API Security Top 10 [EB/OL].
[8] WGS-84 与 GCJ-02 坐标系差异及双向转换实践 [EB/OL].
[9] LiuYiGL. RunWorldSchoolMod — 运动世界校园辅助模块 [EB/OL]. GitHub.
[10] FengLi666. sports — 跑步数据生成脚本 [EB/OL]. GitHub.















