ProCreations commited on
Commit
0ab38d4
·
verified ·
1 Parent(s): 61731db

Add live-site benchmark (147 rendered pages) and chart to the card

Browse files
Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +39 -2
  3. assets/zap-benchmark-doodle.png +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ assets/zap-benchmark-doodle.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -39,6 +39,10 @@ So your autofill can answer questions like *"is this a login form?"*, *"which bo
39
  ¹ A baseline of the kind password managers ship: it trusts a valid `autocomplete`, then the input `type`, then multilingual keyword regexes over name/id/placeholder/label/nearby text.
40
  ² On 1,622 test fields where the site's `autocomplete` token agrees with the reference label, the token is deleted from the model input. The score is how often Zap still predicts it. This check doesn't depend on the LLM teacher.
41
 
 
 
 
 
42
  ## Quick start (Python)
43
 
44
  ```python
@@ -149,6 +153,37 @@ Real payment-card forms are very rare in a static crawl: the real test has 8 pay
149
 
150
  The full reports are in `eval/`, including per-class precision/recall and confusions for bf16 and fp32 weights and the heuristic baseline.
151
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
152
  ## How it was made
153
 
154
  **Data.** 62 random WARC files from Common Crawl `CC-MAIN-2026-39` (September 2026) were streamed in full, one request per file. That's 998,946 pages with at least one fillable field, 2.30M forms, and **948,742 unique forms** after deduplicating by form signature. Pages are re-serialized through an HTML5-spec parser (lexbor) before extraction so the tree matches what a browser builds. Fields outside any `<form>` are grouped the way a single-page app lays them out (nearest ancestor with a button).
@@ -175,9 +210,10 @@ The published weights are bf16. The whole project took about 6 hours of workstat
175
  ## Limitations
176
 
177
  * Zap reads the DOM text and attributes, not pixels or CSS. In a browser, drop invisible inputs (honeypots) before classifying, e.g. with `offsetParent === null` or `getComputedStyle`. The training crawl only saw static HTML, where some CSS-hidden honeypots remained, so Zap usually labels those `other`, but not always.
178
- * The training pages come from a static crawl, so heavily scripted flows (identifier-first sign-in, checkout iframes) are covered mainly by synthetic examples. Real-world accuracy on payment forms is not well measured (8 test forms).
 
179
  * The labels come from an LLM teacher. It is audited, but it isn't human annotation, so expect a small amount of label noise, especially on ambiguous `name`/`given-name` and `search`/`other` cases.
180
- * Fields inside cross-origin iframes (e.g. hosted card fields) are invisible to any content script, including `zap-extract.js`.
181
 
182
  ## Files
183
 
@@ -189,6 +225,7 @@ The published weights are bf16. The whole project took about 6 hours of workstat
189
  | `zap_extract.py` / `zap-extract.js` | form extractor (Python canonical / DOM port) |
190
  | `zap_infer.py` / `zap-infer.js` | windowing + inference helpers |
191
  | `eval/` | evaluation reports, export parity, dataset statistics |
 
192
  | `training/` | the code used to harvest, label, synthesize, train and export |
193
 
194
  ## License
 
39
  ¹ A baseline of the kind password managers ship: it trusts a valid `autocomplete`, then the input `type`, then multilingual keyword regexes over name/id/placeholder/label/nearby text.
40
  ² On 1,622 test fields where the site's `autocomplete` token agrees with the reference label, the token is deleted from the model input. The score is how often Zap still predicts it. This check doesn't depend on the LLM teacher.
41
 
42
+ **Live websites:** 147 real login, sign-up and password-reset pages rendered in Chromium. Every visible field was labeled blind from a marked screenshot, and Zap's predictions were compared with the same keyword heuristic ([details](#live-site-benchmark)):
43
+
44
+ ![Zap vs keyword rules on 147 live login, sign-up and reset pages](https://huggingface.co/ProCreations/Zap/resolve/main/assets/zap-benchmark-doodle.png)
45
+
46
  ## Quick start (Python)
47
 
48
  ```python
 
153
 
154
  The full reports are in `eval/`, including per-class precision/recall and confusions for bf16 and fp32 weights and the heuristic baseline.
155
 
156
+ ### Live-site benchmark
157
+
158
+ This test checks Zap outside the static crawl, on the pages a password manager actually meets.
159
+
160
+ **Setup:**
161
+ - 207 login, sign-up and password-reset URLs of popular sites were rendered in headless Chromium 153 in September 2026.
162
+ - `zap-extract.js` ran in every frame, as an extension content script with `all_frames` would. Nothing was typed or submitted.
163
+ - 147 pages showed a visible form: 182 forms and 260 visible fields.
164
+ - Each visible field was labeled from a screenshot with numbered field markers, plus its attributes. The labelers were separate AI labelers, not the Qwen training teacher, and they never saw either system's predictions.
165
+
166
+ | | **Zap** | keyword heuristic¹ |
167
+ |---|---|---|
168
+ | field accuracy | **96.5%** (251/260) | 72.3% (188/260) |
169
+ | form-purpose accuracy | **91.2%** (166/182) | 54.4% (99/182) |
170
+ | login-form F1 (96 login forms) | **0.960** | 0.610 |
171
+ | pages with every field and form right | **90.5%** (133/147) | 29.9% (44/147) |
172
+ | current vs new password (56 password fields) | **96.4%** | 91.1% |
173
+ | password-field detection F1 | 1.000 | 1.000 |
174
+ | login-identifier (`username`/`email`) F1 | **0.993** | 0.977 |
175
+
176
+ Both systems find every password box. The difference is everything around it:
177
+ - **Keyword heuristic:** misses 53 of the 96 login forms. 49 of those are identifier-first steps ("Email or phone", then Next) with no password field yet, and the heuristic calls them `newsletter` or `other`. Zap finds all 96, at precision 0.923.
178
+ - **Where Zap goes wrong:**
179
+ - The Facebook and Instagram sign-up forms get split into one-field groups by the extractor, because every field wrapper holds an icon-only button. That accounts for 4 of Zap's 9 field errors and 8 of its 16 form errors.
180
+ - Sign-up or reset steps that only ask for an email are sometimes read as `login` (Coinbase, Dropbox, Proton) or `newsletter` (Notion, Netflix).
181
+ - **Extractor coverage:** on the same page loads, `zap-extract.js` returned 267 of 282 visible text-like inputs.
182
+ - The 8 light-DOM misses are password boxes the page marks `aria-hidden="true"`, which the extractor skips by design. They are pre-rendered second-step or password-manager-only fields.
183
+ - The other 7 misses are archive.org inputs inside shadow DOM, which the extractor does not enter.
184
+ - Median extraction time in Chromium is 1.7 ms per frame. For model speed, see **Speed** above.
185
+ - **Excluded pages:** about 27 URLs blocked the headless browser or failed to load, and about 12 showed only buttons on the first step (e.g. "Continue with email"). These pages are not scored. The sample is popular, mostly English-language sites.
186
+
187
  ## How it was made
188
 
189
  **Data.** 62 random WARC files from Common Crawl `CC-MAIN-2026-39` (September 2026) were streamed in full, one request per file. That's 998,946 pages with at least one fillable field, 2.30M forms, and **948,742 unique forms** after deduplicating by form signature. Pages are re-serialized through an HTML5-spec parser (lexbor) before extraction so the tree matches what a browser builds. Fields outside any `<form>` are grouped the way a single-page app lays them out (nearest ancestor with a button).
 
210
  ## Limitations
211
 
212
  * Zap reads the DOM text and attributes, not pixels or CSS. In a browser, drop invisible inputs (honeypots) before classifying, e.g. with `offsetParent === null` or `getComputedStyle`. The training crawl only saw static HTML, where some CSS-hidden honeypots remained, so Zap usually labels those `other`, but not always.
213
+ * The training pages come from a static crawl, so heavily scripted flows (identifier-first sign-in, checkout iframes) are covered mainly by synthetic examples. On live sites, email-only first steps of sign-up and reset flows are the most common purpose mistake (see [Live-site benchmark](#live-site-benchmark)). Real-world accuracy on payment forms is not well measured (8 test forms).
214
+ * The extractor does not enter shadow DOM, and it skips `aria-hidden` inputs.
215
  * The labels come from an LLM teacher. It is audited, but it isn't human annotation, so expect a small amount of label noise, especially on ambiguous `name`/`given-name` and `search`/`other` cases.
216
+ * Each frame is extracted on its own. Fields inside iframes, such as hosted card fields or embedded sign-in boxes, need the content script to run with `all_frames: true`, and they are classified without the parent page's form context.
217
 
218
  ## Files
219
 
 
225
  | `zap_extract.py` / `zap-extract.js` | form extractor (Python canonical / DOM port) |
226
  | `zap_infer.py` / `zap-infer.js` | windowing + inference helpers |
227
  | `eval/` | evaluation reports, export parity, dataset statistics |
228
+ | `assets/` | the live-site benchmark chart |
229
  | `training/` | the code used to harvest, label, synthesize, train and export |
230
 
231
  ## License
assets/zap-benchmark-doodle.png ADDED

Git LFS Details

  • SHA256: 58d8f4d4514712d95b6b719d29b67ba2877033ed2af733d8b7035f2b238edc8b
  • Pointer size: 131 Bytes
  • Size of remote file: 446 kB