Xunzhuo commited on
Commit
794c758
·
verified ·
1 Parent(s): 887f39b

Add Hallucination playground with Vela Halu and verified examples

Browse files
README.md CHANGED
@@ -8,7 +8,7 @@ app_port: 7860
8
  header: mini
9
  pinned: false
10
  license: apache-2.0
11
- short_description: Explore ten Vela models for intelligent routing.
12
  models:
13
  - llm-semantic-router/Vela-1.0-Encoder-307M-Domain
14
  - llm-semantic-router/Vela-1.0-Encoder-307M-PII
@@ -16,6 +16,7 @@ models:
16
  - llm-semantic-router/Vela-1.0-Encoder-307M-Safety
17
  - llm-semantic-router/Vela-1.0-Encoder-307M-Hazard
18
  - llm-semantic-router/Vela-1.0-Encoder-307M-FactCheck
 
19
  - llm-semantic-router/Vela-1.0-Encoder-307M-Modality
20
  - llm-semantic-router/Vela-1.0-Encoder-307M-Feedback
21
  - llm-semantic-router/Vela-1.0-Encoder-307M-Embedding
@@ -24,7 +25,7 @@ models:
24
 
25
  # Vela Studio
26
 
27
- A public playground for the [Vela 1.0 model family](https://huggingface.co/collections/llm-semantic-router/vela-10), by [vLLM Semantic Router](https://vllm-sr.ai). Results come from the selected model's pinned checkpoint on CPU. Exact built-in examples can reuse a prior successful inference, clearly marked as a cached example result. The interface includes twelve interactive demos powered by ten model checkpoints: classification, PII highlights and redaction, content-risk thresholds, semantic similarity, dynamic prompt routing, passage reranking, and semantic clustering.
28
 
29
  ## Runtime
30
 
@@ -32,19 +33,19 @@ The Docker image uses Python 3.11, CPU PyTorch 2.8.0, Transformers 4.57.6, and F
32
 
33
  Two 307M FP32 checkpoints require approximately 2.46 GB of parameter memory, plus runtime and activation memory. Execution remains sequential, with a single input or query-passage pair at a time. Set `VELA_MODEL_CACHE_SIZE=1` to retain only one model on smaller hosts; supported values are `1` and `2`, with `2` as the default.
34
 
35
- Most demos accept up to **512 tokens including special tokens**; Clustering accepts **3–24 texts**, each up to **128 tokens including special tokens** and 2,048 characters. Similarity and Reranker allow at most **six candidates**; Routing allows **2–10 categories**. Reranker applies the limit to each query-and-passage pair; Embedding applies it independently to each text. Inputs exceeding this limit are rejected without truncation. The model family has a larger declared input capacity; this shorter public limit keeps CPU Basic usage bounded.
36
 
37
  One inference worker serves a FIFO queue with up to **eight waiting requests** in addition to the running request. The interface shows queue position, model loading, and inference progress and allows cancellation. HTTP 429 is returned only when the waiting queue is full, with `Retry-After: 5`. Cancelling pending work removes it from the queue immediately. Cancelling running work discards its eventual result; CPU execution finishes before the next request starts. A cold download can still take several minutes, and a long queue does not increase CPU throughput.
38
 
39
- Only exact matches to the public examples in `catalog.py`, `routing_catalog.py`, `clustering_catalog.py`, and `similarity_catalog.py` can reuse successful inference results. The key includes task, text, candidates, the complete ordered category configuration, or the complete ordered clustering texts, model ID, and pinned revision. Each cache hit gets a separate random job ID and `cache_hit: true`; original inference timings are preserved. These public results stay in an in-process cache capped at 80 entries until eviction or restart. Edited examples and other user inputs are never added to this cache.
40
 
41
  No user text, predictions, or exception payloads are logged or written to disk by this application. Pending input remains in memory until execution or cancellation; active input is released when execution returns. Private job results expire about five minutes after completion, failure, or cancellation, with cleanup at most 30 seconds later when idle. At most 64 job records are retained, so older completed results can expire sooner under load. Job IDs are unguessable bearer tokens; keep the ID private if the input is private. Hugging Face provides the surrounding hosting infrastructure. Use fictional data when trying PII examples.
42
 
43
  ## Interface
44
 
45
- All twelve demos share an input → model → result canvas, with the same compass model card, connecting paths, run-button placement, and example toolbar. Routing, Similarity, Reranker, and Clustering appear first in the horizontal navigation. Desktop uses three columns; narrower screens adapt the same flow without hiding task controls. The compass shows the selected model's identity and actual queue, loading, execution, or cache status.
46
 
47
- Results retain their task-specific meaning: classifications show label scores, PII offers complete highlighted and redacted text, Hazard preserves all twelve independently thresholded categories, and Similarity and Reranker show ordered passages with their respective cosine and raw relevance scores. Routing destinations remain editable in place. The global animation control pauses the ship, compass, and connecting paths.
48
 
49
  Similarity offers three example groups in the same tab: **Paraphrases**, **Question answering**, and **Duplicate questions**, with two verified examples per group. Paraphrases retains the original delivery-status and password-reset inputs. Question answering finds password-reset and order-return instructions. Duplicate questions compares Git-commit and Python-list questions against related but different questions. Changing the group or choosing an example prepares the inputs; click **Compare meanings** to run the model. Every example’s intended candidate ranked first in real pinned-model inference before inclusion. These checks establish the displayed rankings, not general model accuracy or a duplicate threshold. The earlier library example and its cross-language/time-detail limitations remain in the separate acceptance records.
50
 
@@ -54,6 +55,12 @@ PII offers seven examples covering 13 entity types: contact details, date and ag
54
 
55
  Each selected PII example was checked for complete entity spans and redaction; each Hazard example was checked against the full detected-label set using the unchanged published thresholds. Candidate IBAN and ZIP-code examples with incomplete spans or incorrect labels, and Hazard wordings with unrelated extra detections, were excluded; their raw results remain in the separate acceptance records. These curated cases demonstrate behavior on the displayed inputs, not general model accuracy.
56
 
 
 
 
 
 
 
57
  ## Routing
58
 
59
  Routing opens by default. Its shared canvas connects the prompt, the Vela Encoder compass card, and compact editable destinations. Scenario selection, examples, and reset share a toolbar below the canvas.
@@ -78,16 +85,17 @@ Clustering has the same queue, cancellation, and model-cache behavior as the oth
78
 
79
  ## Model contracts
80
 
81
- The September 14, 2026 semantic audit of the 20 built-in prompts found 19 aligned with their expected meaning and one miss: Guard's first, short instruction-override example returned `benign` (66.37%) instead of `jailbreak`. After that audit, the user requested a clearer positive Guard demo. Its new default query, a detailed developer-mode override, returned jailbreak 99.9923% on two identical real runs. Guard now offers two examples: the clear override and an ordinary explanatory request. At that point the ten model demos contained 20 examples, alongside 30 Routing examples. Subsequent curated PII and Hazard additions brought the total to 63 examples; three Clustering batches brought it to 66, and four additional Similarity examples bring the current total to 70. The original missed prompt and its raw result remain in the acceptance records; replacing a demonstration does not resolve the model’s short-override limitation. These example checks do not estimate general model accuracy.
82
 
83
  - Domain, FactCheck, Modality, Feedback, Guard, and Safety return softmax label scores.
84
  - FactCheck identifies a need for external factual knowledge or retrieval; it does not decide whether a claim is true. Guard concerns instruction attacks, while Safety and Hazard concern content risk. A risk signal alone does not prescribe refusal.
 
85
  - PII returns 17 entity types from 35 BIO labels. `start` and `end` are Unicode code-point offsets into the original text. Entity scores average their token probabilities.
86
  - Hazard uses independent sigmoid scores and the saved per-category thresholds from its [pinned operating point](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Hazard/blob/188325c44aaf929eeea4e24463ed2543549339b2/operating_point.json). The demo's 512-token maximum is inside one 2048-token window. Full-length clients must implement the documented overlapping-window policy.
87
  - Embedding uses its native ModernBERT checkpoint, attention-mask mean pooling in FP32, and L2 normalization, matching its SentenceTransformers configuration. Similarities are cosine scores between 768-dimensional vectors.
88
  - Reranker returns raw relevance logits. Its [pinned custom loader](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Reranker/blob/a388e41cbbd5dc5f16b6389fa76d0b8b8a38a8bf/modeling_vela_reranker.py) was inspected before enabling `trust_remote_code`; it loads the native encoder and separate trained FP32 classification heads. It uses all 22 layers and 768 dimensions.
89
 
90
- Model IDs, immutable revisions, examples, and Hazard thresholds live in `catalog.py`; Routing scenarios and their verified configurations live in `routing_catalog.py`; Clustering examples, per-item limits, and the shared default cutoff live in `clustering_catalog.py`. Similarity use-case groups and verified configurations live in `similarity_catalog.py`. Upgrading a model requires reviewing its contract and rerunning real inference checks.
91
 
92
  ## Local run
93
 
@@ -112,6 +120,8 @@ uvicorn app:app --host 0.0.0.0 --port 7860 --workers 1 --no-access-log
112
 
113
  `POST /api/jobs` accepts `{"task":"domain","text":"Why do bond prices fall?","candidates":[]}` and returns HTTP 202 with `{id, status, position, queue_ms}`. Poll `GET /api/jobs/{id}` for status: `queued`, `loading`, `running`, `completed`, `failed`, or `cancelled`. Queued positions start at one; active and terminal jobs have position zero. A completed snapshot adds `result`, containing the analysis response. A failed snapshot adds a safe `error` string and `error_status` (422 for input limits, otherwise 503). Missing or expired IDs return HTTP 404. `DELETE /api/jobs/{id}` cancels pending or running work; terminal jobs retain their final status. A cached example can already be completed in the initial HTTP 202 response.
114
 
 
 
115
  Routing jobs use `{"task":"routing","text":"Plan a weekend trip.","categories":[{"id":"travel","name":"Travel","description":"Trip planning and itineraries"},{"id":"coding","name":"Coding","description":"Programming and debugging"}]}`. Routing returns `kind: "routing"`, `top_category_id`, sorted `items` with original category IDs and indices, `dimensions: 768`, `score_type: "cosine"`, and the top-two `margin`. Similarity and Reranker retain their existing six-candidate limit.
116
 
117
  Clustering jobs use `{"task":"clustering","texts":["I forgot my password.","How can I reset my password?","How do I grow tomatoes?"]}`. They return `kind: "clustering"`, indexed original `items`, the cosine `similarities` matrix, a full average-link `merges` hierarchy, the `default_threshold`, and `groups` containing member `indices` and the representative index. Tree leaves are the original indices; merge `i` creates node `n + i`. Groups are ordered by their first input index, with deterministic tie breaking. Nonempty query text, candidates, or categories are rejected for this task; `texts` are rejected for other tasks.
@@ -132,4 +142,4 @@ python smoke_test.py --task domain
132
  python smoke_test.py --task all
133
  ```
134
 
135
- The all-model smoke test invokes ten pinned checkpoints when their example results are not cached and may download about 12.3 GB. It runs requests sequentially. Unit tests and output-shape checks alone do not establish prediction quality.
 
8
  header: mini
9
  pinned: false
10
  license: apache-2.0
11
+ short_description: Explore eleven Vela models for intelligent routing.
12
  models:
13
  - llm-semantic-router/Vela-1.0-Encoder-307M-Domain
14
  - llm-semantic-router/Vela-1.0-Encoder-307M-PII
 
16
  - llm-semantic-router/Vela-1.0-Encoder-307M-Safety
17
  - llm-semantic-router/Vela-1.0-Encoder-307M-Hazard
18
  - llm-semantic-router/Vela-1.0-Encoder-307M-FactCheck
19
+ - llm-semantic-router/Vela-1.0-Encoder-307M-Halu
20
  - llm-semantic-router/Vela-1.0-Encoder-307M-Modality
21
  - llm-semantic-router/Vela-1.0-Encoder-307M-Feedback
22
  - llm-semantic-router/Vela-1.0-Encoder-307M-Embedding
 
25
 
26
  # Vela Studio
27
 
28
+ A public playground for the [Vela 1.0 model family](https://huggingface.co/collections/llm-semantic-router/vela-10), by [vLLM Semantic Router](https://vllm-sr.ai). Results come from the selected model's pinned checkpoint on CPU. Exact built-in examples can reuse a prior successful inference, clearly marked as a cached example result. The interface includes thirteen interactive demos powered by eleven model checkpoints: classification, evidence-based hallucination highlights, PII highlights and redaction, content-risk thresholds, semantic similarity, dynamic prompt routing, passage reranking, and semantic clustering.
29
 
30
  ## Runtime
31
 
 
33
 
34
  Two 307M FP32 checkpoints require approximately 2.46 GB of parameter memory, plus runtime and activation memory. Execution remains sequential, with a single input or query-passage pair at a time. Set `VELA_MODEL_CACHE_SIZE=1` to retain only one model on smaller hosts; supported values are `1` and `2`, with `2` as the default.
35
 
36
+ Most demos accept up to **512 tokens including special tokens**; Clustering accepts **3–24 texts**, each up to **128 tokens including special tokens** and 2,048 characters. Similarity and Reranker allow at most **six candidates**; Routing allows **2–10 categories**. Hallucination applies the limit to the complete question, evidence, and answer pair; Reranker applies the limit to each query-and-passage pair; Embedding applies it independently to each text. Inputs exceeding this limit are rejected without truncation. The model family has a larger declared input capacity; this shorter public limit keeps CPU Basic usage bounded.
37
 
38
  One inference worker serves a FIFO queue with up to **eight waiting requests** in addition to the running request. The interface shows queue position, model loading, and inference progress and allows cancellation. HTTP 429 is returned only when the waiting queue is full, with `Retry-After: 5`. Cancelling pending work removes it from the queue immediately. Cancelling running work discards its eventual result; CPU execution finishes before the next request starts. A cold download can still take several minutes, and a long queue does not increase CPU throughput.
39
 
40
+ Only exact matches to the public examples in `catalog.py`, `routing_catalog.py`, `clustering_catalog.py`, `similarity_catalog.py`, and `hallucination_catalog.py` can reuse successful inference results. The key includes task, text, evidence and answer when supplied, candidates, the complete ordered category configuration, or the complete ordered clustering texts, model ID, and pinned revision. Each cache hit gets a separate random job ID and `cache_hit: true`; original inference timings are preserved. These public results stay in an in-process cache capped at 80 entries until eviction or restart. Edited examples and other user inputs are never added to this cache.
41
 
42
  No user text, predictions, or exception payloads are logged or written to disk by this application. Pending input remains in memory until execution or cancellation; active input is released when execution returns. Private job results expire about five minutes after completion, failure, or cancellation, with cleanup at most 30 seconds later when idle. At most 64 job records are retained, so older completed results can expire sooner under load. Job IDs are unguessable bearer tokens; keep the ID private if the input is private. Hugging Face provides the surrounding hosting infrastructure. Use fictional data when trying PII examples.
43
 
44
  ## Interface
45
 
46
+ All thirteen demos share an input → model → result canvas, with the same compass model card, connecting paths, run-button placement, and example toolbar. Routing, Similarity, Reranker, and Clustering appear first in the horizontal navigation. Desktop uses three columns; narrower screens adapt the same flow without hiding task controls. The compass shows the selected model's identity and actual queue, loading, execution, or cache status.
47
 
48
+ Results retain their task-specific meaning: classifications show label scores, Hallucination highlights unsupported answer spans, PII offers complete highlighted and redacted text, Hazard preserves all twelve independently thresholded categories, and Similarity and Reranker show ordered passages with their respective cosine and raw relevance scores. Routing destinations remain editable in place. The global animation control pauses the ship, compass, and connecting paths.
49
 
50
  Similarity offers three example groups in the same tab: **Paraphrases**, **Question answering**, and **Duplicate questions**, with two verified examples per group. Paraphrases retains the original delivery-status and password-reset inputs. Question answering finds password-reset and order-return instructions. Duplicate questions compares Git-commit and Python-list questions against related but different questions. Changing the group or choosing an example prepares the inputs; click **Compare meanings** to run the model. Every example’s intended candidate ranked first in real pinned-model inference before inclusion. These checks establish the displayed rankings, not general model accuracy or a duplicate threshold. The earlier library example and its cross-language/time-detail limitations remain in the separate acceptance records.
51
 
 
55
 
56
  Each selected PII example was checked for complete entity spans and redaction; each Hazard example was checked against the full detected-label set using the unchanged published thresholds. Candidate IBAN and ZIP-code examples with incomplete spans or incorrect labels, and Hazard wordings with unrelated extra detections, were excluded; their raw results remain in the separate acceptance records. These curated cases demonstrate behavior on the displayed inputs, not general model accuracy.
57
 
58
+ ## Hallucination
59
+
60
+ Choose **Hallucination**, enter a question, the evidence to use, and an answer, then click **Check the answer**. Vela Halu highlights answer spans that are unsupported by the supplied evidence. A result with no highlighted spans means no unsupported spans were detected; it is not a guarantee of factual correctness. FactCheck is a separate task that identifies requests needing factual knowledge.
61
+
62
+ Five examples cover a wrong opening time, a supported answer, a changed quantity, an invented facility, and a tool-result mismatch. All five were checked with the published model in real CPU inference. Editing any of the three fields clears the previous result and cancels pending work. The combined input must fit 512 tokens including the prompt format and special tokens; each field accepts up to 16,000 characters. Overlong inputs are rejected without truncation.
63
+
64
  ## Routing
65
 
66
  Routing opens by default. Its shared canvas connects the prompt, the Vela Encoder compass card, and compact editable destinations. Scenario selection, examples, and reset share a toolbar below the canvas.
 
85
 
86
  ## Model contracts
87
 
88
+ The September 14, 2026 semantic audit of the 20 built-in prompts found 19 aligned with their expected meaning and one miss: Guard's first, short instruction-override example returned `benign` (66.37%) instead of `jailbreak`. After that audit, the user requested a clearer positive Guard demo. Its new default query, a detailed developer-mode override, returned jailbreak 99.9923% on two identical real runs. Guard now offers two examples: the clear override and an ordinary explanatory request. At that point the ten model demos contained 20 examples, alongside 30 Routing examples. Subsequent curated PII and Hazard additions brought the total to 63 examples; three Clustering batches brought it to 66, and four additional Similarity examples brought the total to 70. Five Hallucination examples bring the current total to 75. The original missed prompt and its raw result remain in the acceptance records; replacing a demonstration does not resolve the model’s short-override limitation. These example checks do not estimate general model accuracy.
89
 
90
  - Domain, FactCheck, Modality, Feedback, Guard, and Safety return softmax label scores.
91
  - FactCheck identifies a need for external factual knowledge or retrieval; it does not decide whether a claim is true. Guard concerns instruction attacks, while Safety and Hazard concern content risk. A risk signal alone does not prescribe refusal.
92
+ - Hallucination pairs the question and supplied evidence with an answer and highlights answer tokens whose hallucination score is strictly above 0.5. It preserves the published model’s native YaRN configuration. Span offsets are Unicode code points into the original answer; each span score is the maximum of its detected token scores.
93
  - PII returns 17 entity types from 35 BIO labels. `start` and `end` are Unicode code-point offsets into the original text. Entity scores average their token probabilities.
94
  - Hazard uses independent sigmoid scores and the saved per-category thresholds from its [pinned operating point](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Hazard/blob/188325c44aaf929eeea4e24463ed2543549339b2/operating_point.json). The demo's 512-token maximum is inside one 2048-token window. Full-length clients must implement the documented overlapping-window policy.
95
  - Embedding uses its native ModernBERT checkpoint, attention-mask mean pooling in FP32, and L2 normalization, matching its SentenceTransformers configuration. Similarities are cosine scores between 768-dimensional vectors.
96
  - Reranker returns raw relevance logits. Its [pinned custom loader](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Reranker/blob/a388e41cbbd5dc5f16b6389fa76d0b8b8a38a8bf/modeling_vela_reranker.py) was inspected before enabling `trust_remote_code`; it loads the native encoder and separate trained FP32 classification heads. It uses all 22 layers and 768 dimensions.
97
 
98
+ Model IDs, immutable revisions, examples, and Hazard thresholds live in `catalog.py`; Routing scenarios and their verified configurations live in `routing_catalog.py`; Clustering examples, per-item limits, and the shared default cutoff live in `clustering_catalog.py`. Similarity use-case groups and verified configurations live in `similarity_catalog.py`. Hallucination examples and the fixed threshold live in `hallucination_catalog.py`. Upgrading a model requires reviewing its contract and rerunning real inference checks.
99
 
100
  ## Local run
101
 
 
120
 
121
  `POST /api/jobs` accepts `{"task":"domain","text":"Why do bond prices fall?","candidates":[]}` and returns HTTP 202 with `{id, status, position, queue_ms}`. Poll `GET /api/jobs/{id}` for status: `queued`, `loading`, `running`, `completed`, `failed`, or `cancelled`. Queued positions start at one; active and terminal jobs have position zero. A completed snapshot adds `result`, containing the analysis response. A failed snapshot adds a safe `error` string and `error_status` (422 for input limits, otherwise 503). Missing or expired IDs return HTTP 404. `DELETE /api/jobs/{id}` cancels pending or running work; terminal jobs retain their final status. A cached example can already be completed in the initial HTTP 202 response.
122
 
123
+ Hallucination jobs use `{"task":"hallucination","text":"When does the museum open?","context":"The museum opens at 10:00.","answer":"The museum opens at 09:00."}`. All three text fields are required and preserved verbatim. The result has `kind: "hallucination"`, the original `answer`, `spans` with `start`, `end`, `text`, and `score`, plus `max_hallucination_score` across answer tokens. Nonempty `context` or `answer` fields are rejected for other tasks. Only exact matches of all three fields can reuse a public example result.
124
+
125
  Routing jobs use `{"task":"routing","text":"Plan a weekend trip.","categories":[{"id":"travel","name":"Travel","description":"Trip planning and itineraries"},{"id":"coding","name":"Coding","description":"Programming and debugging"}]}`. Routing returns `kind: "routing"`, `top_category_id`, sorted `items` with original category IDs and indices, `dimensions: 768`, `score_type: "cosine"`, and the top-two `margin`. Similarity and Reranker retain their existing six-candidate limit.
126
 
127
  Clustering jobs use `{"task":"clustering","texts":["I forgot my password.","How can I reset my password?","How do I grow tomatoes?"]}`. They return `kind: "clustering"`, indexed original `items`, the cosine `similarities` matrix, a full average-link `merges` hierarchy, the `default_threshold`, and `groups` containing member `indices` and the representative index. Tree leaves are the original indices; merge `i` creates node `n + i`. Groups are ordered by their first input index, with deterministic tie breaking. Nonempty query text, candidates, or categories are rejected for this task; `texts` are rejected for other tasks.
 
142
  python smoke_test.py --task all
143
  ```
144
 
145
+ The all-model smoke test invokes eleven pinned checkpoints when their example results are not cached and may download about 13.5 GB. It runs requests sequentially. Unit tests and output-shape checks alone do not establish prediction quality.
app.py CHANGED
@@ -14,6 +14,7 @@ from starlette.concurrency import run_in_threadpool
14
 
15
  from catalog import MAX_CANDIDATES, MAX_CLUSTER_ITEMS, MAX_ROUTING_CATEGORIES, TASKS, public_catalog
16
  from inference import InferenceEngine, InputLimitError
 
17
  from job_queue import InferenceQueue, JobFailure, QueueFull
18
  from request_limits import RequestSizeLimitMiddleware
19
 
@@ -26,7 +27,8 @@ def run_inference(payload, set_stage):
26
  with engine.slot:
27
  return engine.analyze_locked(payload["task"], payload["text"], payload["candidates"],
28
  on_stage=set_stage, categories=payload.get("categories", []),
29
- texts=payload.get("texts", []))
 
30
  except InputLimitError as exc:
31
  raise JobFailure(str(exc), 422) from None
32
  except Exception as exc:
@@ -79,6 +81,8 @@ class AnalyzeRequest(BaseModel):
79
  candidates: list[Candidate] = Field(default_factory=list, max_length=MAX_CANDIDATES)
80
  categories: list[RoutingCategory] = Field(default_factory=list, max_length=MAX_ROUTING_CATEGORIES)
81
  texts: list[ClusterText] = Field(default_factory=list, max_length=MAX_CLUSTER_ITEMS)
 
 
82
 
83
  @field_validator("task")
84
  @classmethod
@@ -89,6 +93,14 @@ class AnalyzeRequest(BaseModel):
89
 
90
  @model_validator(mode="after")
91
  def valid_content(self):
 
 
 
 
 
 
 
 
92
  if self.task == "clustering":
93
  if self.text or self.candidates or self.categories:
94
  raise ValueError("Clustering accepts only the texts list.")
 
14
 
15
  from catalog import MAX_CANDIDATES, MAX_CLUSTER_ITEMS, MAX_ROUTING_CATEGORIES, TASKS, public_catalog
16
  from inference import InferenceEngine, InputLimitError
17
+ from hallucination_catalog import MAX_HALLUCINATION_FIELD_CHARS
18
  from job_queue import InferenceQueue, JobFailure, QueueFull
19
  from request_limits import RequestSizeLimitMiddleware
20
 
 
27
  with engine.slot:
28
  return engine.analyze_locked(payload["task"], payload["text"], payload["candidates"],
29
  on_stage=set_stage, categories=payload.get("categories", []),
30
+ texts=payload.get("texts", []),
31
+ context=payload.get("context", ""), answer=payload.get("answer", ""))
32
  except InputLimitError as exc:
33
  raise JobFailure(str(exc), 422) from None
34
  except Exception as exc:
 
81
  candidates: list[Candidate] = Field(default_factory=list, max_length=MAX_CANDIDATES)
82
  categories: list[RoutingCategory] = Field(default_factory=list, max_length=MAX_ROUTING_CATEGORIES)
83
  texts: list[ClusterText] = Field(default_factory=list, max_length=MAX_CLUSTER_ITEMS)
84
+ context: str = Field(default="", max_length=MAX_HALLUCINATION_FIELD_CHARS)
85
+ answer: str = Field(default="", max_length=MAX_HALLUCINATION_FIELD_CHARS)
86
 
87
  @field_validator("task")
88
  @classmethod
 
93
 
94
  @model_validator(mode="after")
95
  def valid_content(self):
96
+ if self.task == "hallucination":
97
+ if self.candidates or self.categories or self.texts:
98
+ raise ValueError("Hallucination accepts only a question, evidence, and answer.")
99
+ if not all(value.strip() for value in (self.text, self.context, self.answer)):
100
+ raise ValueError("Question, evidence, and answer cannot be blank.")
101
+ return self
102
+ if self.context or self.answer:
103
+ raise ValueError("Evidence and answer are only supported by Hallucination.")
104
  if self.task == "clustering":
105
  if self.text or self.candidates or self.categories:
106
  raise ValueError("Clustering accepts only the texts list.")
catalog.py CHANGED
@@ -8,6 +8,7 @@ from routing_catalog import MAX_ROUTING_CATEGORIES, ROUTING_SCENARIOS
8
  from similarity_catalog import SIMILARITY_EXAMPLES, SIMILARITY_EXAMPLE_GROUPS
9
  from clustering_catalog import (CLUSTERING_EXAMPLES, CLUSTER_DISTANCE_THRESHOLD,
10
  MAX_CLUSTER_ITEMS, MAX_CLUSTER_TOKENS)
 
11
 
12
  _TASKS = [
13
  ("domain", "Domain", "Find the subject of a request across 14 domains.", "classification", "f6354f54adcf38770f635ad903be2b00577f6c11", [
@@ -73,6 +74,14 @@ TASKS = {
73
 
74
  TASKS["embedding"]["example_groups"] = SIMILARITY_EXAMPLE_GROUPS
75
 
 
 
 
 
 
 
 
 
76
  # Routing is another demo of the same Embedding checkpoint, not a new model.
77
  TASKS["routing"] = {
78
  "id": "routing", "name": "Routing", "description": "Route prompts to categories you define.",
@@ -117,4 +126,5 @@ def public_catalog():
117
  "max_tokens": MAX_TOKENS, "max_candidates": MAX_CANDIDATES,
118
  "max_routing_categories": MAX_ROUTING_CATEGORIES,
119
  "max_cluster_items": MAX_CLUSTER_ITEMS, "max_cluster_tokens": MAX_CLUSTER_TOKENS,
 
120
  }}
 
8
  from similarity_catalog import SIMILARITY_EXAMPLES, SIMILARITY_EXAMPLE_GROUPS
9
  from clustering_catalog import (CLUSTERING_EXAMPLES, CLUSTER_DISTANCE_THRESHOLD,
10
  MAX_CLUSTER_ITEMS, MAX_CLUSTER_TOKENS)
11
+ from hallucination_catalog import HALLUCINATION_EXAMPLES, MAX_HALLUCINATION_FIELD_CHARS
12
 
13
  _TASKS = [
14
  ("domain", "Domain", "Find the subject of a request across 14 domains.", "classification", "f6354f54adcf38770f635ad903be2b00577f6c11", [
 
74
 
75
  TASKS["embedding"]["example_groups"] = SIMILARITY_EXAMPLE_GROUPS
76
 
77
+ TASKS["hallucination"] = {
78
+ "id": "hallucination", "name": "Hallucination",
79
+ "description": "Find answer spans unsupported by the supplied evidence.",
80
+ "kind": "hallucination", "model": MODEL_PREFIX + "Halu",
81
+ "revision": "521cd05d15e1959e120d002663c3b775649ae4dd",
82
+ "examples": HALLUCINATION_EXAMPLES,
83
+ }
84
+
85
  # Routing is another demo of the same Embedding checkpoint, not a new model.
86
  TASKS["routing"] = {
87
  "id": "routing", "name": "Routing", "description": "Route prompts to categories you define.",
 
126
  "max_tokens": MAX_TOKENS, "max_candidates": MAX_CANDIDATES,
127
  "max_routing_categories": MAX_ROUTING_CATEGORIES,
128
  "max_cluster_items": MAX_CLUSTER_ITEMS, "max_cluster_tokens": MAX_CLUSTER_TOKENS,
129
+ "max_hallucination_field_chars": MAX_HALLUCINATION_FIELD_CHARS,
130
  }}
hallucination.py ADDED
@@ -0,0 +1,87 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Native Vela Halu input and answer-span contracts."""
2
+
3
+ import math
4
+
5
+ from hallucination_catalog import HALLUCINATION_THRESHOLD
6
+
7
+
8
+ def evidence_prompt(question, context):
9
+ return f"User request: {question}\n\n{context}"
10
+
11
+
12
+ def validate_hallucination_model(model):
13
+ config = model.config
14
+ if (type(model).__name__ != "ModernBertForTokenClassification"
15
+ or config.model_type != "modernbert"
16
+ or config.id2label != {0: "supported", 1: "hallucinated"}
17
+ or config.label2id != {"supported": 0, "hallucinated": 1}
18
+ or config.num_labels != 2
19
+ or config.reference_compile is not False
20
+ or config._attn_implementation != "sdpa"):
21
+ raise ValueError("Halu checkpoint differs from its native inference contract")
22
+ if (getattr(config, "rope_scaling", None) or {}).get("rope_type") != "yarn":
23
+ raise ValueError("Halu requires the published YaRN configuration")
24
+ count = 0
25
+ for _, module in model.named_modules():
26
+ if not hasattr(module, "rotary_emb"):
27
+ continue
28
+ rotary = module.rotary_emb
29
+ if (getattr(rotary, "rope_type", None) != "yarn"
30
+ or getattr(getattr(rotary, "rope_init_fn", None), "__name__", None)
31
+ != "_compute_yarn_parameters"):
32
+ raise ValueError("Halu did not instantiate its published YaRN rotary")
33
+ count += 1
34
+ if count != config.num_hidden_layers or count != 22:
35
+ raise ValueError("Halu rotary layer count differs from its foundation")
36
+
37
+
38
+ def hallucination_result(answer, sequence_ids, offsets, scores):
39
+ """Use answer-only Unicode offsets and the published strict > 0.5 rule."""
40
+ if not len(sequence_ids) == len(offsets) == len(scores):
41
+ raise ValueError("Halu token scores and offsets differ in length")
42
+ spans, current = [], None
43
+ max_score = 0.0
44
+ previous_start = -1
45
+ answer_tokens = 0
46
+
47
+ def finish():
48
+ nonlocal current
49
+ if current is not None:
50
+ # Byte-fallback tokens can share a character while straddling the
51
+ # threshold. Present their positive character union once; ordinary
52
+ # nonoverlapping spans separated by a supported token stay distinct.
53
+ if spans and current["start"] < spans[-1]["end"]:
54
+ span = spans[-1]
55
+ span["end"] = max(span["end"], current["end"])
56
+ span["score"] = max(span["score"], current["score"])
57
+ else:
58
+ span = current
59
+ spans.append(span)
60
+ span["text"] = answer[span["start"]:span["end"]]
61
+ current = None
62
+
63
+ for sequence, (start, end), score in zip(sequence_ids, offsets, scores, strict=True):
64
+ score = float(score)
65
+ if not math.isfinite(score) or not 0 <= score <= 1:
66
+ raise ValueError("Invalid Halu token probability")
67
+ if sequence != 1 or end <= start:
68
+ continue
69
+ if (type(start) is not int or type(end) is not int
70
+ or not 0 <= start < end <= len(answer) or start < previous_start):
71
+ raise ValueError("Invalid Halu answer offsets")
72
+ previous_start = start
73
+ answer_tokens += 1
74
+ max_score = max(max_score, score)
75
+ if score > HALLUCINATION_THRESHOLD:
76
+ if current is None:
77
+ current = {"start": start, "end": end, "score": score}
78
+ else:
79
+ current["end"] = max(current["end"], end)
80
+ current["score"] = max(current["score"], score)
81
+ else:
82
+ finish()
83
+ finish()
84
+ if not answer_tokens:
85
+ raise ValueError("Halu found no evaluable answer tokens")
86
+ return {"kind": "hallucination", "answer": answer, "spans": spans,
87
+ "max_hallucination_score": max_score}
hallucination_catalog.py ADDED
@@ -0,0 +1,37 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Public evidence-and-answer examples for Vela Halu."""
2
+
3
+ MAX_HALLUCINATION_FIELD_CHARS = 16000
4
+ HALLUCINATION_THRESHOLD = 0.5
5
+
6
+ HALLUCINATION_EXAMPLES = [
7
+ {
8
+ "name": "Wrong opening time",
9
+ "text": "When does the museum open on Tuesday?",
10
+ "context": "The museum opens at 10:00 on Tuesday.",
11
+ "answer": "The museum opens at 09:00 on Tuesday.",
12
+ },
13
+ {
14
+ "name": "Supported answer",
15
+ "text": "When does the museum open on Tuesday?",
16
+ "context": "The museum opens at 10:00 on Tuesday.",
17
+ "answer": "The museum opens at 10:00 on Tuesday.",
18
+ },
19
+ {
20
+ "name": "Changed number",
21
+ "text": "How many seats does the shuttle have?",
22
+ "context": "The shuttle has 24 seats and runs every 30 minutes.",
23
+ "answer": "The shuttle has 42 seats and runs every 30 minutes.",
24
+ },
25
+ {
26
+ "name": "Invented detail",
27
+ "text": "What facilities does the library offer?",
28
+ "context": "The library offers free Wi-Fi and a reading room.",
29
+ "answer": "The library offers free Wi-Fi, a reading room, and a rooftop café.",
30
+ },
31
+ {
32
+ "name": "Tool output mismatch",
33
+ "text": "Summarize the update tool's result.",
34
+ "context": 'The update tool returned: {"status": "success", "files_updated": 3, "tests_passed": 12}.',
35
+ "answer": "The tool updated 7 files and all 12 tests passed.",
36
+ },
37
+ ]
inference.py CHANGED
@@ -10,6 +10,7 @@ from collections import OrderedDict
10
  from catalog import (HAZARD_THRESHOLDS, MAX_CANDIDATES, MAX_CLUSTER_ITEMS, MAX_CLUSTER_TOKENS,
11
  MAX_ROUTING_CATEGORIES, MAX_TOKENS, TASKS)
12
  from clustering import clustering_result
 
13
 
14
 
15
  class InputLimitError(ValueError):
@@ -154,10 +155,12 @@ class InferenceEngine:
154
  return AutoTokenizer.from_pretrained(spec["model"], revision=spec["revision"],
155
  trust_remote_code=False, use_fast=True)
156
 
157
- def _check_tokens(self, tokenizer, task, text, candidates, texts=None):
158
  limit = MAX_CLUSTER_TOKENS if task == "clustering" else MAX_TOKENS
159
  if task == "clustering":
160
  inputs = [(item, None) for item in texts or []]
 
 
161
  elif task == "reranker":
162
  inputs = [(text, candidate) for candidate in candidates]
163
  elif task in {"embedding", "routing"}:
@@ -167,7 +170,8 @@ class InferenceEngine:
167
  for index, (first, second) in enumerate(inputs):
168
  count = len(tokenizer(first, text_pair=second, truncation=False)["input_ids"])
169
  if count > limit:
170
- subject = (f"Text {index + 1}" if task == "clustering" else
 
171
  f"Query + passage {index + 1}" if task == "reranker"
172
  else "Input" if index == 0 else
173
  f"Category {index}" if task == "routing" else f"Candidate {index}")
@@ -202,10 +206,12 @@ class InferenceEngine:
202
  code_revision=spec["revision"], **arguments)
203
  else:
204
  model_class = (AutoModel if task == "embedding" else
205
- AutoModelForTokenClassification if task == "pii" else
206
  AutoModelForSequenceClassification)
207
  model = model_class.from_pretrained(spec["model"], trust_remote_code=False,
208
  reference_compile=False, **arguments)
 
 
209
  model.eval()
210
  with self._cache_lock:
211
  self._cache[task] = (model, tokenizer)
@@ -252,7 +258,8 @@ class InferenceEngine:
252
  vectors.append(vector)
253
  return vectors, hits
254
 
255
- def analyze_locked(self, task, text, candidates, on_stage=None, categories=None, texts=None):
 
256
  """The caller owns slot for the entire load and inference lifecycle."""
257
  started = time.perf_counter()
258
  model_task = self._model_task(task)
@@ -262,7 +269,8 @@ class InferenceEngine:
262
  if on_stage:
263
  on_stage("loading" if cold else "running")
264
  tokenizer = self._get_tokenizer(task)
265
- self._check_tokens(tokenizer, task, text, representations, texts=texts)
 
266
  self._ensure_model(task, tokenizer)
267
  load_ms = (time.perf_counter() - started) * 1000 if cold else 0.0
268
  if on_stage:
@@ -306,6 +314,15 @@ class InferenceEngine:
306
  tokens = [(model.config.id2label[index], score, start, end)
307
  for index, score, (start, end) in zip(indices.tolist(), scores.tolist(), offsets)]
308
  result = pii_result(text, tokens)
 
 
 
 
 
 
 
 
 
309
  else:
310
  inputs = self._inputs(tokenizer, text)
311
  logits = model(**inputs).logits[0].float().tolist()
 
10
  from catalog import (HAZARD_THRESHOLDS, MAX_CANDIDATES, MAX_CLUSTER_ITEMS, MAX_CLUSTER_TOKENS,
11
  MAX_ROUTING_CATEGORIES, MAX_TOKENS, TASKS)
12
  from clustering import clustering_result
13
+ from hallucination import evidence_prompt, hallucination_result, validate_hallucination_model
14
 
15
 
16
  class InputLimitError(ValueError):
 
155
  return AutoTokenizer.from_pretrained(spec["model"], revision=spec["revision"],
156
  trust_remote_code=False, use_fast=True)
157
 
158
+ def _check_tokens(self, tokenizer, task, text, candidates, texts=None, context="", answer=""):
159
  limit = MAX_CLUSTER_TOKENS if task == "clustering" else MAX_TOKENS
160
  if task == "clustering":
161
  inputs = [(item, None) for item in texts or []]
162
+ elif task == "hallucination":
163
+ inputs = [(evidence_prompt(text, context), answer)]
164
  elif task == "reranker":
165
  inputs = [(text, candidate) for candidate in candidates]
166
  elif task in {"embedding", "routing"}:
 
170
  for index, (first, second) in enumerate(inputs):
171
  count = len(tokenizer(first, text_pair=second, truncation=False)["input_ids"])
172
  if count > limit:
173
+ subject = ("Question + evidence + answer" if task == "hallucination" else
174
+ f"Text {index + 1}" if task == "clustering" else
175
  f"Query + passage {index + 1}" if task == "reranker"
176
  else "Input" if index == 0 else
177
  f"Category {index}" if task == "routing" else f"Candidate {index}")
 
206
  code_revision=spec["revision"], **arguments)
207
  else:
208
  model_class = (AutoModel if task == "embedding" else
209
+ AutoModelForTokenClassification if task in {"pii", "hallucination"} else
210
  AutoModelForSequenceClassification)
211
  model = model_class.from_pretrained(spec["model"], trust_remote_code=False,
212
  reference_compile=False, **arguments)
213
+ if task == "hallucination":
214
+ validate_hallucination_model(model)
215
  model.eval()
216
  with self._cache_lock:
217
  self._cache[task] = (model, tokenizer)
 
258
  vectors.append(vector)
259
  return vectors, hits
260
 
261
+ def analyze_locked(self, task, text, candidates, on_stage=None, categories=None, texts=None,
262
+ context="", answer=""):
263
  """The caller owns slot for the entire load and inference lifecycle."""
264
  started = time.perf_counter()
265
  model_task = self._model_task(task)
 
269
  if on_stage:
270
  on_stage("loading" if cold else "running")
271
  tokenizer = self._get_tokenizer(task)
272
+ self._check_tokens(tokenizer, task, text, representations, texts=texts,
273
+ context=context, answer=answer)
274
  self._ensure_model(task, tokenizer)
275
  load_ms = (time.perf_counter() - started) * 1000 if cold else 0.0
276
  if on_stage:
 
314
  tokens = [(model.config.id2label[index], score, start, end)
315
  for index, score, (start, end) in zip(indices.tolist(), scores.tolist(), offsets)]
316
  result = pii_result(text, tokens)
317
+ elif task == "hallucination":
318
+ inputs = self._inputs(tokenizer, evidence_prompt(text, context), answer, offsets=True)
319
+ sequence_ids = inputs.sequence_ids(0)
320
+ offsets = inputs.pop("offset_mapping")[0].tolist()
321
+ logits = model(**inputs).logits[0]
322
+ if logits.shape[-1] != 2:
323
+ raise ValueError("Halu must return exactly two token labels")
324
+ scores = logits.float().softmax(dim=-1)[:, 1].tolist()
325
+ result = hallucination_result(answer, sequence_ids, offsets, scores)
326
  else:
327
  inputs = self._inputs(tokenizer, text)
328
  logits = model(**inputs).logits[0].float().tolist()
job_queue.py CHANGED
@@ -91,12 +91,16 @@ class InferenceQueue:
91
  candidates = tuple(payload.get("candidates", []))
92
  categories = self._category_key(payload.get("categories", []))
93
  texts = tuple(payload.get("texts", []))
 
94
  for example in spec.get("examples", []):
95
  if (payload.get("text", "") == example.get("text", "")
96
  and candidates == tuple(example.get("candidates", []))
97
  and categories == self._category_key(example.get("categories", []))
98
- and texts == tuple(example.get("texts", []))):
99
- return (payload["task"], spec["model"], spec["revision"], payload.get("text", ""), candidates, categories, texts)
 
 
 
100
  return None
101
 
102
  @staticmethod
 
91
  candidates = tuple(payload.get("candidates", []))
92
  categories = self._category_key(payload.get("categories", []))
93
  texts = tuple(payload.get("texts", []))
94
+ context, answer = payload.get("context", ""), payload.get("answer", "")
95
  for example in spec.get("examples", []):
96
  if (payload.get("text", "") == example.get("text", "")
97
  and candidates == tuple(example.get("candidates", []))
98
  and categories == self._category_key(example.get("categories", []))
99
+ and texts == tuple(example.get("texts", []))
100
+ and context == example.get("context", "")
101
+ and answer == example.get("answer", "")):
102
+ return (payload["task"], spec["model"], spec["revision"], payload.get("text", ""),
103
+ candidates, categories, texts, context, answer)
104
  return None
105
 
106
  @staticmethod
smoke_test.py CHANGED
@@ -2,7 +2,7 @@
2
 
3
  Usage: python smoke_test.py --base-url http://localhost:7860 --task all
4
  Uncached examples execute the pinned model; cached responses are reported as
5
- such. This can download ~12.3 GB of weights on its first run. This checks output
6
  contracts, not semantic correctness; see the separately recorded example audit.
7
  """
8
 
@@ -27,7 +27,7 @@ def request(base, path, payload=None):
27
  def example_payload(task, example):
28
  # Catalog examples can also include display names and other metadata.
29
  return {"task": task["id"], **{key: value for key, value in example.items()
30
- if key in {"text", "candidates", "categories", "texts"}}}
31
 
32
 
33
  def verify_clustering(result, task, example):
@@ -88,6 +88,15 @@ def verify(response, task, example):
88
  assert result["entities"], "The public example should contain detected PII"
89
  for entity in result["entities"]:
90
  assert example["text"][entity["start"]:entity["end"]] == entity["text"]
 
 
 
 
 
 
 
 
 
91
  elif result["kind"] == "routing":
92
  categories = example["categories"]
93
  assert len(result["items"]) == len(categories)
@@ -121,7 +130,7 @@ def main():
121
  parser.error("Unknown task")
122
  example_count = 0
123
  for task in selected:
124
- examples = task["examples"] if task["kind"] == "clustering" else task["examples"][:1]
125
  for example in examples:
126
  payload = example_payload(task, example)
127
  response = request(args.base_url, "/api/analyze", payload)
 
2
 
3
  Usage: python smoke_test.py --base-url http://localhost:7860 --task all
4
  Uncached examples execute the pinned model; cached responses are reported as
5
+ such. This can download ~13.5 GB of weights on its first run. This checks output
6
  contracts, not semantic correctness; see the separately recorded example audit.
7
  """
8
 
 
27
  def example_payload(task, example):
28
  # Catalog examples can also include display names and other metadata.
29
  return {"task": task["id"], **{key: value for key, value in example.items()
30
+ if key in {"text", "context", "answer", "candidates", "categories", "texts"}}}
31
 
32
 
33
  def verify_clustering(result, task, example):
 
88
  assert result["entities"], "The public example should contain detected PII"
89
  for entity in result["entities"]:
90
  assert example["text"][entity["start"]:entity["end"]] == entity["text"]
91
+ elif result["kind"] == "hallucination":
92
+ assert result["answer"] == example["answer"]
93
+ assert 0 <= result["max_hallucination_score"] <= 1
94
+ previous_end = 0
95
+ for span in result["spans"]:
96
+ assert previous_end <= span["start"] < span["end"] <= len(result["answer"])
97
+ assert result["answer"][span["start"]:span["end"]] == span["text"]
98
+ assert 0.5 < span["score"] <= 1
99
+ previous_end = span["end"]
100
  elif result["kind"] == "routing":
101
  categories = example["categories"]
102
  assert len(result["items"]) == len(categories)
 
130
  parser.error("Unknown task")
131
  example_count = 0
132
  for task in selected:
133
+ examples = task["examples"] if task["kind"] in {"clustering", "hallucination"} else task["examples"][:1]
134
  for example in examples:
135
  payload = example_payload(task, example)
136
  response = request(args.base_url, "/api/analyze", payload)
static/app.js CHANGED
@@ -8,6 +8,7 @@ const presentation = {
8
  clustering: {name:'Clustering',group:'Retrieve',icon:'⁙',title:'Shared meaning. Natural company.',description:'Discover groups in a collection of texts, without predefined categories.',action:'Find the groups',note:'Adjust granularity to merge or separate groups. Open a group to inspect every text.',examples:[]},
9
  domain: {name:'Domain',group:'Understand',icon:'◈',title:'Every question has a world.',description:'Find the subject behind a request, across 14 knowledge domains.',action:'Find the domain',note:'Scores describe the model’s distribution across its domain labels.',examples:['A curious question','Another field']},
10
  factcheck: {name:'Fact check',group:'Understand',icon:'⊙',title:'Know when to look it up.',description:'Recognize requests that call for external factual knowledge.',action:'Check the need',note:'This predicts whether a fact check is needed. It does not verify whether a claim is true.',examples:['Factual knowledge','Creative writing']},
 
11
  modality: {name:'Modality',group:'Understand',icon:'▧',title:'The right medium for the idea.',description:'Explore whether a request calls for text, an image, or both.',action:'Find the medium',note:'AR means text generation; DIFFUSION means image generation; BOTH means both.',examples:['An image idea','Words and pictures']},
12
  feedback: {name:'Feedback',group:'Understand',icon:'☷',title:'Listen a little closer.',description:'Understand the feedback hidden in a conversational response.',action:'Read the feedback',note:'Predictions distinguish satisfaction, clarification, correction, a different answer, and no feedback.',examples:['A little gratitude','A correction']},
13
  pii: {name:'Personal data',group:'Protect',icon:'◎',title:'Find what should stay private.',description:'Locate personal information, then see it highlighted or redacted.',action:'Detect PII',note:'Highlights show model predictions. Review them before relying on redaction.',examples:["Contact details","Date & age","Street address","Network log","Test card","Synthetic ID","Work profile"]},
@@ -35,16 +36,18 @@ function renderEmpty() {
35
  FlowUI.setTargets([]);
36
  if(selected==='routing' && catalog.routing) { RoutingUI.renderEmpty(); return; }
37
  if(selected==='clustering') { ClusteringUI.renderEmpty(); return; }
 
38
  $('output').innerHTML = '<div class="empty-state"><span class="flow-empty-label">A SIGNAL, WAITING TO BE FOUND</span><strong>Ready when you are.</strong><p>Choose an example or bring your own text. Run the model to see what it notices.</p></div>';
39
  $('result-meta').textContent=''; $('output-tabs').hidden=true; $('copy-button').hidden=true;
40
  $('output').setAttribute('aria-live','polite');
41
  }
42
  function setModelState(message) { $('model-state').replaceChildren(textNode('span','','hollow-dot'),textNode('span',message)); }
43
  function selectTask(id) {
44
- if (!presentation[id] || (busy && !['routing','clustering'].includes(selected))) return;
45
  if(busy) invalidateRouting();
46
  RoutingUI.leave();
47
  ClusteringUI.leave();
 
48
  FlowUI.leave();
49
  selected=id; lastResult=null; edited=false; view='highlight'; $('error').hidden=true;
50
  $('result-meta').classList.remove('stale-note');
@@ -55,13 +58,14 @@ function selectTask(id) {
55
  $('task-title').textContent=p.title; $('task-description').textContent=p.description; $('task-symbol').textContent=p.icon;
56
  $('run-label').textContent=p.action; $('task-note').textContent=p.note;
57
  const retrieval=['embedding','reranker'].includes(id);
58
- $('candidates-wrap').hidden=!retrieval; $('input-label').textContent=id==='routing'?'YOUR PROMPT':id==='clustering'?'YOUR TEXTS':retrieval?'YOUR QUERY':'YOUR TEXT';
59
  $('candidates-label').textContent=id==='embedding'?'CANDIDATES':'PASSAGES';
60
- $('input-hint').textContent=id==='clustering'?'3–24 texts · 128 tokens each':'Up to 512 tokens';
61
- $('input').maxLength=id==='clustering'?98351:12000;
62
- $('input').placeholder=id==='clustering'?'One text per line. Let shared meaning find its shape.':'Write a little. Discover a lot.';
63
- document.querySelector('.examples>span').textContent=id==='clustering'?'TRY A COLLECTION':'TRY A PROMPT';
64
- $('input').value=''; $('candidates').value='';
 
65
  $('model-link').textContent=`Vela-1.0 · ${task ? task.model.split('-').pop() : p.name} ↗`;
66
  $('model-link').href=`https://huggingface.co/${task?.model || 'collections/llm-semantic-router/vela-10'}`;
67
  setModelState(cachedTasks.includes(task?.model_task || id)?'Ready on CPU':'Loads on first run');
@@ -74,17 +78,18 @@ function selectTask(id) {
74
  renderExamples();
75
  FlowUI.enter(id,p,task);
76
  if(id==='clustering')ClusteringUI.enter();
 
77
  if(id==='routing' && task?.scenarios) RoutingUI.enter(task);
78
  else if (task?.examples?.length) loadExample(task.examples.find(example=>!grouped || example.group===similarityGroup));
79
  countCharacters(); renderEmpty();
80
  const activeButton=$('tasks').querySelector(`[data-task="${id}"]`), nav=$('tasks');
81
  if(activeButton && nav.scrollWidth>nav.clientWidth) nav.scrollLeft+=activeButton.getBoundingClientRect().left-nav.getBoundingClientRect().left-(nav.clientWidth-activeButton.clientWidth)/2;
82
  }
83
- function countCharacters() { $('char-count').textContent=selected==='clustering'?`${ClusteringUI.texts().length} / 24 texts`:`${Array.from($('input').value).length.toLocaleString()} characters`; }
84
  function inputChanged() {
85
  countCharacters(); updateExampleNote(); $('error').hidden=true;
86
  if(selected==='routing'){RoutingUI.inputChanged();return;}
87
- if(selected==='clustering'){invalidateRouting();return;}
88
  if(selected==='embedding'){
89
  lastResult=null;edited=false;$('result-meta').classList.remove('stale-note');renderEmpty();
90
  if(!busy)setModelState(cachedTasks.includes('embedding')?'Ready on CPU':'Loads on first run');
@@ -108,8 +113,10 @@ function renderExamples() {
108
  }
109
  }
110
  function loadExample(example) {
111
- if(!example || (busy && !['routing','clustering'].includes(selected)))return;
112
- $('input').value=selected==='clustering'?(example.texts || []).join('\n'):example.text;$('candidates').value=(example.candidates || []).join('\n');inputChanged();
 
 
113
  }
114
  function selectSimilarityGroup() {
115
  if(selected!=='embedding')return;
@@ -120,10 +127,11 @@ function selectSimilarityGroup() {
120
  }
121
  function setBusy(value) {
122
  busy=value; document.body.classList.toggle('is-running',value);
123
- const editable=['routing','clustering'].includes(selected);
124
  for (const el of document.querySelectorAll('.task-button,.example-button,#clear-button,#similarity-group')) el.disabled=value && !editable;
125
  $('run-button').disabled=value;
126
  $('input').readOnly=value && !editable; $('candidates').readOnly=value;
 
127
  $('run-label').textContent=value?'Finding the signal…':presentation[selected].action;
128
  $('run-arrow').textContent=value?'◌':'↗'; $('output').setAttribute('aria-busy',String(value));
129
  }
@@ -135,6 +143,7 @@ function renderLoading() {
135
  $('output-tabs').hidden=true; $('copy-button').hidden=true; $('result-meta').textContent='';
136
  $('output').innerHTML='<div class="empty-state loading"><span class="flow-empty-label">FOLLOWING THE SIGNAL</span><strong id="loading-title">Finding you a place…</strong><p id="loading-copy">Your request will wait its turn on the shared CPU.</p><div class="loading-track" aria-hidden="true"></div><button id="cancel-button" class="text-button cancel-button" type="button">Cancel request</button></div>';
137
  if(selected==='clustering')$('loading-title').textContent='Gathering shared meaning…';
 
138
  $('cancel-button').addEventListener('click',cancelRun);
139
  setModelState('Submitting request…');
140
  }
@@ -176,7 +185,7 @@ function cancelRun() {
176
  }
177
  }
178
  function invalidateRouting() {
179
- // Routing and Clustering allow edits while waiting. An old request must not
180
  // publish into new inputs, even when its submission response arrives late.
181
  if(busy) {
182
  cancelRun();
@@ -190,16 +199,17 @@ function invalidateRouting() {
190
  $('error').hidden=true;
191
  $('result-meta').classList.remove('stale-note');
192
  renderEmpty();
193
- setModelState(cachedTasks.includes(catalog[selected]?.model_task || selected)?'Ready on CPU':selected==='clustering'?'Ready to group':'Ready to route');
194
  }
195
  async function analyze() {
196
  if (busy) return;
197
  const taskId=selected, input=$('input').value, candidates=$('candidates').value.split('\n').map(t=>t.trim()).filter(Boolean);
 
198
  if (!input.trim()) { showError(taskId==='clustering'?'Add a collection of texts or choose an example first.':'Write a prompt or choose an example first.'); $('input').focus(); return; }
199
  if (['embedding','reranker'].includes(taskId) && (!candidates.length || candidates.length>6)) { showError(`Add between 1 and 6 ${taskId==='embedding'?'candidates':'passages'}, one per line.`); return; }
200
  if(taskId==='routing') { const error=RoutingUI.validate(); if(error){showError(error);return;} }
201
  if(taskId==='clustering') { const error=ClusteringUI.validate(); if(error){showError(error);return;} }
202
- const payload=taskId==='clustering'?{task:taskId,texts:ClusteringUI.texts()}:{task:taskId,text:input,candidates:['embedding','reranker'].includes(taskId)?candidates:[]};
203
  if(taskId==='routing')payload.categories=RoutingUI.categories();
204
  $('error').hidden=true; lastResult=null; edited=false; activeJob=null; cancelRequested=false;
205
  const controller=new AbortController(); runController=controller;
@@ -231,6 +241,7 @@ async function analyze() {
231
  let pollFailures=0;
232
  while(job && !cancelRequested && isCurrent() && Date.now()<deadline) {
233
  if(job.status==='completed') {
 
234
  lastResult=job.result; lastInput=input; view='highlight'; clearInterval(runTimer);
235
  renderResult(); setModelState(lastResult.cache_hit?'Example result reused':'Ready on CPU'); break;
236
  }
@@ -275,8 +286,8 @@ async function analyze() {
275
  }
276
  function updateExampleNote() {
277
  const candidates=$('candidates').value.split('\n').map(text=>text.trim()).filter(Boolean);
278
- const match=catalog[selected]?.examples?.findIndex(example=>selected==='clustering'?JSON.stringify(example.texts || [])===JSON.stringify(ClusteringUI.texts()):example.text===$('input').value && (selected==='embedding'?(example.candidates || []).join('\n')===$('candidates').value:JSON.stringify(example.candidates || [])===JSON.stringify(candidates)));
279
- if(['embedding','clustering'].includes(selected))$('examples').querySelectorAll('.example-button').forEach(button=>button.setAttribute('aria-pressed',String(Number(button.dataset.exampleIndex)===match)));
280
  const message=catalog[selected]?.example_notes?.[String(match)] || '';
281
  const note=$('example-note'); note.hidden=!message; note.textContent=message;
282
  }
@@ -291,6 +302,7 @@ function renderResult() {
291
  else if (['ranking','similarity'].includes(result.kind)) renderRanking(result);
292
  else if (result.kind==='routing') RoutingUI.renderResult(result);
293
  else if (result.kind==='clustering') ClusteringUI.renderResult(result);
 
294
  else { showError('This result format could not be displayed. Please try again.'); return; }
295
  if(!['routing','clustering'].includes(result.kind)) {
296
  updateResultConnections(result);
@@ -303,6 +315,10 @@ function renderResult() {
303
  }
304
  function updateResultConnections(result) {
305
  if(lastResult?.result!==result || selected==='routing') return;
 
 
 
 
306
  if(result.kind==='pii') {
307
  const text=$('output').querySelector('.pii-text,.redacted') || $('output');
308
  FlowUI.setTargets([{element:text,winner:!edited}]);
@@ -357,8 +373,9 @@ function renderRanking(result) {
357
  if(result.kind==='similarity') $('output').append(textNode('div',`${result.dimensions} dimensions · normalized embeddings`,'result-caption'));
358
  }
359
  async function copyText(value,button) { try {await navigator.clipboard.writeText(value); const before=button.textContent;button.textContent='Copied';setTimeout(()=>button.textContent=before,1500);}catch{showError('Clipboard access is unavailable. Select the result text to copy it.');} }
360
- $('run-button').addEventListener('click',analyze); $('clear-button').addEventListener('click',()=>{$('input').value='';$('candidates').value='';inputChanged();$('input').focus();});
361
  $('input').addEventListener('input',inputChanged);$('candidates').addEventListener('input',inputChanged);
 
362
  $('similarity-group').addEventListener('change',selectSimilarityGroup);
363
  $('highlight-button').addEventListener('click',()=>{view='highlight';renderResult();});$('redact-button').addEventListener('click',()=>{view='redact';renderResult();});
364
  $('copy-button').addEventListener('click',()=>{if(lastResult)copyText(JSON.stringify(lastResult.result.kind==='clustering'?ClusteringUI.exportResult():lastResult.result,null,2),$('copy-button'));});
 
8
  clustering: {name:'Clustering',group:'Retrieve',icon:'⁙',title:'Shared meaning. Natural company.',description:'Discover groups in a collection of texts, without predefined categories.',action:'Find the groups',note:'Adjust granularity to merge or separate groups. Open a group to inspect every text.',examples:[]},
9
  domain: {name:'Domain',group:'Understand',icon:'◈',title:'Every question has a world.',description:'Find the subject behind a request, across 14 knowledge domains.',action:'Find the domain',note:'Scores describe the model’s distribution across its domain labels.',examples:['A curious question','Another field']},
10
  factcheck: {name:'Fact check',group:'Understand',icon:'⊙',title:'Know when to look it up.',description:'Recognize requests that call for external factual knowledge.',action:'Check the need',note:'This predicts whether a fact check is needed. It does not verify whether a claim is true.',examples:['Factual knowledge','Creative writing']},
11
+ hallucination: {name:'Hallucination',group:'Understand',icon:'⌁',title:'An answer, grounded in evidence.',description:'Compare an answer with its evidence. Find passages that may not be supported.',action:'Check the answer',note:'Highlights reflect the evidence you provide, not a guarantee of real-world truth.',examples:[]},
12
  modality: {name:'Modality',group:'Understand',icon:'▧',title:'The right medium for the idea.',description:'Explore whether a request calls for text, an image, or both.',action:'Find the medium',note:'AR means text generation; DIFFUSION means image generation; BOTH means both.',examples:['An image idea','Words and pictures']},
13
  feedback: {name:'Feedback',group:'Understand',icon:'☷',title:'Listen a little closer.',description:'Understand the feedback hidden in a conversational response.',action:'Read the feedback',note:'Predictions distinguish satisfaction, clarification, correction, a different answer, and no feedback.',examples:['A little gratitude','A correction']},
14
  pii: {name:'Personal data',group:'Protect',icon:'◎',title:'Find what should stay private.',description:'Locate personal information, then see it highlighted or redacted.',action:'Detect PII',note:'Highlights show model predictions. Review them before relying on redaction.',examples:["Contact details","Date & age","Street address","Network log","Test card","Synthetic ID","Work profile"]},
 
36
  FlowUI.setTargets([]);
37
  if(selected==='routing' && catalog.routing) { RoutingUI.renderEmpty(); return; }
38
  if(selected==='clustering') { ClusteringUI.renderEmpty(); return; }
39
+ if(selected==='hallucination') { HallucinationUI.renderEmpty(); return; }
40
  $('output').innerHTML = '<div class="empty-state"><span class="flow-empty-label">A SIGNAL, WAITING TO BE FOUND</span><strong>Ready when you are.</strong><p>Choose an example or bring your own text. Run the model to see what it notices.</p></div>';
41
  $('result-meta').textContent=''; $('output-tabs').hidden=true; $('copy-button').hidden=true;
42
  $('output').setAttribute('aria-live','polite');
43
  }
44
  function setModelState(message) { $('model-state').replaceChildren(textNode('span','','hollow-dot'),textNode('span',message)); }
45
  function selectTask(id) {
46
+ if (!presentation[id] || (busy && !['routing','clustering','hallucination'].includes(selected))) return;
47
  if(busy) invalidateRouting();
48
  RoutingUI.leave();
49
  ClusteringUI.leave();
50
+ HallucinationUI.leave();
51
  FlowUI.leave();
52
  selected=id; lastResult=null; edited=false; view='highlight'; $('error').hidden=true;
53
  $('result-meta').classList.remove('stale-note');
 
58
  $('task-title').textContent=p.title; $('task-description').textContent=p.description; $('task-symbol').textContent=p.icon;
59
  $('run-label').textContent=p.action; $('task-note').textContent=p.note;
60
  const retrieval=['embedding','reranker'].includes(id);
61
+ $('candidates-wrap').hidden=!retrieval; $('input-label').textContent=id==='routing'?'YOUR PROMPT':id==='clustering'?'YOUR TEXTS':id==='hallucination'?'YOUR QUESTION':retrieval?'YOUR QUERY':'YOUR TEXT';
62
  $('candidates-label').textContent=id==='embedding'?'CANDIDATES':'PASSAGES';
63
+ $('input-hint').textContent=id==='clustering'?'3–24 texts · 128 tokens each':id==='hallucination'?'512 tokens total · inputs are never shortened':'Up to 512 tokens';
64
+ if(id==='hallucination')$('input').removeAttribute('maxlength');
65
+ else $('input').maxLength=id==='clustering'?98351:12000;
66
+ $('input').placeholder=id==='clustering'?'One text per line. Let shared meaning find its shape.':id==='hallucination'?'What question is the answer responding to?':'Write a little. Discover a lot.';
67
+ document.querySelector('.examples>span').textContent=id==='clustering'?'TRY A COLLECTION':id==='hallucination'?'TRY AN EXAMPLE':'TRY A PROMPT';
68
+ $('input').value=''; $('candidates').value=''; HallucinationUI.clear();
69
  $('model-link').textContent=`Vela-1.0 · ${task ? task.model.split('-').pop() : p.name} ↗`;
70
  $('model-link').href=`https://huggingface.co/${task?.model || 'collections/llm-semantic-router/vela-10'}`;
71
  setModelState(cachedTasks.includes(task?.model_task || id)?'Ready on CPU':'Loads on first run');
 
78
  renderExamples();
79
  FlowUI.enter(id,p,task);
80
  if(id==='clustering')ClusteringUI.enter();
81
+ if(id==='hallucination')HallucinationUI.enter();
82
  if(id==='routing' && task?.scenarios) RoutingUI.enter(task);
83
  else if (task?.examples?.length) loadExample(task.examples.find(example=>!grouped || example.group===similarityGroup));
84
  countCharacters(); renderEmpty();
85
  const activeButton=$('tasks').querySelector(`[data-task="${id}"]`), nav=$('tasks');
86
  if(activeButton && nav.scrollWidth>nav.clientWidth) nav.scrollLeft+=activeButton.getBoundingClientRect().left-nav.getBoundingClientRect().left-(nav.clientWidth-activeButton.clientWidth)/2;
87
  }
88
+ function countCharacters() { $('char-count').textContent=selected==='hallucination'?`${HallucinationUI.characterCount().toLocaleString()} characters total`:selected==='clustering'?`${ClusteringUI.texts().length} / 24 texts`:`${Array.from($('input').value).length.toLocaleString()} characters`; }
89
  function inputChanged() {
90
  countCharacters(); updateExampleNote(); $('error').hidden=true;
91
  if(selected==='routing'){RoutingUI.inputChanged();return;}
92
+ if(['clustering','hallucination'].includes(selected)){invalidateRouting();return;}
93
  if(selected==='embedding'){
94
  lastResult=null;edited=false;$('result-meta').classList.remove('stale-note');renderEmpty();
95
  if(!busy)setModelState(cachedTasks.includes('embedding')?'Ready on CPU':'Loads on first run');
 
113
  }
114
  }
115
  function loadExample(example) {
116
+ if(!example || (busy && !['routing','clustering','hallucination'].includes(selected)))return;
117
+ $('input').value=selected==='clustering'?(example.texts || []).join('\n'):example.text;$('candidates').value=(example.candidates || []).join('\n');
118
+ if(selected==='hallucination')HallucinationUI.fillExample(example);
119
+ inputChanged();
120
  }
121
  function selectSimilarityGroup() {
122
  if(selected!=='embedding')return;
 
127
  }
128
  function setBusy(value) {
129
  busy=value; document.body.classList.toggle('is-running',value);
130
+ const editable=['routing','clustering','hallucination'].includes(selected);
131
  for (const el of document.querySelectorAll('.task-button,.example-button,#clear-button,#similarity-group')) el.disabled=value && !editable;
132
  $('run-button').disabled=value;
133
  $('input').readOnly=value && !editable; $('candidates').readOnly=value;
134
+ $('hallucination-context').readOnly=value && !editable; $('hallucination-answer').readOnly=value && !editable;
135
  $('run-label').textContent=value?'Finding the signal…':presentation[selected].action;
136
  $('run-arrow').textContent=value?'◌':'↗'; $('output').setAttribute('aria-busy',String(value));
137
  }
 
143
  $('output-tabs').hidden=true; $('copy-button').hidden=true; $('result-meta').textContent='';
144
  $('output').innerHTML='<div class="empty-state loading"><span class="flow-empty-label">FOLLOWING THE SIGNAL</span><strong id="loading-title">Finding you a place…</strong><p id="loading-copy">Your request will wait its turn on the shared CPU.</p><div class="loading-track" aria-hidden="true"></div><button id="cancel-button" class="text-button cancel-button" type="button">Cancel request</button></div>';
145
  if(selected==='clustering')$('loading-title').textContent='Gathering shared meaning…';
146
+ if(selected==='hallucination')$('loading-title').textContent='Reading against the evidence…';
147
  $('cancel-button').addEventListener('click',cancelRun);
148
  setModelState('Submitting request…');
149
  }
 
185
  }
186
  }
187
  function invalidateRouting() {
188
+ // Editable tasks allow changes while waiting. An old request must not
189
  // publish into new inputs, even when its submission response arrives late.
190
  if(busy) {
191
  cancelRun();
 
199
  $('error').hidden=true;
200
  $('result-meta').classList.remove('stale-note');
201
  renderEmpty();
202
+ setModelState(cachedTasks.includes(catalog[selected]?.model_task || selected)?'Ready on CPU':selected==='hallucination'?'Loads on first run':selected==='clustering'?'Ready to group':'Ready to route');
203
  }
204
  async function analyze() {
205
  if (busy) return;
206
  const taskId=selected, input=$('input').value, candidates=$('candidates').value.split('\n').map(t=>t.trim()).filter(Boolean);
207
+ if(taskId==='hallucination') { const error=HallucinationUI.validate(); if(error){showError(error);return;} }
208
  if (!input.trim()) { showError(taskId==='clustering'?'Add a collection of texts or choose an example first.':'Write a prompt or choose an example first.'); $('input').focus(); return; }
209
  if (['embedding','reranker'].includes(taskId) && (!candidates.length || candidates.length>6)) { showError(`Add between 1 and 6 ${taskId==='embedding'?'candidates':'passages'}, one per line.`); return; }
210
  if(taskId==='routing') { const error=RoutingUI.validate(); if(error){showError(error);return;} }
211
  if(taskId==='clustering') { const error=ClusteringUI.validate(); if(error){showError(error);return;} }
212
+ const payload=taskId==='hallucination'?{task:taskId,text:input,...HallucinationUI.values()}:taskId==='clustering'?{task:taskId,texts:ClusteringUI.texts()}:{task:taskId,text:input,candidates:['embedding','reranker'].includes(taskId)?candidates:[]};
213
  if(taskId==='routing')payload.categories=RoutingUI.categories();
214
  $('error').hidden=true; lastResult=null; edited=false; activeJob=null; cancelRequested=false;
215
  const controller=new AbortController(); runController=controller;
 
241
  let pollFailures=0;
242
  while(job && !cancelRequested && isCurrent() && Date.now()<deadline) {
243
  if(job.status==='completed') {
244
+ if(taskId==='hallucination' && (job.result?.result?.kind!=='hallucination' || job.result.result.answer!==payload.answer)) throw new Error('The result did not match this answer. Please run it again.');
245
  lastResult=job.result; lastInput=input; view='highlight'; clearInterval(runTimer);
246
  renderResult(); setModelState(lastResult.cache_hit?'Example result reused':'Ready on CPU'); break;
247
  }
 
286
  }
287
  function updateExampleNote() {
288
  const candidates=$('candidates').value.split('\n').map(text=>text.trim()).filter(Boolean);
289
+ const match=catalog[selected]?.examples?.findIndex(example=>selected==='hallucination'?HallucinationUI.matchesExample(example):selected==='clustering'?JSON.stringify(example.texts || [])===JSON.stringify(ClusteringUI.texts()):example.text===$('input').value && (selected==='embedding'?(example.candidates || []).join('\n')===$('candidates').value:JSON.stringify(example.candidates || [])===JSON.stringify(candidates)));
290
+ if(['embedding','clustering','hallucination'].includes(selected))$('examples').querySelectorAll('.example-button').forEach(button=>button.setAttribute('aria-pressed',String(Number(button.dataset.exampleIndex)===match)));
291
  const message=catalog[selected]?.example_notes?.[String(match)] || '';
292
  const note=$('example-note'); note.hidden=!message; note.textContent=message;
293
  }
 
302
  else if (['ranking','similarity'].includes(result.kind)) renderRanking(result);
303
  else if (result.kind==='routing') RoutingUI.renderResult(result);
304
  else if (result.kind==='clustering') ClusteringUI.renderResult(result);
305
+ else if (result.kind==='hallucination') HallucinationUI.renderResult(result);
306
  else { showError('This result format could not be displayed. Please try again.'); return; }
307
  if(!['routing','clustering'].includes(result.kind)) {
308
  updateResultConnections(result);
 
315
  }
316
  function updateResultConnections(result) {
317
  if(lastResult?.result!==result || selected==='routing') return;
318
+ if(result.kind==='hallucination') {
319
+ FlowUI.setTargets([{element:$('output').querySelector('.hallucination-answer'),winner:!edited}]);
320
+ return;
321
+ }
322
  if(result.kind==='pii') {
323
  const text=$('output').querySelector('.pii-text,.redacted') || $('output');
324
  FlowUI.setTargets([{element:text,winner:!edited}]);
 
373
  if(result.kind==='similarity') $('output').append(textNode('div',`${result.dimensions} dimensions · normalized embeddings`,'result-caption'));
374
  }
375
  async function copyText(value,button) { try {await navigator.clipboard.writeText(value); const before=button.textContent;button.textContent='Copied';setTimeout(()=>button.textContent=before,1500);}catch{showError('Clipboard access is unavailable. Select the result text to copy it.');} }
376
+ $('run-button').addEventListener('click',analyze); $('clear-button').addEventListener('click',()=>{$('input').value='';$('candidates').value='';HallucinationUI.clear();inputChanged();$('input').focus();});
377
  $('input').addEventListener('input',inputChanged);$('candidates').addEventListener('input',inputChanged);
378
+ $('hallucination-context').addEventListener('input',inputChanged);$('hallucination-answer').addEventListener('input',inputChanged);
379
  $('similarity-group').addEventListener('change',selectSimilarityGroup);
380
  $('highlight-button').addEventListener('click',()=>{view='highlight';renderResult();});$('redact-button').addEventListener('click',()=>{view='redact';renderResult();});
381
  $('copy-button').addEventListener('click',()=>{if(lastResult)copyText(JSON.stringify(lastResult.result.kind==='clustering'?ClusteringUI.exportResult():lastResult.result,null,2),$('copy-button'));});
static/flow.js CHANGED
@@ -1,4 +1,4 @@
1
- /* One shared canvas and model node for all twelve Vela demos. */
2
  'use strict';
3
  window.FlowUI = (() => {
4
  const el = id => document.getElementById(id);
@@ -85,7 +85,7 @@ window.FlowUI = (() => {
85
  const body = node('div', '', 'flow-model-body');
86
  el('model-link').textContent = `Vela ${name} ↗`;
87
  const metadata = node('div', '', 'flow-model-metadata');
88
- const variant = ['embedding', 'routing', 'clustering'].includes(id) ? '768D' : {classification:'Classifier', pii:'Token tags', hazard:'Multi-label', ranking:'Cross-encoder'}[task?.kind] || 'Encoder';
89
  metadata.append(node('span', '307M'), node('i', '·'), node('span', variant));
90
  body.append(el('model-link'), metadata, el('model-state'));
91
  modelCard.append(art, body);
 
1
+ /* One shared canvas and model node for all Vela demos. */
2
  'use strict';
3
  window.FlowUI = (() => {
4
  const el = id => document.getElementById(id);
 
85
  const body = node('div', '', 'flow-model-body');
86
  el('model-link').textContent = `Vela ${name} ↗`;
87
  const metadata = node('div', '', 'flow-model-metadata');
88
+ const variant = ['embedding', 'routing', 'clustering'].includes(id) ? '768D' : {classification:'Classifier', pii:'Token tags', hazard:'Multi-label', hallucination:'Grounding', ranking:'Cross-encoder'}[task?.kind] || 'Encoder';
89
  metadata.append(node('span', '307M'), node('i', '·'), node('span', variant));
90
  body.append(el('model-link'), metadata, el('model-state'));
91
  modelCard.append(art, body);
static/hallucination.css ADDED
@@ -0,0 +1,75 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /* Three source fields and quiet copper annotations share the Vela field guide. */
2
+ .hallucination-mode #input {
3
+ min-height: 82px;
4
+ max-height: 170px;
5
+ font-size: 13px;
6
+ }
7
+ .hallucination-fields label {
8
+ display: block;
9
+ border-top: 1px solid #e2e7dc;
10
+ padding: 12px 16px 0;
11
+ color: #61725a;
12
+ font: 8px var(--mono);
13
+ letter-spacing: 1px;
14
+ }
15
+ .hallucination-mode .hallucination-fields textarea {
16
+ display: block;
17
+ min-height: 110px;
18
+ max-height: 240px;
19
+ width: 100%;
20
+ padding: 10px 16px 14px;
21
+ font-size: 12px;
22
+ line-height: 1.8;
23
+ resize: vertical;
24
+ }
25
+ .hallucination-mode .editor-foot {
26
+ align-items: flex-start;
27
+ gap: 12px;
28
+ line-height: 1.65;
29
+ }
30
+ .hallucination-mode #char-count { white-space: nowrap; }
31
+ .hallucination-mode .hallucination-result-title {
32
+ font-size: 25px;
33
+ line-height: 1.3;
34
+ }
35
+ .hallucination-answer {
36
+ border: 1px solid #d9dfcf;
37
+ border-radius: 5px;
38
+ padding: 18px;
39
+ background: #fcfaf5;
40
+ color: var(--ink);
41
+ font-size: 13px;
42
+ line-height: 2;
43
+ white-space: pre-wrap;
44
+ overflow-wrap: anywhere;
45
+ max-height: 520px;
46
+ overflow: auto;
47
+ }
48
+ .hallucination-span {
49
+ background: #ead5b8;
50
+ color: #5b402b;
51
+ border-bottom: 1px solid #b78053;
52
+ border-radius: 2px;
53
+ box-decoration-break: clone;
54
+ -webkit-box-decoration-break: clone;
55
+ }
56
+ .hallucination-legend {
57
+ margin-top: 12px;
58
+ padding-left: 11px;
59
+ border-left: 3px solid #bb895d;
60
+ color: #795437;
61
+ font-size: 9px;
62
+ line-height: 1.7;
63
+ }
64
+ .hallucination-result-note {
65
+ margin: 13px 0 0;
66
+ color: #66735d;
67
+ font-size: 10px;
68
+ line-height: 1.8;
69
+ }
70
+ @media (max-width: 700px) {
71
+ .hallucination-mode #input { min-height: 90px; font-size: 14px; }
72
+ .hallucination-mode .hallucination-fields textarea { font-size: 14px; min-height: 115px; }
73
+ .hallucination-answer { padding: 15px; }
74
+ .hallucination-mode .editor-foot { flex-direction: column; gap: 3px; }
75
+ }
static/hallucination.js ADDED
@@ -0,0 +1,94 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /* Grounding highlights use the answer's Unicode code-point offsets. */
2
+ 'use strict';
3
+ window.HallucinationUI = (() => {
4
+ const el = id => document.getElementById(id);
5
+ const node = (tag, text = '', className = '') => {
6
+ const element = document.createElement(tag);
7
+ element.textContent = text;
8
+ if (className) element.className = className;
9
+ return element;
10
+ };
11
+ function enter() {
12
+ document.querySelector('.task-main').classList.add('hallucination-mode');
13
+ el('hallucination-fields').hidden = false;
14
+ }
15
+ function leave() {
16
+ document.querySelector('.task-main').classList.remove('hallucination-mode');
17
+ el('hallucination-fields').hidden = true;
18
+ }
19
+ function values() {
20
+ return {context: el('hallucination-context').value, answer: el('hallucination-answer').value};
21
+ }
22
+ function fillExample(example) {
23
+ el('hallucination-context').value = example.context || '';
24
+ el('hallucination-answer').value = example.answer || '';
25
+ }
26
+ function clear() { fillExample({}); }
27
+ function matchesExample(example) {
28
+ const current = values();
29
+ return example.text === el('input').value && example.context === current.context && example.answer === current.answer;
30
+ }
31
+ function validate() {
32
+ for (const [id, message] of [
33
+ ['input', 'Add the question or choose an example first.'],
34
+ ['hallucination-context', 'Add the evidence the answer should be based on.'],
35
+ ['hallucination-answer', 'Add the answer you want to check.'],
36
+ ]) {
37
+ if (!el(id).value.trim()) { el(id).focus(); return message; }
38
+ }
39
+ return null;
40
+ }
41
+ function characterCount() {
42
+ const current = values();
43
+ return [el('input').value, current.context, current.answer].reduce((total, value) => total + Array.from(value).length, 0);
44
+ }
45
+ function renderEmpty() {
46
+ const empty = node('div', '', 'empty-state hallucination-empty');
47
+ empty.append(node('span', 'AN ANSWER, IN THE LIGHT OF EVIDENCE', 'flow-empty-label'));
48
+ empty.append(node('strong', 'See what the evidence supports.'));
49
+ empty.append(node('p', 'Add a question, its evidence, and an answer. Potentially unsupported passages will appear here.'));
50
+ el('output').replaceChildren(empty);
51
+ el('result-meta').textContent = '';
52
+ el('output-tabs').hidden = true;
53
+ el('copy-button').hidden = true;
54
+ el('output').setAttribute('aria-live', 'polite');
55
+ }
56
+ function segments(answer, spans) {
57
+ if (typeof answer !== 'string' || !answer.trim() || !Array.isArray(spans)) throw new Error('The answer highlights could not be read. Please try again.');
58
+ const chars = Array.from(answer), result = [];
59
+ let previous = 0;
60
+ for (const span of [...spans].sort((a, b) => a.start - b.start)) {
61
+ if (!Number.isInteger(span.start) || !Number.isInteger(span.end) || span.start < previous || span.end <= span.start || span.end > chars.length || !Number.isFinite(span.score) || span.score < 0 || span.score > 1) {
62
+ throw new Error('The answer highlights did not align with the text. Please try again.');
63
+ }
64
+ const text = chars.slice(span.start, span.end).join('');
65
+ if (span.text !== undefined && span.text !== text) throw new Error('The answer highlights did not match the text. Please try again.');
66
+ if (span.start > previous) result.push({text: chars.slice(previous, span.start).join(''), highlighted: false});
67
+ result.push({text, highlighted: true, score: span.score});
68
+ previous = span.end;
69
+ }
70
+ if (previous < chars.length) result.push({text: chars.slice(previous).join(''), highlighted: false});
71
+ return result;
72
+ }
73
+ function renderResult(result) {
74
+ // Validate the entire response before creating a reassuring empty-span state.
75
+ const parts = segments(result.answer, result.spans);
76
+ const count = result.spans.length;
77
+ const title = count ? `${count} potentially unsupported ${count === 1 ? 'span' : 'spans'}` : 'No unsupported spans detected';
78
+ const heading = node('div', title, 'result-title hallucination-result-title');
79
+ const answer = node('div', '', 'hallucination-answer');
80
+ answer.setAttribute('aria-label', 'Answer with potentially unsupported spans highlighted');
81
+ for (const part of parts) {
82
+ if (!part.highlighted) { answer.append(document.createTextNode(part.text)); continue; }
83
+ const mark = node('mark', part.text, 'hallucination-span');
84
+ mark.title = `Potentially unsupported · model score ${(part.score * 100).toFixed(1)}%`;
85
+ answer.append(mark);
86
+ }
87
+ const note = node('p', 'Assessed against the evidence you provided. This does not guarantee factual accuracy beyond that evidence.', 'hallucination-result-note');
88
+ const children = [heading, answer];
89
+ if (count) children.push(node('div', 'Highlighted · potentially unsupported by the evidence', 'hallucination-legend'));
90
+ children.push(note);
91
+ el('output').replaceChildren(...children);
92
+ }
93
+ return {enter, leave, values, fillExample, clear, matchesExample, validate, characterCount, renderEmpty, renderResult, segments};
94
+ })();
static/index.html CHANGED
@@ -4,7 +4,7 @@
4
  <meta charset="utf-8">
5
  <meta name="viewport" content="width=device-width, initial-scale=1">
6
  <meta name="theme-color" content="#f5f2e9">
7
- <meta name="description" content="Explore Vela 1.0: ten specialized models for understanding, protecting, and retrieving language.">
8
  <title>Vela Studio — A better direction.</title>
9
  <link rel="icon" href="/static/favicon.svg" type="image/svg+xml">
10
  <link rel="preconnect" href="https://fonts.googleapis.com">
@@ -14,10 +14,12 @@
14
  <link rel="stylesheet" href="/static/flow.css?v=e7df90552db0">
15
  <link rel="stylesheet" href="/static/routing.css?v=d2bea5f416">
16
  <link rel="stylesheet" href="/static/clustering.css?v=1a774c530cb1">
17
- <script src="/static/flow.js?v=7b6d347f2d93" defer></script>
 
 
18
  <script src="/static/routing.js?v=c3f8df099d" defer></script>
19
  <script src="/static/clustering.js?v=531c6c50bc4f" defer></script>
20
- <script src="/static/app.js?v=0cd9c7ee6f1a" defer></script>
21
  </head>
22
  <body>
23
  <a class="skip-link" href="#studio">Skip to playground</a>
@@ -41,7 +43,7 @@
41
  <span class="art-caption">FIG. 01 — A LITTLE INTELLIGENCE GOES A LONG WAY.</span>
42
  </div>
43
  </section>
44
- <div class="specimen-strip"><span><b>307M</b> parameters per model</span><span><b>10</b> specialized models</span><span><i class="status-dot"></i> Open weights, open possibilities</span><a href="https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M" target="_blank" rel="noopener">Meet the encoder <span>↗</span></a></div>
45
  <main id="studio" tabindex="-1">
46
  <div class="section-title"><div><span class="eyebrow">THE INTERACTIVE FIELD GUIDE</span><h2>The signal room<span>.</span></h2></div><div class="session-badge"><i class="status-dot"></i><span>Vela 1.0</span><span class="badge-divider">/</span> CPU</div></div>
47
  <div class="workbench">
@@ -51,7 +53,7 @@
51
  <div class="model-strip"><a id="model-link" target="_blank" rel="noopener">Vela-1.0 · Embedding <span>↗</span></a><span id="model-state" role="status" aria-live="polite"><span class="hollow-dot"></span> Loads on first run</span></div>
52
  <section id="routing-scenarios" class="routing-scenarios" aria-label="Routing scenario" hidden></section>
53
  <div class="panels">
54
- <section class="input-panel"><div class="panel-heading"><label id="input-label" for="input">YOUR PROMPT</label><button class="text-button" id="clear-button">Clear <span>×</span></button></div><textarea id="input" maxlength="12000" spellcheck="false" placeholder="Write a little. Discover a lot." aria-describedby="input-hint"></textarea><div id="candidates-wrap" hidden><label for="candidates"><span id="candidates-label">PASSAGES</span> <span>One per line · up to 6</span></label><textarea id="candidates" maxlength="36000" rows="5" spellcheck="false"></textarea></div><div class="editor-foot"><span id="input-hint">Up to 512 tokens</span><span id="char-count">0 characters</span></div></section>
55
  <section class="output-panel"><div class="panel-heading"><span>THE SIGNAL</span><div id="output-tabs" class="output-tabs" aria-label="PII output view" hidden><button id="highlight-button" aria-pressed="true">Highlight</button><button id="redact-button" aria-pressed="false">Redact</button></div><button class="text-button" id="copy-button" hidden>Copy</button></div><div id="output" class="output" aria-live="polite" aria-atomic="true"></div><div id="result-meta" class="result-meta"></div></section>
56
  </div>
57
  <div id="example-note" class="example-note" role="note" hidden></div>
 
4
  <meta charset="utf-8">
5
  <meta name="viewport" content="width=device-width, initial-scale=1">
6
  <meta name="theme-color" content="#f5f2e9">
7
+ <meta name="description" content="Explore Vela 1.0: eleven specialized models for understanding, protecting, grounding, and retrieving language.">
8
  <title>Vela Studio — A better direction.</title>
9
  <link rel="icon" href="/static/favicon.svg" type="image/svg+xml">
10
  <link rel="preconnect" href="https://fonts.googleapis.com">
 
14
  <link rel="stylesheet" href="/static/flow.css?v=e7df90552db0">
15
  <link rel="stylesheet" href="/static/routing.css?v=d2bea5f416">
16
  <link rel="stylesheet" href="/static/clustering.css?v=1a774c530cb1">
17
+ <link rel="stylesheet" href="/static/hallucination.css?v=ddd852a221a6">
18
+ <script src="/static/hallucination.js?v=76e36d418177" defer></script>
19
+ <script src="/static/flow.js?v=0d15a0d480c5" defer></script>
20
  <script src="/static/routing.js?v=c3f8df099d" defer></script>
21
  <script src="/static/clustering.js?v=531c6c50bc4f" defer></script>
22
+ <script src="/static/app.js?v=8eeeb11f8795" defer></script>
23
  </head>
24
  <body>
25
  <a class="skip-link" href="#studio">Skip to playground</a>
 
43
  <span class="art-caption">FIG. 01 — A LITTLE INTELLIGENCE GOES A LONG WAY.</span>
44
  </div>
45
  </section>
46
+ <div class="specimen-strip"><span><b>307M</b> parameters per model</span><span><b>11</b> specialized models</span><span><i class="status-dot"></i> Open weights, open possibilities</span><a href="https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M" target="_blank" rel="noopener">Meet the encoder <span>↗</span></a></div>
47
  <main id="studio" tabindex="-1">
48
  <div class="section-title"><div><span class="eyebrow">THE INTERACTIVE FIELD GUIDE</span><h2>The signal room<span>.</span></h2></div><div class="session-badge"><i class="status-dot"></i><span>Vela 1.0</span><span class="badge-divider">/</span> CPU</div></div>
49
  <div class="workbench">
 
53
  <div class="model-strip"><a id="model-link" target="_blank" rel="noopener">Vela-1.0 · Embedding <span>↗</span></a><span id="model-state" role="status" aria-live="polite"><span class="hollow-dot"></span> Loads on first run</span></div>
54
  <section id="routing-scenarios" class="routing-scenarios" aria-label="Routing scenario" hidden></section>
55
  <div class="panels">
56
+ <section class="input-panel"><div class="panel-heading"><label id="input-label" for="input">YOUR PROMPT</label><button class="text-button" id="clear-button">Clear <span>×</span></button></div><textarea id="input" maxlength="12000" spellcheck="false" placeholder="Write a little. Discover a lot." aria-describedby="input-hint"></textarea><div id="hallucination-fields" class="hallucination-fields" hidden><label for="hallucination-context">EVIDENCE</label><textarea id="hallucination-context" rows="4" spellcheck="false" placeholder="Paste the source material the answer should be based on." aria-describedby="input-hint"></textarea><label for="hallucination-answer">ANSWER TO CHECK</label><textarea id="hallucination-answer" rows="4" spellcheck="false" placeholder="Paste an answer to compare with the evidence." aria-describedby="input-hint"></textarea></div><div id="candidates-wrap" hidden><label for="candidates"><span id="candidates-label">PASSAGES</span> <span>One per line · up to 6</span></label><textarea id="candidates" maxlength="36000" rows="5" spellcheck="false"></textarea></div><div class="editor-foot"><span id="input-hint">Up to 512 tokens</span><span id="char-count">0 characters</span></div></section>
57
  <section class="output-panel"><div class="panel-heading"><span>THE SIGNAL</span><div id="output-tabs" class="output-tabs" aria-label="PII output view" hidden><button id="highlight-button" aria-pressed="true">Highlight</button><button id="redact-button" aria-pressed="false">Redact</button></div><button class="text-button" id="copy-button" hidden>Copy</button></div><div id="output" class="output" aria-live="polite" aria-atomic="true"></div><div id="result-meta" class="result-meta"></div></section>
58
  </div>
59
  <div id="example-note" class="example-note" role="note" hidden></div>
tests/test_hallucination.py ADDED
@@ -0,0 +1,238 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Evidence-pair, Unicode span, queue cache, and API contracts without weights."""
2
+
3
+ from copy import deepcopy
4
+ import math
5
+ from types import SimpleNamespace
6
+ import unittest
7
+ from unittest.mock import Mock, patch
8
+
9
+ from fastapi.testclient import TestClient
10
+
11
+ import app as server
12
+ from catalog import TASKS
13
+ from hallucination import evidence_prompt, hallucination_result, validate_hallucination_model
14
+ from hallucination_catalog import HALLUCINATION_EXAMPLES
15
+ from inference import InferenceEngine, InputLimitError
16
+ from job_queue import InferenceQueue
17
+
18
+
19
+ def public_request(index=0):
20
+ example = HALLUCINATION_EXAMPLES[index]
21
+ return {"task": "hallucination", **{key: example[key] for key in ("text", "context", "answer")}}
22
+
23
+
24
+ class HallucinationSpanTests(unittest.TestCase):
25
+ def test_only_answer_tokens_contribute_and_unicode_offsets_preserve_original(self):
26
+ answer = "😀 李雷在上海。"
27
+ result = hallucination_result(answer,
28
+ [None, 0, None, 1, 1, 1, 1, 1, 1, None],
29
+ [[0, 0], [0, 100], [0, 0], [0, 1], [2, 3], [3, 4], [4, 5], [5, 7], [7, 8], [0, 0]],
30
+ [1, .99, 1, .1, .8, .9, .5, .7, .2, 1])
31
+ self.assertEqual(result["kind"], "hallucination")
32
+ self.assertEqual(result["answer"], answer)
33
+ self.assertEqual(result["spans"], [
34
+ {"start": 2, "end": 4, "text": "李雷", "score": .9},
35
+ {"start": 5, "end": 7, "text": "上海", "score": .7},
36
+ ])
37
+ self.assertEqual(result["max_hallucination_score"], .9)
38
+
39
+ def test_strict_threshold_and_supported_token_separate_spans(self):
40
+ result = hallucination_result("ABC", [1, 1, 1], [[0, 1], [1, 2], [2, 3]], [.8, .5, .9])
41
+ self.assertEqual([span["text"] for span in result["spans"]], ["A", "C"])
42
+ self.assertEqual(hallucination_result("A", [1], [[0, 1]], [.5])["spans"], [])
43
+
44
+ def test_adjacent_positive_tokens_keep_internal_space_and_byte_fallback(self):
45
+ result = hallucination_result("𐍈 B", [1] * 5,
46
+ [[0, 1]] * 4 + [[2, 3]], [.7, .8, .9, .8, .6])
47
+ self.assertEqual(result["spans"], [{"start": 0, "end": 3, "text": "𐍈 B", "score": .9}])
48
+
49
+ def test_alternating_byte_fallback_scores_union_only_overlapping_characters(self):
50
+ result = hallucination_result("𐍈", [1] * 4, [[0, 1]] * 4, [.8, .1, .9, .5])
51
+ self.assertEqual(result["spans"], [{"start": 0, "end": 1, "text": "𐍈", "score": .9}])
52
+ # A supported byte token closes its span. A later positive token over
53
+ # another character remains distinct, even when the ranges just touch.
54
+ separate = hallucination_result("𐍈B", [1] * 3, [[0, 1], [0, 1], [1, 2]], [.8, .5, .9])
55
+ self.assertEqual(separate["spans"], [
56
+ {"start": 0, "end": 1, "text": "𐍈", "score": .8},
57
+ {"start": 1, "end": 2, "text": "B", "score": .9},
58
+ ])
59
+
60
+ def test_invalid_scores_offsets_or_empty_answer_mapping_fail_closed(self):
61
+ cases = [([1], [[0, 2]], [.7]), ([1], [[-1, 1]], [.7]),
62
+ ([1], [[0, 1]], [math.nan]), ([1], [[0, 1]], [1.1]),
63
+ ([0], [[0, 1]], [.9]), ([1, 1], [[0, 1]], [.7])]
64
+ for sequences, offsets, scores in cases:
65
+ with self.subTest(case=(sequences, offsets, scores)), self.assertRaises(ValueError):
66
+ hallucination_result("A", sequences, offsets, scores)
67
+
68
+
69
+ class HallucinationAPITests(unittest.TestCase):
70
+ def setUp(self):
71
+ self.engine_patch = patch.object(server, "engine", InferenceEngine())
72
+ self.engine_patch.start()
73
+ self.addCleanup(self.engine_patch.stop)
74
+ self.queue = InferenceQueue(server.run_inference, tasks=TASKS)
75
+ self.queue_patch = patch.object(server, "jobs", self.queue)
76
+ self.queue_patch.start()
77
+ self.addCleanup(self.queue_patch.stop)
78
+ self.client = TestClient(server.app).__enter__()
79
+ self.addCleanup(self.client.__exit__, None, None, None)
80
+
81
+ def test_all_three_original_fields_reach_both_endpoints_unchanged(self):
82
+ payload = {"task": "hallucination", "text": " Question?\n", "context": " Evidence. ",
83
+ "answer": " 😀 Answer.\n"}
84
+ with patch.object(server.engine, "analyze_locked", return_value={"result": {"kind": "hallucination"}}) as analyze:
85
+ created = self.client.post("/api/jobs", json=payload)
86
+ self.assertEqual(created.status_code, 202)
87
+ self.assertEqual(self.queue.wait(created.json()["id"])["status"], "completed")
88
+ self.assertEqual(analyze.call_args.args, ("hallucination", payload["text"], []))
89
+ self.assertEqual(analyze.call_args.kwargs["context"], payload["context"])
90
+ self.assertEqual(analyze.call_args.kwargs["answer"], payload["answer"])
91
+ self.assertEqual(self.client.post("/api/analyze", json=payload).status_code, 200)
92
+
93
+ def test_required_fields_char_limits_and_unrelated_fields_reject_without_echo(self):
94
+ invalid = []
95
+ for field in ("text", "context", "answer"):
96
+ for value in ("", " \n", "x" * 16001, 123, None):
97
+ invalid.append(dict(public_request(), **{field: value}))
98
+ missing = public_request()
99
+ del missing[field]
100
+ invalid.append(missing)
101
+ invalid += [dict(public_request(), candidates=["passage"]),
102
+ dict(public_request(), categories=[{"id": "a", "name": "A"}]),
103
+ dict(public_request(), texts=["first", "second", "third"])]
104
+ with patch.object(server.engine, "analyze_locked") as analyze:
105
+ for payload in invalid:
106
+ for path in ("/api/jobs", "/api/analyze"):
107
+ response = self.client.post(path, json=payload)
108
+ self.assertEqual(response.status_code, 422)
109
+ self.assertNotIn("The museum opens", response.text)
110
+ analyze.assert_not_called()
111
+
112
+ def test_existing_tasks_reject_nonempty_evidence_answer_and_keep_empty_defaults(self):
113
+ with patch.object(server.engine, "analyze_locked", return_value={"result": {}}) as analyze:
114
+ for task, spec in TASKS.items():
115
+ if task == "hallucination":
116
+ continue
117
+ payload = {key: value for key, value in spec["examples"][0].items()
118
+ if key in {"text", "candidates", "categories", "texts"}}
119
+ payload.update(task=task, context="", answer="")
120
+ self.assertEqual(self.client.post("/api/analyze", json=payload).status_code, 200, task)
121
+ for field in ("context", "answer"):
122
+ count = analyze.call_count
123
+ response = self.client.post("/api/analyze", json=dict(payload, **{field: "private@example.com"}))
124
+ self.assertEqual(response.status_code, 422, task)
125
+ self.assertNotIn("private@example.com", response.text)
126
+ self.assertEqual(analyze.call_count, count)
127
+
128
+ def test_combined_token_error_is_a_422_and_releases_worker_slot(self):
129
+ message = "Question + evidence + answer has 513 tokens. This CPU demo accepts at most 512."
130
+ with patch.object(server.engine, "analyze_locked", side_effect=InputLimitError(message)):
131
+ response = self.client.post("/api/analyze", json=public_request())
132
+ self.assertEqual(response.status_code, 422)
133
+ self.assertEqual(response.json()["detail"], message)
134
+ self.assertFalse(server.engine.slot.locked())
135
+
136
+ def test_catalog_has_halu_pin_examples_and_demo_limits_without_loading(self):
137
+ catalog = self.client.get("/api/catalog").json()
138
+ spec = next(task for task in catalog["tasks"] if task["id"] == "hallucination")
139
+ self.assertEqual(spec["model"], "llm-semantic-router/Vela-1.0-Encoder-307M-Halu")
140
+ self.assertEqual(spec["revision"], "521cd05d15e1959e120d002663c3b775649ae4dd")
141
+ self.assertEqual(spec["kind"], "hallucination")
142
+ self.assertEqual(len(spec["examples"]), 5)
143
+ self.assertTrue(all({"name", "text", "context", "answer"} <= example.keys()
144
+ for example in spec["examples"]))
145
+ self.assertEqual(catalog["limits"]["max_tokens"], 512)
146
+ self.assertEqual(catalog["limits"]["max_hallucination_field_chars"], 16000)
147
+ self.assertIsNone(self.client.get("/api/status").json()["resident_task"])
148
+
149
+
150
+ class HallucinationInferenceTests(unittest.TestCase):
151
+ def test_512_token_pair_budget_preserves_exact_prompt_and_rejects_before_load(self):
152
+ observed = []
153
+
154
+ def tokenizer(text, text_pair=None, **kwargs):
155
+ observed.append((text, text_pair, kwargs))
156
+ return {"input_ids": [0] * (len(text.split()) + len(text_pair.split()) + 3)}
157
+
158
+ engine = InferenceEngine()
159
+ question, context, answer = " Q? ", " evidence " * 250, " answer " * 256
160
+ engine._check_tokens(tokenizer, "hallucination", question, [], context=context, answer=answer)
161
+ self.assertEqual(observed[-1][0], f"User request: {question}\n\n{context}")
162
+ self.assertEqual(observed[-1][1], answer)
163
+ self.assertIs(observed[-1][2]["truncation"], False)
164
+ with patch.object(engine, "_get_tokenizer", return_value=tokenizer), patch.object(engine, "_ensure_model") as load:
165
+ with self.assertRaisesRegex(InputLimitError, r"Question \+ evidence \+ answer has 513 tokens.*at most 512"):
166
+ engine.analyze_locked("hallucination", question, [], context=context, answer=answer + " extra")
167
+ load.assert_not_called()
168
+
169
+ def test_halu_loads_native_token_classification_cpu_contract_and_reuses_lru(self):
170
+ model = SimpleNamespace(eval=Mock())
171
+ token_loader = SimpleNamespace(from_pretrained=Mock(return_value=model))
172
+ wrong_loader = SimpleNamespace(from_pretrained=Mock(side_effect=AssertionError("wrong model class")))
173
+ transformers = SimpleNamespace(AutoModel=wrong_loader, AutoModelForSequenceClassification=wrong_loader,
174
+ AutoModelForTokenClassification=token_loader)
175
+ torch = SimpleNamespace(float32="float32", set_num_threads=lambda count: None)
176
+ engine, tokenizer = InferenceEngine(), object()
177
+ with patch.dict("sys.modules", {"torch": torch, "transformers": transformers}), \
178
+ patch("inference.validate_hallucination_model") as validate:
179
+ engine._ensure_model("hallucination", tokenizer)
180
+ engine._ensure_model("hallucination", tokenizer)
181
+ validate.assert_called_once_with(model)
182
+ token_loader.from_pretrained.assert_called_once_with(TASKS["hallucination"]["model"],
183
+ revision=TASKS["hallucination"]["revision"], torch_dtype="float32",
184
+ attn_implementation="sdpa", trust_remote_code=False, reference_compile=False)
185
+ self.assertEqual(engine.status()["cached_tasks"], ["hallucination"])
186
+ self.assertEqual(engine.model_cache_size, 2)
187
+
188
+ def test_wrong_labels_fail_before_a_model_can_enter_the_cache(self):
189
+ config = SimpleNamespace(model_type="modernbert", id2label={0: "hallucinated", 1: "supported"},
190
+ label2id={"supported": 1, "hallucinated": 0}, num_labels=2,
191
+ reference_compile=False, _attn_implementation="sdpa")
192
+ model = type("ModernBertForTokenClassification", (), {"config": config})()
193
+ with self.assertRaisesRegex(ValueError, "native inference contract"):
194
+ validate_hallucination_model(model)
195
+
196
+
197
+ class HallucinationCacheTests(unittest.TestCase):
198
+ def test_only_full_exact_public_question_evidence_answer_can_hit_cache(self):
199
+ spec = deepcopy(TASKS["hallucination"])
200
+ calls = []
201
+
202
+ def runner(payload, stage):
203
+ calls.append(deepcopy(payload))
204
+ return {"task": "hallucination", "model": spec["model"], "revision": spec["revision"],
205
+ "result": {"kind": "hallucination", "answer": payload["answer"], "spans": []}}
206
+
207
+ now = [100.0]
208
+ queue = InferenceQueue(runner, tasks={"hallucination": spec}, clock=lambda: now[0])
209
+ queue.start()
210
+ self.addCleanup(queue.stop)
211
+
212
+ def run(payload):
213
+ created = queue.submit(payload)
214
+ self.assertTrue(queue._jobs[created["id"]].done.wait(2))
215
+ return queue.get(created["id"])
216
+
217
+ self.assertFalse(run(public_request())["result"]["cache_hit"])
218
+ self.assertTrue(run(public_request())["result"]["cache_hit"])
219
+ # Same question/evidence with a different public answer is a separate key.
220
+ self.assertFalse(run(public_request(1))["result"]["cache_hit"])
221
+ self.assertTrue(run(public_request(1))["result"]["cache_hit"])
222
+ self.assertEqual(len(calls), 2)
223
+ for field in ("text", "context", "answer"):
224
+ custom = dict(public_request(), **{field: public_request()[field] + " private@example.com"})
225
+ for _ in range(2):
226
+ self.assertFalse(run(custom)["result"]["cache_hit"])
227
+ self.assertEqual(len(calls), 8)
228
+ self.assertEqual(queue.status()["example_cache_entries"], 2)
229
+ self.assertNotIn("private@example.com", str(queue._examples))
230
+ self.assertTrue(all(job.payload is None for job in queue._jobs.values()))
231
+ now[0] += 301
232
+ self.assertTrue(run(public_request())["result"]["cache_hit"])
233
+ spec["revision"] = "new-pinned-revision"
234
+ self.assertFalse(run(public_request())["result"]["cache_hit"])
235
+
236
+
237
+ if __name__ == "__main__":
238
+ unittest.main()