Spaces:
Running
Running
Add Hallucination playground with Vela Halu and verified examples
Browse files- README.md +19 -9
- app.py +13 -1
- catalog.py +10 -0
- hallucination.py +87 -0
- hallucination_catalog.py +37 -0
- inference.py +22 -5
- job_queue.py +6 -2
- smoke_test.py +12 -3
- static/app.js +35 -18
- static/flow.js +2 -2
- static/hallucination.css +75 -0
- static/hallucination.js +94 -0
- static/index.html +7 -5
- tests/test_hallucination.py +238 -0
README.md
CHANGED
|
@@ -8,7 +8,7 @@ app_port: 7860
|
|
| 8 |
header: mini
|
| 9 |
pinned: false
|
| 10 |
license: apache-2.0
|
| 11 |
-
short_description: Explore
|
| 12 |
models:
|
| 13 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Domain
|
| 14 |
- llm-semantic-router/Vela-1.0-Encoder-307M-PII
|
|
@@ -16,6 +16,7 @@ models:
|
|
| 16 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Safety
|
| 17 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Hazard
|
| 18 |
- llm-semantic-router/Vela-1.0-Encoder-307M-FactCheck
|
|
|
|
| 19 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Modality
|
| 20 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Feedback
|
| 21 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Embedding
|
|
@@ -24,7 +25,7 @@ models:
|
|
| 24 |
|
| 25 |
# Vela Studio
|
| 26 |
|
| 27 |
-
A public playground for the [Vela 1.0 model family](https://huggingface.co/collections/llm-semantic-router/vela-10), by [vLLM Semantic Router](https://vllm-sr.ai). Results come from the selected model's pinned checkpoint on CPU. Exact built-in examples can reuse a prior successful inference, clearly marked as a cached example result. The interface includes
|
| 28 |
|
| 29 |
## Runtime
|
| 30 |
|
|
@@ -32,19 +33,19 @@ The Docker image uses Python 3.11, CPU PyTorch 2.8.0, Transformers 4.57.6, and F
|
|
| 32 |
|
| 33 |
Two 307M FP32 checkpoints require approximately 2.46 GB of parameter memory, plus runtime and activation memory. Execution remains sequential, with a single input or query-passage pair at a time. Set `VELA_MODEL_CACHE_SIZE=1` to retain only one model on smaller hosts; supported values are `1` and `2`, with `2` as the default.
|
| 34 |
|
| 35 |
-
Most demos accept up to **512 tokens including special tokens**; Clustering accepts **3–24 texts**, each up to **128 tokens including special tokens** and 2,048 characters. Similarity and Reranker allow at most **six candidates**; Routing allows **2–10 categories**. Reranker applies the limit to each query-and-passage pair; Embedding applies it independently to each text. Inputs exceeding this limit are rejected without truncation. The model family has a larger declared input capacity; this shorter public limit keeps CPU Basic usage bounded.
|
| 36 |
|
| 37 |
One inference worker serves a FIFO queue with up to **eight waiting requests** in addition to the running request. The interface shows queue position, model loading, and inference progress and allows cancellation. HTTP 429 is returned only when the waiting queue is full, with `Retry-After: 5`. Cancelling pending work removes it from the queue immediately. Cancelling running work discards its eventual result; CPU execution finishes before the next request starts. A cold download can still take several minutes, and a long queue does not increase CPU throughput.
|
| 38 |
|
| 39 |
-
Only exact matches to the public examples in `catalog.py`, `routing_catalog.py`, `clustering_catalog.py`, and `
|
| 40 |
|
| 41 |
No user text, predictions, or exception payloads are logged or written to disk by this application. Pending input remains in memory until execution or cancellation; active input is released when execution returns. Private job results expire about five minutes after completion, failure, or cancellation, with cleanup at most 30 seconds later when idle. At most 64 job records are retained, so older completed results can expire sooner under load. Job IDs are unguessable bearer tokens; keep the ID private if the input is private. Hugging Face provides the surrounding hosting infrastructure. Use fictional data when trying PII examples.
|
| 42 |
|
| 43 |
## Interface
|
| 44 |
|
| 45 |
-
All
|
| 46 |
|
| 47 |
-
Results retain their task-specific meaning: classifications show label scores, PII offers complete highlighted and redacted text, Hazard preserves all twelve independently thresholded categories, and Similarity and Reranker show ordered passages with their respective cosine and raw relevance scores. Routing destinations remain editable in place. The global animation control pauses the ship, compass, and connecting paths.
|
| 48 |
|
| 49 |
Similarity offers three example groups in the same tab: **Paraphrases**, **Question answering**, and **Duplicate questions**, with two verified examples per group. Paraphrases retains the original delivery-status and password-reset inputs. Question answering finds password-reset and order-return instructions. Duplicate questions compares Git-commit and Python-list questions against related but different questions. Changing the group or choosing an example prepares the inputs; click **Compare meanings** to run the model. Every example’s intended candidate ranked first in real pinned-model inference before inclusion. These checks establish the displayed rankings, not general model accuracy or a duplicate threshold. The earlier library example and its cross-language/time-detail limitations remain in the separate acceptance records.
|
| 50 |
|
|
@@ -54,6 +55,12 @@ PII offers seven examples covering 13 entity types: contact details, date and ag
|
|
| 54 |
|
| 55 |
Each selected PII example was checked for complete entity spans and redaction; each Hazard example was checked against the full detected-label set using the unchanged published thresholds. Candidate IBAN and ZIP-code examples with incomplete spans or incorrect labels, and Hazard wordings with unrelated extra detections, were excluded; their raw results remain in the separate acceptance records. These curated cases demonstrate behavior on the displayed inputs, not general model accuracy.
|
| 56 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 57 |
## Routing
|
| 58 |
|
| 59 |
Routing opens by default. Its shared canvas connects the prompt, the Vela Encoder compass card, and compact editable destinations. Scenario selection, examples, and reset share a toolbar below the canvas.
|
|
@@ -78,16 +85,17 @@ Clustering has the same queue, cancellation, and model-cache behavior as the oth
|
|
| 78 |
|
| 79 |
## Model contracts
|
| 80 |
|
| 81 |
-
The September 14, 2026 semantic audit of the 20 built-in prompts found 19 aligned with their expected meaning and one miss: Guard's first, short instruction-override example returned `benign` (66.37%) instead of `jailbreak`. After that audit, the user requested a clearer positive Guard demo. Its new default query, a detailed developer-mode override, returned jailbreak 99.9923% on two identical real runs. Guard now offers two examples: the clear override and an ordinary explanatory request. At that point the ten model demos contained 20 examples, alongside 30 Routing examples. Subsequent curated PII and Hazard additions brought the total to 63 examples; three Clustering batches brought it to 66, and four additional Similarity examples bring the current total to
|
| 82 |
|
| 83 |
- Domain, FactCheck, Modality, Feedback, Guard, and Safety return softmax label scores.
|
| 84 |
- FactCheck identifies a need for external factual knowledge or retrieval; it does not decide whether a claim is true. Guard concerns instruction attacks, while Safety and Hazard concern content risk. A risk signal alone does not prescribe refusal.
|
|
|
|
| 85 |
- PII returns 17 entity types from 35 BIO labels. `start` and `end` are Unicode code-point offsets into the original text. Entity scores average their token probabilities.
|
| 86 |
- Hazard uses independent sigmoid scores and the saved per-category thresholds from its [pinned operating point](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Hazard/blob/188325c44aaf929eeea4e24463ed2543549339b2/operating_point.json). The demo's 512-token maximum is inside one 2048-token window. Full-length clients must implement the documented overlapping-window policy.
|
| 87 |
- Embedding uses its native ModernBERT checkpoint, attention-mask mean pooling in FP32, and L2 normalization, matching its SentenceTransformers configuration. Similarities are cosine scores between 768-dimensional vectors.
|
| 88 |
- Reranker returns raw relevance logits. Its [pinned custom loader](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Reranker/blob/a388e41cbbd5dc5f16b6389fa76d0b8b8a38a8bf/modeling_vela_reranker.py) was inspected before enabling `trust_remote_code`; it loads the native encoder and separate trained FP32 classification heads. It uses all 22 layers and 768 dimensions.
|
| 89 |
|
| 90 |
-
Model IDs, immutable revisions, examples, and Hazard thresholds live in `catalog.py`; Routing scenarios and their verified configurations live in `routing_catalog.py`; Clustering examples, per-item limits, and the shared default cutoff live in `clustering_catalog.py`. Similarity use-case groups and verified configurations live in `similarity_catalog.py`. Upgrading a model requires reviewing its contract and rerunning real inference checks.
|
| 91 |
|
| 92 |
## Local run
|
| 93 |
|
|
@@ -112,6 +120,8 @@ uvicorn app:app --host 0.0.0.0 --port 7860 --workers 1 --no-access-log
|
|
| 112 |
|
| 113 |
`POST /api/jobs` accepts `{"task":"domain","text":"Why do bond prices fall?","candidates":[]}` and returns HTTP 202 with `{id, status, position, queue_ms}`. Poll `GET /api/jobs/{id}` for status: `queued`, `loading`, `running`, `completed`, `failed`, or `cancelled`. Queued positions start at one; active and terminal jobs have position zero. A completed snapshot adds `result`, containing the analysis response. A failed snapshot adds a safe `error` string and `error_status` (422 for input limits, otherwise 503). Missing or expired IDs return HTTP 404. `DELETE /api/jobs/{id}` cancels pending or running work; terminal jobs retain their final status. A cached example can already be completed in the initial HTTP 202 response.
|
| 114 |
|
|
|
|
|
|
|
| 115 |
Routing jobs use `{"task":"routing","text":"Plan a weekend trip.","categories":[{"id":"travel","name":"Travel","description":"Trip planning and itineraries"},{"id":"coding","name":"Coding","description":"Programming and debugging"}]}`. Routing returns `kind: "routing"`, `top_category_id`, sorted `items` with original category IDs and indices, `dimensions: 768`, `score_type: "cosine"`, and the top-two `margin`. Similarity and Reranker retain their existing six-candidate limit.
|
| 116 |
|
| 117 |
Clustering jobs use `{"task":"clustering","texts":["I forgot my password.","How can I reset my password?","How do I grow tomatoes?"]}`. They return `kind: "clustering"`, indexed original `items`, the cosine `similarities` matrix, a full average-link `merges` hierarchy, the `default_threshold`, and `groups` containing member `indices` and the representative index. Tree leaves are the original indices; merge `i` creates node `n + i`. Groups are ordered by their first input index, with deterministic tie breaking. Nonempty query text, candidates, or categories are rejected for this task; `texts` are rejected for other tasks.
|
|
@@ -132,4 +142,4 @@ python smoke_test.py --task domain
|
|
| 132 |
python smoke_test.py --task all
|
| 133 |
```
|
| 134 |
|
| 135 |
-
The all-model smoke test invokes
|
|
|
|
| 8 |
header: mini
|
| 9 |
pinned: false
|
| 10 |
license: apache-2.0
|
| 11 |
+
short_description: Explore eleven Vela models for intelligent routing.
|
| 12 |
models:
|
| 13 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Domain
|
| 14 |
- llm-semantic-router/Vela-1.0-Encoder-307M-PII
|
|
|
|
| 16 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Safety
|
| 17 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Hazard
|
| 18 |
- llm-semantic-router/Vela-1.0-Encoder-307M-FactCheck
|
| 19 |
+
- llm-semantic-router/Vela-1.0-Encoder-307M-Halu
|
| 20 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Modality
|
| 21 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Feedback
|
| 22 |
- llm-semantic-router/Vela-1.0-Encoder-307M-Embedding
|
|
|
|
| 25 |
|
| 26 |
# Vela Studio
|
| 27 |
|
| 28 |
+
A public playground for the [Vela 1.0 model family](https://huggingface.co/collections/llm-semantic-router/vela-10), by [vLLM Semantic Router](https://vllm-sr.ai). Results come from the selected model's pinned checkpoint on CPU. Exact built-in examples can reuse a prior successful inference, clearly marked as a cached example result. The interface includes thirteen interactive demos powered by eleven model checkpoints: classification, evidence-based hallucination highlights, PII highlights and redaction, content-risk thresholds, semantic similarity, dynamic prompt routing, passage reranking, and semantic clustering.
|
| 29 |
|
| 30 |
## Runtime
|
| 31 |
|
|
|
|
| 33 |
|
| 34 |
Two 307M FP32 checkpoints require approximately 2.46 GB of parameter memory, plus runtime and activation memory. Execution remains sequential, with a single input or query-passage pair at a time. Set `VELA_MODEL_CACHE_SIZE=1` to retain only one model on smaller hosts; supported values are `1` and `2`, with `2` as the default.
|
| 35 |
|
| 36 |
+
Most demos accept up to **512 tokens including special tokens**; Clustering accepts **3–24 texts**, each up to **128 tokens including special tokens** and 2,048 characters. Similarity and Reranker allow at most **six candidates**; Routing allows **2–10 categories**. Hallucination applies the limit to the complete question, evidence, and answer pair; Reranker applies the limit to each query-and-passage pair; Embedding applies it independently to each text. Inputs exceeding this limit are rejected without truncation. The model family has a larger declared input capacity; this shorter public limit keeps CPU Basic usage bounded.
|
| 37 |
|
| 38 |
One inference worker serves a FIFO queue with up to **eight waiting requests** in addition to the running request. The interface shows queue position, model loading, and inference progress and allows cancellation. HTTP 429 is returned only when the waiting queue is full, with `Retry-After: 5`. Cancelling pending work removes it from the queue immediately. Cancelling running work discards its eventual result; CPU execution finishes before the next request starts. A cold download can still take several minutes, and a long queue does not increase CPU throughput.
|
| 39 |
|
| 40 |
+
Only exact matches to the public examples in `catalog.py`, `routing_catalog.py`, `clustering_catalog.py`, `similarity_catalog.py`, and `hallucination_catalog.py` can reuse successful inference results. The key includes task, text, evidence and answer when supplied, candidates, the complete ordered category configuration, or the complete ordered clustering texts, model ID, and pinned revision. Each cache hit gets a separate random job ID and `cache_hit: true`; original inference timings are preserved. These public results stay in an in-process cache capped at 80 entries until eviction or restart. Edited examples and other user inputs are never added to this cache.
|
| 41 |
|
| 42 |
No user text, predictions, or exception payloads are logged or written to disk by this application. Pending input remains in memory until execution or cancellation; active input is released when execution returns. Private job results expire about five minutes after completion, failure, or cancellation, with cleanup at most 30 seconds later when idle. At most 64 job records are retained, so older completed results can expire sooner under load. Job IDs are unguessable bearer tokens; keep the ID private if the input is private. Hugging Face provides the surrounding hosting infrastructure. Use fictional data when trying PII examples.
|
| 43 |
|
| 44 |
## Interface
|
| 45 |
|
| 46 |
+
All thirteen demos share an input → model → result canvas, with the same compass model card, connecting paths, run-button placement, and example toolbar. Routing, Similarity, Reranker, and Clustering appear first in the horizontal navigation. Desktop uses three columns; narrower screens adapt the same flow without hiding task controls. The compass shows the selected model's identity and actual queue, loading, execution, or cache status.
|
| 47 |
|
| 48 |
+
Results retain their task-specific meaning: classifications show label scores, Hallucination highlights unsupported answer spans, PII offers complete highlighted and redacted text, Hazard preserves all twelve independently thresholded categories, and Similarity and Reranker show ordered passages with their respective cosine and raw relevance scores. Routing destinations remain editable in place. The global animation control pauses the ship, compass, and connecting paths.
|
| 49 |
|
| 50 |
Similarity offers three example groups in the same tab: **Paraphrases**, **Question answering**, and **Duplicate questions**, with two verified examples per group. Paraphrases retains the original delivery-status and password-reset inputs. Question answering finds password-reset and order-return instructions. Duplicate questions compares Git-commit and Python-list questions against related but different questions. Changing the group or choosing an example prepares the inputs; click **Compare meanings** to run the model. Every example’s intended candidate ranked first in real pinned-model inference before inclusion. These checks establish the displayed rankings, not general model accuracy or a duplicate threshold. The earlier library example and its cross-language/time-detail limitations remain in the separate acceptance records.
|
| 51 |
|
|
|
|
| 55 |
|
| 56 |
Each selected PII example was checked for complete entity spans and redaction; each Hazard example was checked against the full detected-label set using the unchanged published thresholds. Candidate IBAN and ZIP-code examples with incomplete spans or incorrect labels, and Hazard wordings with unrelated extra detections, were excluded; their raw results remain in the separate acceptance records. These curated cases demonstrate behavior on the displayed inputs, not general model accuracy.
|
| 57 |
|
| 58 |
+
## Hallucination
|
| 59 |
+
|
| 60 |
+
Choose **Hallucination**, enter a question, the evidence to use, and an answer, then click **Check the answer**. Vela Halu highlights answer spans that are unsupported by the supplied evidence. A result with no highlighted spans means no unsupported spans were detected; it is not a guarantee of factual correctness. FactCheck is a separate task that identifies requests needing factual knowledge.
|
| 61 |
+
|
| 62 |
+
Five examples cover a wrong opening time, a supported answer, a changed quantity, an invented facility, and a tool-result mismatch. All five were checked with the published model in real CPU inference. Editing any of the three fields clears the previous result and cancels pending work. The combined input must fit 512 tokens including the prompt format and special tokens; each field accepts up to 16,000 characters. Overlong inputs are rejected without truncation.
|
| 63 |
+
|
| 64 |
## Routing
|
| 65 |
|
| 66 |
Routing opens by default. Its shared canvas connects the prompt, the Vela Encoder compass card, and compact editable destinations. Scenario selection, examples, and reset share a toolbar below the canvas.
|
|
|
|
| 85 |
|
| 86 |
## Model contracts
|
| 87 |
|
| 88 |
+
The September 14, 2026 semantic audit of the 20 built-in prompts found 19 aligned with their expected meaning and one miss: Guard's first, short instruction-override example returned `benign` (66.37%) instead of `jailbreak`. After that audit, the user requested a clearer positive Guard demo. Its new default query, a detailed developer-mode override, returned jailbreak 99.9923% on two identical real runs. Guard now offers two examples: the clear override and an ordinary explanatory request. At that point the ten model demos contained 20 examples, alongside 30 Routing examples. Subsequent curated PII and Hazard additions brought the total to 63 examples; three Clustering batches brought it to 66, and four additional Similarity examples brought the total to 70. Five Hallucination examples bring the current total to 75. The original missed prompt and its raw result remain in the acceptance records; replacing a demonstration does not resolve the model’s short-override limitation. These example checks do not estimate general model accuracy.
|
| 89 |
|
| 90 |
- Domain, FactCheck, Modality, Feedback, Guard, and Safety return softmax label scores.
|
| 91 |
- FactCheck identifies a need for external factual knowledge or retrieval; it does not decide whether a claim is true. Guard concerns instruction attacks, while Safety and Hazard concern content risk. A risk signal alone does not prescribe refusal.
|
| 92 |
+
- Hallucination pairs the question and supplied evidence with an answer and highlights answer tokens whose hallucination score is strictly above 0.5. It preserves the published model’s native YaRN configuration. Span offsets are Unicode code points into the original answer; each span score is the maximum of its detected token scores.
|
| 93 |
- PII returns 17 entity types from 35 BIO labels. `start` and `end` are Unicode code-point offsets into the original text. Entity scores average their token probabilities.
|
| 94 |
- Hazard uses independent sigmoid scores and the saved per-category thresholds from its [pinned operating point](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Hazard/blob/188325c44aaf929eeea4e24463ed2543549339b2/operating_point.json). The demo's 512-token maximum is inside one 2048-token window. Full-length clients must implement the documented overlapping-window policy.
|
| 95 |
- Embedding uses its native ModernBERT checkpoint, attention-mask mean pooling in FP32, and L2 normalization, matching its SentenceTransformers configuration. Similarities are cosine scores between 768-dimensional vectors.
|
| 96 |
- Reranker returns raw relevance logits. Its [pinned custom loader](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Reranker/blob/a388e41cbbd5dc5f16b6389fa76d0b8b8a38a8bf/modeling_vela_reranker.py) was inspected before enabling `trust_remote_code`; it loads the native encoder and separate trained FP32 classification heads. It uses all 22 layers and 768 dimensions.
|
| 97 |
|
| 98 |
+
Model IDs, immutable revisions, examples, and Hazard thresholds live in `catalog.py`; Routing scenarios and their verified configurations live in `routing_catalog.py`; Clustering examples, per-item limits, and the shared default cutoff live in `clustering_catalog.py`. Similarity use-case groups and verified configurations live in `similarity_catalog.py`. Hallucination examples and the fixed threshold live in `hallucination_catalog.py`. Upgrading a model requires reviewing its contract and rerunning real inference checks.
|
| 99 |
|
| 100 |
## Local run
|
| 101 |
|
|
|
|
| 120 |
|
| 121 |
`POST /api/jobs` accepts `{"task":"domain","text":"Why do bond prices fall?","candidates":[]}` and returns HTTP 202 with `{id, status, position, queue_ms}`. Poll `GET /api/jobs/{id}` for status: `queued`, `loading`, `running`, `completed`, `failed`, or `cancelled`. Queued positions start at one; active and terminal jobs have position zero. A completed snapshot adds `result`, containing the analysis response. A failed snapshot adds a safe `error` string and `error_status` (422 for input limits, otherwise 503). Missing or expired IDs return HTTP 404. `DELETE /api/jobs/{id}` cancels pending or running work; terminal jobs retain their final status. A cached example can already be completed in the initial HTTP 202 response.
|
| 122 |
|
| 123 |
+
Hallucination jobs use `{"task":"hallucination","text":"When does the museum open?","context":"The museum opens at 10:00.","answer":"The museum opens at 09:00."}`. All three text fields are required and preserved verbatim. The result has `kind: "hallucination"`, the original `answer`, `spans` with `start`, `end`, `text`, and `score`, plus `max_hallucination_score` across answer tokens. Nonempty `context` or `answer` fields are rejected for other tasks. Only exact matches of all three fields can reuse a public example result.
|
| 124 |
+
|
| 125 |
Routing jobs use `{"task":"routing","text":"Plan a weekend trip.","categories":[{"id":"travel","name":"Travel","description":"Trip planning and itineraries"},{"id":"coding","name":"Coding","description":"Programming and debugging"}]}`. Routing returns `kind: "routing"`, `top_category_id`, sorted `items` with original category IDs and indices, `dimensions: 768`, `score_type: "cosine"`, and the top-two `margin`. Similarity and Reranker retain their existing six-candidate limit.
|
| 126 |
|
| 127 |
Clustering jobs use `{"task":"clustering","texts":["I forgot my password.","How can I reset my password?","How do I grow tomatoes?"]}`. They return `kind: "clustering"`, indexed original `items`, the cosine `similarities` matrix, a full average-link `merges` hierarchy, the `default_threshold`, and `groups` containing member `indices` and the representative index. Tree leaves are the original indices; merge `i` creates node `n + i`. Groups are ordered by their first input index, with deterministic tie breaking. Nonempty query text, candidates, or categories are rejected for this task; `texts` are rejected for other tasks.
|
|
|
|
| 142 |
python smoke_test.py --task all
|
| 143 |
```
|
| 144 |
|
| 145 |
+
The all-model smoke test invokes eleven pinned checkpoints when their example results are not cached and may download about 13.5 GB. It runs requests sequentially. Unit tests and output-shape checks alone do not establish prediction quality.
|
app.py
CHANGED
|
@@ -14,6 +14,7 @@ from starlette.concurrency import run_in_threadpool
|
|
| 14 |
|
| 15 |
from catalog import MAX_CANDIDATES, MAX_CLUSTER_ITEMS, MAX_ROUTING_CATEGORIES, TASKS, public_catalog
|
| 16 |
from inference import InferenceEngine, InputLimitError
|
|
|
|
| 17 |
from job_queue import InferenceQueue, JobFailure, QueueFull
|
| 18 |
from request_limits import RequestSizeLimitMiddleware
|
| 19 |
|
|
@@ -26,7 +27,8 @@ def run_inference(payload, set_stage):
|
|
| 26 |
with engine.slot:
|
| 27 |
return engine.analyze_locked(payload["task"], payload["text"], payload["candidates"],
|
| 28 |
on_stage=set_stage, categories=payload.get("categories", []),
|
| 29 |
-
texts=payload.get("texts", [])
|
|
|
|
| 30 |
except InputLimitError as exc:
|
| 31 |
raise JobFailure(str(exc), 422) from None
|
| 32 |
except Exception as exc:
|
|
@@ -79,6 +81,8 @@ class AnalyzeRequest(BaseModel):
|
|
| 79 |
candidates: list[Candidate] = Field(default_factory=list, max_length=MAX_CANDIDATES)
|
| 80 |
categories: list[RoutingCategory] = Field(default_factory=list, max_length=MAX_ROUTING_CATEGORIES)
|
| 81 |
texts: list[ClusterText] = Field(default_factory=list, max_length=MAX_CLUSTER_ITEMS)
|
|
|
|
|
|
|
| 82 |
|
| 83 |
@field_validator("task")
|
| 84 |
@classmethod
|
|
@@ -89,6 +93,14 @@ class AnalyzeRequest(BaseModel):
|
|
| 89 |
|
| 90 |
@model_validator(mode="after")
|
| 91 |
def valid_content(self):
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 92 |
if self.task == "clustering":
|
| 93 |
if self.text or self.candidates or self.categories:
|
| 94 |
raise ValueError("Clustering accepts only the texts list.")
|
|
|
|
| 14 |
|
| 15 |
from catalog import MAX_CANDIDATES, MAX_CLUSTER_ITEMS, MAX_ROUTING_CATEGORIES, TASKS, public_catalog
|
| 16 |
from inference import InferenceEngine, InputLimitError
|
| 17 |
+
from hallucination_catalog import MAX_HALLUCINATION_FIELD_CHARS
|
| 18 |
from job_queue import InferenceQueue, JobFailure, QueueFull
|
| 19 |
from request_limits import RequestSizeLimitMiddleware
|
| 20 |
|
|
|
|
| 27 |
with engine.slot:
|
| 28 |
return engine.analyze_locked(payload["task"], payload["text"], payload["candidates"],
|
| 29 |
on_stage=set_stage, categories=payload.get("categories", []),
|
| 30 |
+
texts=payload.get("texts", []),
|
| 31 |
+
context=payload.get("context", ""), answer=payload.get("answer", ""))
|
| 32 |
except InputLimitError as exc:
|
| 33 |
raise JobFailure(str(exc), 422) from None
|
| 34 |
except Exception as exc:
|
|
|
|
| 81 |
candidates: list[Candidate] = Field(default_factory=list, max_length=MAX_CANDIDATES)
|
| 82 |
categories: list[RoutingCategory] = Field(default_factory=list, max_length=MAX_ROUTING_CATEGORIES)
|
| 83 |
texts: list[ClusterText] = Field(default_factory=list, max_length=MAX_CLUSTER_ITEMS)
|
| 84 |
+
context: str = Field(default="", max_length=MAX_HALLUCINATION_FIELD_CHARS)
|
| 85 |
+
answer: str = Field(default="", max_length=MAX_HALLUCINATION_FIELD_CHARS)
|
| 86 |
|
| 87 |
@field_validator("task")
|
| 88 |
@classmethod
|
|
|
|
| 93 |
|
| 94 |
@model_validator(mode="after")
|
| 95 |
def valid_content(self):
|
| 96 |
+
if self.task == "hallucination":
|
| 97 |
+
if self.candidates or self.categories or self.texts:
|
| 98 |
+
raise ValueError("Hallucination accepts only a question, evidence, and answer.")
|
| 99 |
+
if not all(value.strip() for value in (self.text, self.context, self.answer)):
|
| 100 |
+
raise ValueError("Question, evidence, and answer cannot be blank.")
|
| 101 |
+
return self
|
| 102 |
+
if self.context or self.answer:
|
| 103 |
+
raise ValueError("Evidence and answer are only supported by Hallucination.")
|
| 104 |
if self.task == "clustering":
|
| 105 |
if self.text or self.candidates or self.categories:
|
| 106 |
raise ValueError("Clustering accepts only the texts list.")
|
catalog.py
CHANGED
|
@@ -8,6 +8,7 @@ from routing_catalog import MAX_ROUTING_CATEGORIES, ROUTING_SCENARIOS
|
|
| 8 |
from similarity_catalog import SIMILARITY_EXAMPLES, SIMILARITY_EXAMPLE_GROUPS
|
| 9 |
from clustering_catalog import (CLUSTERING_EXAMPLES, CLUSTER_DISTANCE_THRESHOLD,
|
| 10 |
MAX_CLUSTER_ITEMS, MAX_CLUSTER_TOKENS)
|
|
|
|
| 11 |
|
| 12 |
_TASKS = [
|
| 13 |
("domain", "Domain", "Find the subject of a request across 14 domains.", "classification", "f6354f54adcf38770f635ad903be2b00577f6c11", [
|
|
@@ -73,6 +74,14 @@ TASKS = {
|
|
| 73 |
|
| 74 |
TASKS["embedding"]["example_groups"] = SIMILARITY_EXAMPLE_GROUPS
|
| 75 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 76 |
# Routing is another demo of the same Embedding checkpoint, not a new model.
|
| 77 |
TASKS["routing"] = {
|
| 78 |
"id": "routing", "name": "Routing", "description": "Route prompts to categories you define.",
|
|
@@ -117,4 +126,5 @@ def public_catalog():
|
|
| 117 |
"max_tokens": MAX_TOKENS, "max_candidates": MAX_CANDIDATES,
|
| 118 |
"max_routing_categories": MAX_ROUTING_CATEGORIES,
|
| 119 |
"max_cluster_items": MAX_CLUSTER_ITEMS, "max_cluster_tokens": MAX_CLUSTER_TOKENS,
|
|
|
|
| 120 |
}}
|
|
|
|
| 8 |
from similarity_catalog import SIMILARITY_EXAMPLES, SIMILARITY_EXAMPLE_GROUPS
|
| 9 |
from clustering_catalog import (CLUSTERING_EXAMPLES, CLUSTER_DISTANCE_THRESHOLD,
|
| 10 |
MAX_CLUSTER_ITEMS, MAX_CLUSTER_TOKENS)
|
| 11 |
+
from hallucination_catalog import HALLUCINATION_EXAMPLES, MAX_HALLUCINATION_FIELD_CHARS
|
| 12 |
|
| 13 |
_TASKS = [
|
| 14 |
("domain", "Domain", "Find the subject of a request across 14 domains.", "classification", "f6354f54adcf38770f635ad903be2b00577f6c11", [
|
|
|
|
| 74 |
|
| 75 |
TASKS["embedding"]["example_groups"] = SIMILARITY_EXAMPLE_GROUPS
|
| 76 |
|
| 77 |
+
TASKS["hallucination"] = {
|
| 78 |
+
"id": "hallucination", "name": "Hallucination",
|
| 79 |
+
"description": "Find answer spans unsupported by the supplied evidence.",
|
| 80 |
+
"kind": "hallucination", "model": MODEL_PREFIX + "Halu",
|
| 81 |
+
"revision": "521cd05d15e1959e120d002663c3b775649ae4dd",
|
| 82 |
+
"examples": HALLUCINATION_EXAMPLES,
|
| 83 |
+
}
|
| 84 |
+
|
| 85 |
# Routing is another demo of the same Embedding checkpoint, not a new model.
|
| 86 |
TASKS["routing"] = {
|
| 87 |
"id": "routing", "name": "Routing", "description": "Route prompts to categories you define.",
|
|
|
|
| 126 |
"max_tokens": MAX_TOKENS, "max_candidates": MAX_CANDIDATES,
|
| 127 |
"max_routing_categories": MAX_ROUTING_CATEGORIES,
|
| 128 |
"max_cluster_items": MAX_CLUSTER_ITEMS, "max_cluster_tokens": MAX_CLUSTER_TOKENS,
|
| 129 |
+
"max_hallucination_field_chars": MAX_HALLUCINATION_FIELD_CHARS,
|
| 130 |
}}
|
hallucination.py
ADDED
|
@@ -0,0 +1,87 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Native Vela Halu input and answer-span contracts."""
|
| 2 |
+
|
| 3 |
+
import math
|
| 4 |
+
|
| 5 |
+
from hallucination_catalog import HALLUCINATION_THRESHOLD
|
| 6 |
+
|
| 7 |
+
|
| 8 |
+
def evidence_prompt(question, context):
|
| 9 |
+
return f"User request: {question}\n\n{context}"
|
| 10 |
+
|
| 11 |
+
|
| 12 |
+
def validate_hallucination_model(model):
|
| 13 |
+
config = model.config
|
| 14 |
+
if (type(model).__name__ != "ModernBertForTokenClassification"
|
| 15 |
+
or config.model_type != "modernbert"
|
| 16 |
+
or config.id2label != {0: "supported", 1: "hallucinated"}
|
| 17 |
+
or config.label2id != {"supported": 0, "hallucinated": 1}
|
| 18 |
+
or config.num_labels != 2
|
| 19 |
+
or config.reference_compile is not False
|
| 20 |
+
or config._attn_implementation != "sdpa"):
|
| 21 |
+
raise ValueError("Halu checkpoint differs from its native inference contract")
|
| 22 |
+
if (getattr(config, "rope_scaling", None) or {}).get("rope_type") != "yarn":
|
| 23 |
+
raise ValueError("Halu requires the published YaRN configuration")
|
| 24 |
+
count = 0
|
| 25 |
+
for _, module in model.named_modules():
|
| 26 |
+
if not hasattr(module, "rotary_emb"):
|
| 27 |
+
continue
|
| 28 |
+
rotary = module.rotary_emb
|
| 29 |
+
if (getattr(rotary, "rope_type", None) != "yarn"
|
| 30 |
+
or getattr(getattr(rotary, "rope_init_fn", None), "__name__", None)
|
| 31 |
+
!= "_compute_yarn_parameters"):
|
| 32 |
+
raise ValueError("Halu did not instantiate its published YaRN rotary")
|
| 33 |
+
count += 1
|
| 34 |
+
if count != config.num_hidden_layers or count != 22:
|
| 35 |
+
raise ValueError("Halu rotary layer count differs from its foundation")
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def hallucination_result(answer, sequence_ids, offsets, scores):
|
| 39 |
+
"""Use answer-only Unicode offsets and the published strict > 0.5 rule."""
|
| 40 |
+
if not len(sequence_ids) == len(offsets) == len(scores):
|
| 41 |
+
raise ValueError("Halu token scores and offsets differ in length")
|
| 42 |
+
spans, current = [], None
|
| 43 |
+
max_score = 0.0
|
| 44 |
+
previous_start = -1
|
| 45 |
+
answer_tokens = 0
|
| 46 |
+
|
| 47 |
+
def finish():
|
| 48 |
+
nonlocal current
|
| 49 |
+
if current is not None:
|
| 50 |
+
# Byte-fallback tokens can share a character while straddling the
|
| 51 |
+
# threshold. Present their positive character union once; ordinary
|
| 52 |
+
# nonoverlapping spans separated by a supported token stay distinct.
|
| 53 |
+
if spans and current["start"] < spans[-1]["end"]:
|
| 54 |
+
span = spans[-1]
|
| 55 |
+
span["end"] = max(span["end"], current["end"])
|
| 56 |
+
span["score"] = max(span["score"], current["score"])
|
| 57 |
+
else:
|
| 58 |
+
span = current
|
| 59 |
+
spans.append(span)
|
| 60 |
+
span["text"] = answer[span["start"]:span["end"]]
|
| 61 |
+
current = None
|
| 62 |
+
|
| 63 |
+
for sequence, (start, end), score in zip(sequence_ids, offsets, scores, strict=True):
|
| 64 |
+
score = float(score)
|
| 65 |
+
if not math.isfinite(score) or not 0 <= score <= 1:
|
| 66 |
+
raise ValueError("Invalid Halu token probability")
|
| 67 |
+
if sequence != 1 or end <= start:
|
| 68 |
+
continue
|
| 69 |
+
if (type(start) is not int or type(end) is not int
|
| 70 |
+
or not 0 <= start < end <= len(answer) or start < previous_start):
|
| 71 |
+
raise ValueError("Invalid Halu answer offsets")
|
| 72 |
+
previous_start = start
|
| 73 |
+
answer_tokens += 1
|
| 74 |
+
max_score = max(max_score, score)
|
| 75 |
+
if score > HALLUCINATION_THRESHOLD:
|
| 76 |
+
if current is None:
|
| 77 |
+
current = {"start": start, "end": end, "score": score}
|
| 78 |
+
else:
|
| 79 |
+
current["end"] = max(current["end"], end)
|
| 80 |
+
current["score"] = max(current["score"], score)
|
| 81 |
+
else:
|
| 82 |
+
finish()
|
| 83 |
+
finish()
|
| 84 |
+
if not answer_tokens:
|
| 85 |
+
raise ValueError("Halu found no evaluable answer tokens")
|
| 86 |
+
return {"kind": "hallucination", "answer": answer, "spans": spans,
|
| 87 |
+
"max_hallucination_score": max_score}
|
hallucination_catalog.py
ADDED
|
@@ -0,0 +1,37 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Public evidence-and-answer examples for Vela Halu."""
|
| 2 |
+
|
| 3 |
+
MAX_HALLUCINATION_FIELD_CHARS = 16000
|
| 4 |
+
HALLUCINATION_THRESHOLD = 0.5
|
| 5 |
+
|
| 6 |
+
HALLUCINATION_EXAMPLES = [
|
| 7 |
+
{
|
| 8 |
+
"name": "Wrong opening time",
|
| 9 |
+
"text": "When does the museum open on Tuesday?",
|
| 10 |
+
"context": "The museum opens at 10:00 on Tuesday.",
|
| 11 |
+
"answer": "The museum opens at 09:00 on Tuesday.",
|
| 12 |
+
},
|
| 13 |
+
{
|
| 14 |
+
"name": "Supported answer",
|
| 15 |
+
"text": "When does the museum open on Tuesday?",
|
| 16 |
+
"context": "The museum opens at 10:00 on Tuesday.",
|
| 17 |
+
"answer": "The museum opens at 10:00 on Tuesday.",
|
| 18 |
+
},
|
| 19 |
+
{
|
| 20 |
+
"name": "Changed number",
|
| 21 |
+
"text": "How many seats does the shuttle have?",
|
| 22 |
+
"context": "The shuttle has 24 seats and runs every 30 minutes.",
|
| 23 |
+
"answer": "The shuttle has 42 seats and runs every 30 minutes.",
|
| 24 |
+
},
|
| 25 |
+
{
|
| 26 |
+
"name": "Invented detail",
|
| 27 |
+
"text": "What facilities does the library offer?",
|
| 28 |
+
"context": "The library offers free Wi-Fi and a reading room.",
|
| 29 |
+
"answer": "The library offers free Wi-Fi, a reading room, and a rooftop café.",
|
| 30 |
+
},
|
| 31 |
+
{
|
| 32 |
+
"name": "Tool output mismatch",
|
| 33 |
+
"text": "Summarize the update tool's result.",
|
| 34 |
+
"context": 'The update tool returned: {"status": "success", "files_updated": 3, "tests_passed": 12}.',
|
| 35 |
+
"answer": "The tool updated 7 files and all 12 tests passed.",
|
| 36 |
+
},
|
| 37 |
+
]
|
inference.py
CHANGED
|
@@ -10,6 +10,7 @@ from collections import OrderedDict
|
|
| 10 |
from catalog import (HAZARD_THRESHOLDS, MAX_CANDIDATES, MAX_CLUSTER_ITEMS, MAX_CLUSTER_TOKENS,
|
| 11 |
MAX_ROUTING_CATEGORIES, MAX_TOKENS, TASKS)
|
| 12 |
from clustering import clustering_result
|
|
|
|
| 13 |
|
| 14 |
|
| 15 |
class InputLimitError(ValueError):
|
|
@@ -154,10 +155,12 @@ class InferenceEngine:
|
|
| 154 |
return AutoTokenizer.from_pretrained(spec["model"], revision=spec["revision"],
|
| 155 |
trust_remote_code=False, use_fast=True)
|
| 156 |
|
| 157 |
-
def _check_tokens(self, tokenizer, task, text, candidates, texts=None):
|
| 158 |
limit = MAX_CLUSTER_TOKENS if task == "clustering" else MAX_TOKENS
|
| 159 |
if task == "clustering":
|
| 160 |
inputs = [(item, None) for item in texts or []]
|
|
|
|
|
|
|
| 161 |
elif task == "reranker":
|
| 162 |
inputs = [(text, candidate) for candidate in candidates]
|
| 163 |
elif task in {"embedding", "routing"}:
|
|
@@ -167,7 +170,8 @@ class InferenceEngine:
|
|
| 167 |
for index, (first, second) in enumerate(inputs):
|
| 168 |
count = len(tokenizer(first, text_pair=second, truncation=False)["input_ids"])
|
| 169 |
if count > limit:
|
| 170 |
-
subject = (
|
|
|
|
| 171 |
f"Query + passage {index + 1}" if task == "reranker"
|
| 172 |
else "Input" if index == 0 else
|
| 173 |
f"Category {index}" if task == "routing" else f"Candidate {index}")
|
|
@@ -202,10 +206,12 @@ class InferenceEngine:
|
|
| 202 |
code_revision=spec["revision"], **arguments)
|
| 203 |
else:
|
| 204 |
model_class = (AutoModel if task == "embedding" else
|
| 205 |
-
AutoModelForTokenClassification if task
|
| 206 |
AutoModelForSequenceClassification)
|
| 207 |
model = model_class.from_pretrained(spec["model"], trust_remote_code=False,
|
| 208 |
reference_compile=False, **arguments)
|
|
|
|
|
|
|
| 209 |
model.eval()
|
| 210 |
with self._cache_lock:
|
| 211 |
self._cache[task] = (model, tokenizer)
|
|
@@ -252,7 +258,8 @@ class InferenceEngine:
|
|
| 252 |
vectors.append(vector)
|
| 253 |
return vectors, hits
|
| 254 |
|
| 255 |
-
def analyze_locked(self, task, text, candidates, on_stage=None, categories=None, texts=None
|
|
|
|
| 256 |
"""The caller owns slot for the entire load and inference lifecycle."""
|
| 257 |
started = time.perf_counter()
|
| 258 |
model_task = self._model_task(task)
|
|
@@ -262,7 +269,8 @@ class InferenceEngine:
|
|
| 262 |
if on_stage:
|
| 263 |
on_stage("loading" if cold else "running")
|
| 264 |
tokenizer = self._get_tokenizer(task)
|
| 265 |
-
self._check_tokens(tokenizer, task, text, representations, texts=texts
|
|
|
|
| 266 |
self._ensure_model(task, tokenizer)
|
| 267 |
load_ms = (time.perf_counter() - started) * 1000 if cold else 0.0
|
| 268 |
if on_stage:
|
|
@@ -306,6 +314,15 @@ class InferenceEngine:
|
|
| 306 |
tokens = [(model.config.id2label[index], score, start, end)
|
| 307 |
for index, score, (start, end) in zip(indices.tolist(), scores.tolist(), offsets)]
|
| 308 |
result = pii_result(text, tokens)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 309 |
else:
|
| 310 |
inputs = self._inputs(tokenizer, text)
|
| 311 |
logits = model(**inputs).logits[0].float().tolist()
|
|
|
|
| 10 |
from catalog import (HAZARD_THRESHOLDS, MAX_CANDIDATES, MAX_CLUSTER_ITEMS, MAX_CLUSTER_TOKENS,
|
| 11 |
MAX_ROUTING_CATEGORIES, MAX_TOKENS, TASKS)
|
| 12 |
from clustering import clustering_result
|
| 13 |
+
from hallucination import evidence_prompt, hallucination_result, validate_hallucination_model
|
| 14 |
|
| 15 |
|
| 16 |
class InputLimitError(ValueError):
|
|
|
|
| 155 |
return AutoTokenizer.from_pretrained(spec["model"], revision=spec["revision"],
|
| 156 |
trust_remote_code=False, use_fast=True)
|
| 157 |
|
| 158 |
+
def _check_tokens(self, tokenizer, task, text, candidates, texts=None, context="", answer=""):
|
| 159 |
limit = MAX_CLUSTER_TOKENS if task == "clustering" else MAX_TOKENS
|
| 160 |
if task == "clustering":
|
| 161 |
inputs = [(item, None) for item in texts or []]
|
| 162 |
+
elif task == "hallucination":
|
| 163 |
+
inputs = [(evidence_prompt(text, context), answer)]
|
| 164 |
elif task == "reranker":
|
| 165 |
inputs = [(text, candidate) for candidate in candidates]
|
| 166 |
elif task in {"embedding", "routing"}:
|
|
|
|
| 170 |
for index, (first, second) in enumerate(inputs):
|
| 171 |
count = len(tokenizer(first, text_pair=second, truncation=False)["input_ids"])
|
| 172 |
if count > limit:
|
| 173 |
+
subject = ("Question + evidence + answer" if task == "hallucination" else
|
| 174 |
+
f"Text {index + 1}" if task == "clustering" else
|
| 175 |
f"Query + passage {index + 1}" if task == "reranker"
|
| 176 |
else "Input" if index == 0 else
|
| 177 |
f"Category {index}" if task == "routing" else f"Candidate {index}")
|
|
|
|
| 206 |
code_revision=spec["revision"], **arguments)
|
| 207 |
else:
|
| 208 |
model_class = (AutoModel if task == "embedding" else
|
| 209 |
+
AutoModelForTokenClassification if task in {"pii", "hallucination"} else
|
| 210 |
AutoModelForSequenceClassification)
|
| 211 |
model = model_class.from_pretrained(spec["model"], trust_remote_code=False,
|
| 212 |
reference_compile=False, **arguments)
|
| 213 |
+
if task == "hallucination":
|
| 214 |
+
validate_hallucination_model(model)
|
| 215 |
model.eval()
|
| 216 |
with self._cache_lock:
|
| 217 |
self._cache[task] = (model, tokenizer)
|
|
|
|
| 258 |
vectors.append(vector)
|
| 259 |
return vectors, hits
|
| 260 |
|
| 261 |
+
def analyze_locked(self, task, text, candidates, on_stage=None, categories=None, texts=None,
|
| 262 |
+
context="", answer=""):
|
| 263 |
"""The caller owns slot for the entire load and inference lifecycle."""
|
| 264 |
started = time.perf_counter()
|
| 265 |
model_task = self._model_task(task)
|
|
|
|
| 269 |
if on_stage:
|
| 270 |
on_stage("loading" if cold else "running")
|
| 271 |
tokenizer = self._get_tokenizer(task)
|
| 272 |
+
self._check_tokens(tokenizer, task, text, representations, texts=texts,
|
| 273 |
+
context=context, answer=answer)
|
| 274 |
self._ensure_model(task, tokenizer)
|
| 275 |
load_ms = (time.perf_counter() - started) * 1000 if cold else 0.0
|
| 276 |
if on_stage:
|
|
|
|
| 314 |
tokens = [(model.config.id2label[index], score, start, end)
|
| 315 |
for index, score, (start, end) in zip(indices.tolist(), scores.tolist(), offsets)]
|
| 316 |
result = pii_result(text, tokens)
|
| 317 |
+
elif task == "hallucination":
|
| 318 |
+
inputs = self._inputs(tokenizer, evidence_prompt(text, context), answer, offsets=True)
|
| 319 |
+
sequence_ids = inputs.sequence_ids(0)
|
| 320 |
+
offsets = inputs.pop("offset_mapping")[0].tolist()
|
| 321 |
+
logits = model(**inputs).logits[0]
|
| 322 |
+
if logits.shape[-1] != 2:
|
| 323 |
+
raise ValueError("Halu must return exactly two token labels")
|
| 324 |
+
scores = logits.float().softmax(dim=-1)[:, 1].tolist()
|
| 325 |
+
result = hallucination_result(answer, sequence_ids, offsets, scores)
|
| 326 |
else:
|
| 327 |
inputs = self._inputs(tokenizer, text)
|
| 328 |
logits = model(**inputs).logits[0].float().tolist()
|
job_queue.py
CHANGED
|
@@ -91,12 +91,16 @@ class InferenceQueue:
|
|
| 91 |
candidates = tuple(payload.get("candidates", []))
|
| 92 |
categories = self._category_key(payload.get("categories", []))
|
| 93 |
texts = tuple(payload.get("texts", []))
|
|
|
|
| 94 |
for example in spec.get("examples", []):
|
| 95 |
if (payload.get("text", "") == example.get("text", "")
|
| 96 |
and candidates == tuple(example.get("candidates", []))
|
| 97 |
and categories == self._category_key(example.get("categories", []))
|
| 98 |
-
and texts == tuple(example.get("texts", []))
|
| 99 |
-
|
|
|
|
|
|
|
|
|
|
| 100 |
return None
|
| 101 |
|
| 102 |
@staticmethod
|
|
|
|
| 91 |
candidates = tuple(payload.get("candidates", []))
|
| 92 |
categories = self._category_key(payload.get("categories", []))
|
| 93 |
texts = tuple(payload.get("texts", []))
|
| 94 |
+
context, answer = payload.get("context", ""), payload.get("answer", "")
|
| 95 |
for example in spec.get("examples", []):
|
| 96 |
if (payload.get("text", "") == example.get("text", "")
|
| 97 |
and candidates == tuple(example.get("candidates", []))
|
| 98 |
and categories == self._category_key(example.get("categories", []))
|
| 99 |
+
and texts == tuple(example.get("texts", []))
|
| 100 |
+
and context == example.get("context", "")
|
| 101 |
+
and answer == example.get("answer", "")):
|
| 102 |
+
return (payload["task"], spec["model"], spec["revision"], payload.get("text", ""),
|
| 103 |
+
candidates, categories, texts, context, answer)
|
| 104 |
return None
|
| 105 |
|
| 106 |
@staticmethod
|
smoke_test.py
CHANGED
|
@@ -2,7 +2,7 @@
|
|
| 2 |
|
| 3 |
Usage: python smoke_test.py --base-url http://localhost:7860 --task all
|
| 4 |
Uncached examples execute the pinned model; cached responses are reported as
|
| 5 |
-
such. This can download ~
|
| 6 |
contracts, not semantic correctness; see the separately recorded example audit.
|
| 7 |
"""
|
| 8 |
|
|
@@ -27,7 +27,7 @@ def request(base, path, payload=None):
|
|
| 27 |
def example_payload(task, example):
|
| 28 |
# Catalog examples can also include display names and other metadata.
|
| 29 |
return {"task": task["id"], **{key: value for key, value in example.items()
|
| 30 |
-
if key in {"text", "candidates", "categories", "texts"}}}
|
| 31 |
|
| 32 |
|
| 33 |
def verify_clustering(result, task, example):
|
|
@@ -88,6 +88,15 @@ def verify(response, task, example):
|
|
| 88 |
assert result["entities"], "The public example should contain detected PII"
|
| 89 |
for entity in result["entities"]:
|
| 90 |
assert example["text"][entity["start"]:entity["end"]] == entity["text"]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 91 |
elif result["kind"] == "routing":
|
| 92 |
categories = example["categories"]
|
| 93 |
assert len(result["items"]) == len(categories)
|
|
@@ -121,7 +130,7 @@ def main():
|
|
| 121 |
parser.error("Unknown task")
|
| 122 |
example_count = 0
|
| 123 |
for task in selected:
|
| 124 |
-
examples = task["examples"] if task["kind"]
|
| 125 |
for example in examples:
|
| 126 |
payload = example_payload(task, example)
|
| 127 |
response = request(args.base_url, "/api/analyze", payload)
|
|
|
|
| 2 |
|
| 3 |
Usage: python smoke_test.py --base-url http://localhost:7860 --task all
|
| 4 |
Uncached examples execute the pinned model; cached responses are reported as
|
| 5 |
+
such. This can download ~13.5 GB of weights on its first run. This checks output
|
| 6 |
contracts, not semantic correctness; see the separately recorded example audit.
|
| 7 |
"""
|
| 8 |
|
|
|
|
| 27 |
def example_payload(task, example):
|
| 28 |
# Catalog examples can also include display names and other metadata.
|
| 29 |
return {"task": task["id"], **{key: value for key, value in example.items()
|
| 30 |
+
if key in {"text", "context", "answer", "candidates", "categories", "texts"}}}
|
| 31 |
|
| 32 |
|
| 33 |
def verify_clustering(result, task, example):
|
|
|
|
| 88 |
assert result["entities"], "The public example should contain detected PII"
|
| 89 |
for entity in result["entities"]:
|
| 90 |
assert example["text"][entity["start"]:entity["end"]] == entity["text"]
|
| 91 |
+
elif result["kind"] == "hallucination":
|
| 92 |
+
assert result["answer"] == example["answer"]
|
| 93 |
+
assert 0 <= result["max_hallucination_score"] <= 1
|
| 94 |
+
previous_end = 0
|
| 95 |
+
for span in result["spans"]:
|
| 96 |
+
assert previous_end <= span["start"] < span["end"] <= len(result["answer"])
|
| 97 |
+
assert result["answer"][span["start"]:span["end"]] == span["text"]
|
| 98 |
+
assert 0.5 < span["score"] <= 1
|
| 99 |
+
previous_end = span["end"]
|
| 100 |
elif result["kind"] == "routing":
|
| 101 |
categories = example["categories"]
|
| 102 |
assert len(result["items"]) == len(categories)
|
|
|
|
| 130 |
parser.error("Unknown task")
|
| 131 |
example_count = 0
|
| 132 |
for task in selected:
|
| 133 |
+
examples = task["examples"] if task["kind"] in {"clustering", "hallucination"} else task["examples"][:1]
|
| 134 |
for example in examples:
|
| 135 |
payload = example_payload(task, example)
|
| 136 |
response = request(args.base_url, "/api/analyze", payload)
|
static/app.js
CHANGED
|
@@ -8,6 +8,7 @@ const presentation = {
|
|
| 8 |
clustering: {name:'Clustering',group:'Retrieve',icon:'⁙',title:'Shared meaning. Natural company.',description:'Discover groups in a collection of texts, without predefined categories.',action:'Find the groups',note:'Adjust granularity to merge or separate groups. Open a group to inspect every text.',examples:[]},
|
| 9 |
domain: {name:'Domain',group:'Understand',icon:'◈',title:'Every question has a world.',description:'Find the subject behind a request, across 14 knowledge domains.',action:'Find the domain',note:'Scores describe the model’s distribution across its domain labels.',examples:['A curious question','Another field']},
|
| 10 |
factcheck: {name:'Fact check',group:'Understand',icon:'⊙',title:'Know when to look it up.',description:'Recognize requests that call for external factual knowledge.',action:'Check the need',note:'This predicts whether a fact check is needed. It does not verify whether a claim is true.',examples:['Factual knowledge','Creative writing']},
|
|
|
|
| 11 |
modality: {name:'Modality',group:'Understand',icon:'▧',title:'The right medium for the idea.',description:'Explore whether a request calls for text, an image, or both.',action:'Find the medium',note:'AR means text generation; DIFFUSION means image generation; BOTH means both.',examples:['An image idea','Words and pictures']},
|
| 12 |
feedback: {name:'Feedback',group:'Understand',icon:'☷',title:'Listen a little closer.',description:'Understand the feedback hidden in a conversational response.',action:'Read the feedback',note:'Predictions distinguish satisfaction, clarification, correction, a different answer, and no feedback.',examples:['A little gratitude','A correction']},
|
| 13 |
pii: {name:'Personal data',group:'Protect',icon:'◎',title:'Find what should stay private.',description:'Locate personal information, then see it highlighted or redacted.',action:'Detect PII',note:'Highlights show model predictions. Review them before relying on redaction.',examples:["Contact details","Date & age","Street address","Network log","Test card","Synthetic ID","Work profile"]},
|
|
@@ -35,16 +36,18 @@ function renderEmpty() {
|
|
| 35 |
FlowUI.setTargets([]);
|
| 36 |
if(selected==='routing' && catalog.routing) { RoutingUI.renderEmpty(); return; }
|
| 37 |
if(selected==='clustering') { ClusteringUI.renderEmpty(); return; }
|
|
|
|
| 38 |
$('output').innerHTML = '<div class="empty-state"><span class="flow-empty-label">A SIGNAL, WAITING TO BE FOUND</span><strong>Ready when you are.</strong><p>Choose an example or bring your own text. Run the model to see what it notices.</p></div>';
|
| 39 |
$('result-meta').textContent=''; $('output-tabs').hidden=true; $('copy-button').hidden=true;
|
| 40 |
$('output').setAttribute('aria-live','polite');
|
| 41 |
}
|
| 42 |
function setModelState(message) { $('model-state').replaceChildren(textNode('span','','hollow-dot'),textNode('span',message)); }
|
| 43 |
function selectTask(id) {
|
| 44 |
-
if (!presentation[id] || (busy && !['routing','clustering'].includes(selected))) return;
|
| 45 |
if(busy) invalidateRouting();
|
| 46 |
RoutingUI.leave();
|
| 47 |
ClusteringUI.leave();
|
|
|
|
| 48 |
FlowUI.leave();
|
| 49 |
selected=id; lastResult=null; edited=false; view='highlight'; $('error').hidden=true;
|
| 50 |
$('result-meta').classList.remove('stale-note');
|
|
@@ -55,13 +58,14 @@ function selectTask(id) {
|
|
| 55 |
$('task-title').textContent=p.title; $('task-description').textContent=p.description; $('task-symbol').textContent=p.icon;
|
| 56 |
$('run-label').textContent=p.action; $('task-note').textContent=p.note;
|
| 57 |
const retrieval=['embedding','reranker'].includes(id);
|
| 58 |
-
$('candidates-wrap').hidden=!retrieval; $('input-label').textContent=id==='routing'?'YOUR PROMPT':id==='clustering'?'YOUR TEXTS':retrieval?'YOUR QUERY':'YOUR TEXT';
|
| 59 |
$('candidates-label').textContent=id==='embedding'?'CANDIDATES':'PASSAGES';
|
| 60 |
-
$('input-hint').textContent=id==='clustering'?'3–24 texts · 128 tokens each':'Up to 512 tokens';
|
| 61 |
-
$('input').
|
| 62 |
-
$('input').
|
| 63 |
-
|
| 64 |
-
|
|
|
|
| 65 |
$('model-link').textContent=`Vela-1.0 · ${task ? task.model.split('-').pop() : p.name} ↗`;
|
| 66 |
$('model-link').href=`https://huggingface.co/${task?.model || 'collections/llm-semantic-router/vela-10'}`;
|
| 67 |
setModelState(cachedTasks.includes(task?.model_task || id)?'Ready on CPU':'Loads on first run');
|
|
@@ -74,17 +78,18 @@ function selectTask(id) {
|
|
| 74 |
renderExamples();
|
| 75 |
FlowUI.enter(id,p,task);
|
| 76 |
if(id==='clustering')ClusteringUI.enter();
|
|
|
|
| 77 |
if(id==='routing' && task?.scenarios) RoutingUI.enter(task);
|
| 78 |
else if (task?.examples?.length) loadExample(task.examples.find(example=>!grouped || example.group===similarityGroup));
|
| 79 |
countCharacters(); renderEmpty();
|
| 80 |
const activeButton=$('tasks').querySelector(`[data-task="${id}"]`), nav=$('tasks');
|
| 81 |
if(activeButton && nav.scrollWidth>nav.clientWidth) nav.scrollLeft+=activeButton.getBoundingClientRect().left-nav.getBoundingClientRect().left-(nav.clientWidth-activeButton.clientWidth)/2;
|
| 82 |
}
|
| 83 |
-
function countCharacters() { $('char-count').textContent=selected==='clustering'?`${ClusteringUI.texts().length} / 24 texts`:`${Array.from($('input').value).length.toLocaleString()} characters`; }
|
| 84 |
function inputChanged() {
|
| 85 |
countCharacters(); updateExampleNote(); $('error').hidden=true;
|
| 86 |
if(selected==='routing'){RoutingUI.inputChanged();return;}
|
| 87 |
-
if(
|
| 88 |
if(selected==='embedding'){
|
| 89 |
lastResult=null;edited=false;$('result-meta').classList.remove('stale-note');renderEmpty();
|
| 90 |
if(!busy)setModelState(cachedTasks.includes('embedding')?'Ready on CPU':'Loads on first run');
|
|
@@ -108,8 +113,10 @@ function renderExamples() {
|
|
| 108 |
}
|
| 109 |
}
|
| 110 |
function loadExample(example) {
|
| 111 |
-
if(!example || (busy && !['routing','clustering'].includes(selected)))return;
|
| 112 |
-
$('input').value=selected==='clustering'?(example.texts || []).join('\n'):example.text;$('candidates').value=(example.candidates || []).join('\n');
|
|
|
|
|
|
|
| 113 |
}
|
| 114 |
function selectSimilarityGroup() {
|
| 115 |
if(selected!=='embedding')return;
|
|
@@ -120,10 +127,11 @@ function selectSimilarityGroup() {
|
|
| 120 |
}
|
| 121 |
function setBusy(value) {
|
| 122 |
busy=value; document.body.classList.toggle('is-running',value);
|
| 123 |
-
const editable=['routing','clustering'].includes(selected);
|
| 124 |
for (const el of document.querySelectorAll('.task-button,.example-button,#clear-button,#similarity-group')) el.disabled=value && !editable;
|
| 125 |
$('run-button').disabled=value;
|
| 126 |
$('input').readOnly=value && !editable; $('candidates').readOnly=value;
|
|
|
|
| 127 |
$('run-label').textContent=value?'Finding the signal…':presentation[selected].action;
|
| 128 |
$('run-arrow').textContent=value?'◌':'↗'; $('output').setAttribute('aria-busy',String(value));
|
| 129 |
}
|
|
@@ -135,6 +143,7 @@ function renderLoading() {
|
|
| 135 |
$('output-tabs').hidden=true; $('copy-button').hidden=true; $('result-meta').textContent='';
|
| 136 |
$('output').innerHTML='<div class="empty-state loading"><span class="flow-empty-label">FOLLOWING THE SIGNAL</span><strong id="loading-title">Finding you a place…</strong><p id="loading-copy">Your request will wait its turn on the shared CPU.</p><div class="loading-track" aria-hidden="true"></div><button id="cancel-button" class="text-button cancel-button" type="button">Cancel request</button></div>';
|
| 137 |
if(selected==='clustering')$('loading-title').textContent='Gathering shared meaning…';
|
|
|
|
| 138 |
$('cancel-button').addEventListener('click',cancelRun);
|
| 139 |
setModelState('Submitting request…');
|
| 140 |
}
|
|
@@ -176,7 +185,7 @@ function cancelRun() {
|
|
| 176 |
}
|
| 177 |
}
|
| 178 |
function invalidateRouting() {
|
| 179 |
-
//
|
| 180 |
// publish into new inputs, even when its submission response arrives late.
|
| 181 |
if(busy) {
|
| 182 |
cancelRun();
|
|
@@ -190,16 +199,17 @@ function invalidateRouting() {
|
|
| 190 |
$('error').hidden=true;
|
| 191 |
$('result-meta').classList.remove('stale-note');
|
| 192 |
renderEmpty();
|
| 193 |
-
setModelState(cachedTasks.includes(catalog[selected]?.model_task || selected)?'Ready on CPU':selected==='clustering'?'Ready to group':'Ready to route');
|
| 194 |
}
|
| 195 |
async function analyze() {
|
| 196 |
if (busy) return;
|
| 197 |
const taskId=selected, input=$('input').value, candidates=$('candidates').value.split('\n').map(t=>t.trim()).filter(Boolean);
|
|
|
|
| 198 |
if (!input.trim()) { showError(taskId==='clustering'?'Add a collection of texts or choose an example first.':'Write a prompt or choose an example first.'); $('input').focus(); return; }
|
| 199 |
if (['embedding','reranker'].includes(taskId) && (!candidates.length || candidates.length>6)) { showError(`Add between 1 and 6 ${taskId==='embedding'?'candidates':'passages'}, one per line.`); return; }
|
| 200 |
if(taskId==='routing') { const error=RoutingUI.validate(); if(error){showError(error);return;} }
|
| 201 |
if(taskId==='clustering') { const error=ClusteringUI.validate(); if(error){showError(error);return;} }
|
| 202 |
-
const payload=taskId==='clustering'?{task:taskId,texts:ClusteringUI.texts()}:{task:taskId,text:input,candidates:['embedding','reranker'].includes(taskId)?candidates:[]};
|
| 203 |
if(taskId==='routing')payload.categories=RoutingUI.categories();
|
| 204 |
$('error').hidden=true; lastResult=null; edited=false; activeJob=null; cancelRequested=false;
|
| 205 |
const controller=new AbortController(); runController=controller;
|
|
@@ -231,6 +241,7 @@ async function analyze() {
|
|
| 231 |
let pollFailures=0;
|
| 232 |
while(job && !cancelRequested && isCurrent() && Date.now()<deadline) {
|
| 233 |
if(job.status==='completed') {
|
|
|
|
| 234 |
lastResult=job.result; lastInput=input; view='highlight'; clearInterval(runTimer);
|
| 235 |
renderResult(); setModelState(lastResult.cache_hit?'Example result reused':'Ready on CPU'); break;
|
| 236 |
}
|
|
@@ -275,8 +286,8 @@ async function analyze() {
|
|
| 275 |
}
|
| 276 |
function updateExampleNote() {
|
| 277 |
const candidates=$('candidates').value.split('\n').map(text=>text.trim()).filter(Boolean);
|
| 278 |
-
const match=catalog[selected]?.examples?.findIndex(example=>selected==='clustering'?JSON.stringify(example.texts || [])===JSON.stringify(ClusteringUI.texts()):example.text===$('input').value && (selected==='embedding'?(example.candidates || []).join('\n')===$('candidates').value:JSON.stringify(example.candidates || [])===JSON.stringify(candidates)));
|
| 279 |
-
if(['embedding','clustering'].includes(selected))$('examples').querySelectorAll('.example-button').forEach(button=>button.setAttribute('aria-pressed',String(Number(button.dataset.exampleIndex)===match)));
|
| 280 |
const message=catalog[selected]?.example_notes?.[String(match)] || '';
|
| 281 |
const note=$('example-note'); note.hidden=!message; note.textContent=message;
|
| 282 |
}
|
|
@@ -291,6 +302,7 @@ function renderResult() {
|
|
| 291 |
else if (['ranking','similarity'].includes(result.kind)) renderRanking(result);
|
| 292 |
else if (result.kind==='routing') RoutingUI.renderResult(result);
|
| 293 |
else if (result.kind==='clustering') ClusteringUI.renderResult(result);
|
|
|
|
| 294 |
else { showError('This result format could not be displayed. Please try again.'); return; }
|
| 295 |
if(!['routing','clustering'].includes(result.kind)) {
|
| 296 |
updateResultConnections(result);
|
|
@@ -303,6 +315,10 @@ function renderResult() {
|
|
| 303 |
}
|
| 304 |
function updateResultConnections(result) {
|
| 305 |
if(lastResult?.result!==result || selected==='routing') return;
|
|
|
|
|
|
|
|
|
|
|
|
|
| 306 |
if(result.kind==='pii') {
|
| 307 |
const text=$('output').querySelector('.pii-text,.redacted') || $('output');
|
| 308 |
FlowUI.setTargets([{element:text,winner:!edited}]);
|
|
@@ -357,8 +373,9 @@ function renderRanking(result) {
|
|
| 357 |
if(result.kind==='similarity') $('output').append(textNode('div',`${result.dimensions} dimensions · normalized embeddings`,'result-caption'));
|
| 358 |
}
|
| 359 |
async function copyText(value,button) { try {await navigator.clipboard.writeText(value); const before=button.textContent;button.textContent='Copied';setTimeout(()=>button.textContent=before,1500);}catch{showError('Clipboard access is unavailable. Select the result text to copy it.');} }
|
| 360 |
-
$('run-button').addEventListener('click',analyze); $('clear-button').addEventListener('click',()=>{$('input').value='';$('candidates').value='';inputChanged();$('input').focus();});
|
| 361 |
$('input').addEventListener('input',inputChanged);$('candidates').addEventListener('input',inputChanged);
|
|
|
|
| 362 |
$('similarity-group').addEventListener('change',selectSimilarityGroup);
|
| 363 |
$('highlight-button').addEventListener('click',()=>{view='highlight';renderResult();});$('redact-button').addEventListener('click',()=>{view='redact';renderResult();});
|
| 364 |
$('copy-button').addEventListener('click',()=>{if(lastResult)copyText(JSON.stringify(lastResult.result.kind==='clustering'?ClusteringUI.exportResult():lastResult.result,null,2),$('copy-button'));});
|
|
|
|
| 8 |
clustering: {name:'Clustering',group:'Retrieve',icon:'⁙',title:'Shared meaning. Natural company.',description:'Discover groups in a collection of texts, without predefined categories.',action:'Find the groups',note:'Adjust granularity to merge or separate groups. Open a group to inspect every text.',examples:[]},
|
| 9 |
domain: {name:'Domain',group:'Understand',icon:'◈',title:'Every question has a world.',description:'Find the subject behind a request, across 14 knowledge domains.',action:'Find the domain',note:'Scores describe the model’s distribution across its domain labels.',examples:['A curious question','Another field']},
|
| 10 |
factcheck: {name:'Fact check',group:'Understand',icon:'⊙',title:'Know when to look it up.',description:'Recognize requests that call for external factual knowledge.',action:'Check the need',note:'This predicts whether a fact check is needed. It does not verify whether a claim is true.',examples:['Factual knowledge','Creative writing']},
|
| 11 |
+
hallucination: {name:'Hallucination',group:'Understand',icon:'⌁',title:'An answer, grounded in evidence.',description:'Compare an answer with its evidence. Find passages that may not be supported.',action:'Check the answer',note:'Highlights reflect the evidence you provide, not a guarantee of real-world truth.',examples:[]},
|
| 12 |
modality: {name:'Modality',group:'Understand',icon:'▧',title:'The right medium for the idea.',description:'Explore whether a request calls for text, an image, or both.',action:'Find the medium',note:'AR means text generation; DIFFUSION means image generation; BOTH means both.',examples:['An image idea','Words and pictures']},
|
| 13 |
feedback: {name:'Feedback',group:'Understand',icon:'☷',title:'Listen a little closer.',description:'Understand the feedback hidden in a conversational response.',action:'Read the feedback',note:'Predictions distinguish satisfaction, clarification, correction, a different answer, and no feedback.',examples:['A little gratitude','A correction']},
|
| 14 |
pii: {name:'Personal data',group:'Protect',icon:'◎',title:'Find what should stay private.',description:'Locate personal information, then see it highlighted or redacted.',action:'Detect PII',note:'Highlights show model predictions. Review them before relying on redaction.',examples:["Contact details","Date & age","Street address","Network log","Test card","Synthetic ID","Work profile"]},
|
|
|
|
| 36 |
FlowUI.setTargets([]);
|
| 37 |
if(selected==='routing' && catalog.routing) { RoutingUI.renderEmpty(); return; }
|
| 38 |
if(selected==='clustering') { ClusteringUI.renderEmpty(); return; }
|
| 39 |
+
if(selected==='hallucination') { HallucinationUI.renderEmpty(); return; }
|
| 40 |
$('output').innerHTML = '<div class="empty-state"><span class="flow-empty-label">A SIGNAL, WAITING TO BE FOUND</span><strong>Ready when you are.</strong><p>Choose an example or bring your own text. Run the model to see what it notices.</p></div>';
|
| 41 |
$('result-meta').textContent=''; $('output-tabs').hidden=true; $('copy-button').hidden=true;
|
| 42 |
$('output').setAttribute('aria-live','polite');
|
| 43 |
}
|
| 44 |
function setModelState(message) { $('model-state').replaceChildren(textNode('span','','hollow-dot'),textNode('span',message)); }
|
| 45 |
function selectTask(id) {
|
| 46 |
+
if (!presentation[id] || (busy && !['routing','clustering','hallucination'].includes(selected))) return;
|
| 47 |
if(busy) invalidateRouting();
|
| 48 |
RoutingUI.leave();
|
| 49 |
ClusteringUI.leave();
|
| 50 |
+
HallucinationUI.leave();
|
| 51 |
FlowUI.leave();
|
| 52 |
selected=id; lastResult=null; edited=false; view='highlight'; $('error').hidden=true;
|
| 53 |
$('result-meta').classList.remove('stale-note');
|
|
|
|
| 58 |
$('task-title').textContent=p.title; $('task-description').textContent=p.description; $('task-symbol').textContent=p.icon;
|
| 59 |
$('run-label').textContent=p.action; $('task-note').textContent=p.note;
|
| 60 |
const retrieval=['embedding','reranker'].includes(id);
|
| 61 |
+
$('candidates-wrap').hidden=!retrieval; $('input-label').textContent=id==='routing'?'YOUR PROMPT':id==='clustering'?'YOUR TEXTS':id==='hallucination'?'YOUR QUESTION':retrieval?'YOUR QUERY':'YOUR TEXT';
|
| 62 |
$('candidates-label').textContent=id==='embedding'?'CANDIDATES':'PASSAGES';
|
| 63 |
+
$('input-hint').textContent=id==='clustering'?'3–24 texts · 128 tokens each':id==='hallucination'?'512 tokens total · inputs are never shortened':'Up to 512 tokens';
|
| 64 |
+
if(id==='hallucination')$('input').removeAttribute('maxlength');
|
| 65 |
+
else $('input').maxLength=id==='clustering'?98351:12000;
|
| 66 |
+
$('input').placeholder=id==='clustering'?'One text per line. Let shared meaning find its shape.':id==='hallucination'?'What question is the answer responding to?':'Write a little. Discover a lot.';
|
| 67 |
+
document.querySelector('.examples>span').textContent=id==='clustering'?'TRY A COLLECTION':id==='hallucination'?'TRY AN EXAMPLE':'TRY A PROMPT';
|
| 68 |
+
$('input').value=''; $('candidates').value=''; HallucinationUI.clear();
|
| 69 |
$('model-link').textContent=`Vela-1.0 · ${task ? task.model.split('-').pop() : p.name} ↗`;
|
| 70 |
$('model-link').href=`https://huggingface.co/${task?.model || 'collections/llm-semantic-router/vela-10'}`;
|
| 71 |
setModelState(cachedTasks.includes(task?.model_task || id)?'Ready on CPU':'Loads on first run');
|
|
|
|
| 78 |
renderExamples();
|
| 79 |
FlowUI.enter(id,p,task);
|
| 80 |
if(id==='clustering')ClusteringUI.enter();
|
| 81 |
+
if(id==='hallucination')HallucinationUI.enter();
|
| 82 |
if(id==='routing' && task?.scenarios) RoutingUI.enter(task);
|
| 83 |
else if (task?.examples?.length) loadExample(task.examples.find(example=>!grouped || example.group===similarityGroup));
|
| 84 |
countCharacters(); renderEmpty();
|
| 85 |
const activeButton=$('tasks').querySelector(`[data-task="${id}"]`), nav=$('tasks');
|
| 86 |
if(activeButton && nav.scrollWidth>nav.clientWidth) nav.scrollLeft+=activeButton.getBoundingClientRect().left-nav.getBoundingClientRect().left-(nav.clientWidth-activeButton.clientWidth)/2;
|
| 87 |
}
|
| 88 |
+
function countCharacters() { $('char-count').textContent=selected==='hallucination'?`${HallucinationUI.characterCount().toLocaleString()} characters total`:selected==='clustering'?`${ClusteringUI.texts().length} / 24 texts`:`${Array.from($('input').value).length.toLocaleString()} characters`; }
|
| 89 |
function inputChanged() {
|
| 90 |
countCharacters(); updateExampleNote(); $('error').hidden=true;
|
| 91 |
if(selected==='routing'){RoutingUI.inputChanged();return;}
|
| 92 |
+
if(['clustering','hallucination'].includes(selected)){invalidateRouting();return;}
|
| 93 |
if(selected==='embedding'){
|
| 94 |
lastResult=null;edited=false;$('result-meta').classList.remove('stale-note');renderEmpty();
|
| 95 |
if(!busy)setModelState(cachedTasks.includes('embedding')?'Ready on CPU':'Loads on first run');
|
|
|
|
| 113 |
}
|
| 114 |
}
|
| 115 |
function loadExample(example) {
|
| 116 |
+
if(!example || (busy && !['routing','clustering','hallucination'].includes(selected)))return;
|
| 117 |
+
$('input').value=selected==='clustering'?(example.texts || []).join('\n'):example.text;$('candidates').value=(example.candidates || []).join('\n');
|
| 118 |
+
if(selected==='hallucination')HallucinationUI.fillExample(example);
|
| 119 |
+
inputChanged();
|
| 120 |
}
|
| 121 |
function selectSimilarityGroup() {
|
| 122 |
if(selected!=='embedding')return;
|
|
|
|
| 127 |
}
|
| 128 |
function setBusy(value) {
|
| 129 |
busy=value; document.body.classList.toggle('is-running',value);
|
| 130 |
+
const editable=['routing','clustering','hallucination'].includes(selected);
|
| 131 |
for (const el of document.querySelectorAll('.task-button,.example-button,#clear-button,#similarity-group')) el.disabled=value && !editable;
|
| 132 |
$('run-button').disabled=value;
|
| 133 |
$('input').readOnly=value && !editable; $('candidates').readOnly=value;
|
| 134 |
+
$('hallucination-context').readOnly=value && !editable; $('hallucination-answer').readOnly=value && !editable;
|
| 135 |
$('run-label').textContent=value?'Finding the signal…':presentation[selected].action;
|
| 136 |
$('run-arrow').textContent=value?'◌':'↗'; $('output').setAttribute('aria-busy',String(value));
|
| 137 |
}
|
|
|
|
| 143 |
$('output-tabs').hidden=true; $('copy-button').hidden=true; $('result-meta').textContent='';
|
| 144 |
$('output').innerHTML='<div class="empty-state loading"><span class="flow-empty-label">FOLLOWING THE SIGNAL</span><strong id="loading-title">Finding you a place…</strong><p id="loading-copy">Your request will wait its turn on the shared CPU.</p><div class="loading-track" aria-hidden="true"></div><button id="cancel-button" class="text-button cancel-button" type="button">Cancel request</button></div>';
|
| 145 |
if(selected==='clustering')$('loading-title').textContent='Gathering shared meaning…';
|
| 146 |
+
if(selected==='hallucination')$('loading-title').textContent='Reading against the evidence…';
|
| 147 |
$('cancel-button').addEventListener('click',cancelRun);
|
| 148 |
setModelState('Submitting request…');
|
| 149 |
}
|
|
|
|
| 185 |
}
|
| 186 |
}
|
| 187 |
function invalidateRouting() {
|
| 188 |
+
// Editable tasks allow changes while waiting. An old request must not
|
| 189 |
// publish into new inputs, even when its submission response arrives late.
|
| 190 |
if(busy) {
|
| 191 |
cancelRun();
|
|
|
|
| 199 |
$('error').hidden=true;
|
| 200 |
$('result-meta').classList.remove('stale-note');
|
| 201 |
renderEmpty();
|
| 202 |
+
setModelState(cachedTasks.includes(catalog[selected]?.model_task || selected)?'Ready on CPU':selected==='hallucination'?'Loads on first run':selected==='clustering'?'Ready to group':'Ready to route');
|
| 203 |
}
|
| 204 |
async function analyze() {
|
| 205 |
if (busy) return;
|
| 206 |
const taskId=selected, input=$('input').value, candidates=$('candidates').value.split('\n').map(t=>t.trim()).filter(Boolean);
|
| 207 |
+
if(taskId==='hallucination') { const error=HallucinationUI.validate(); if(error){showError(error);return;} }
|
| 208 |
if (!input.trim()) { showError(taskId==='clustering'?'Add a collection of texts or choose an example first.':'Write a prompt or choose an example first.'); $('input').focus(); return; }
|
| 209 |
if (['embedding','reranker'].includes(taskId) && (!candidates.length || candidates.length>6)) { showError(`Add between 1 and 6 ${taskId==='embedding'?'candidates':'passages'}, one per line.`); return; }
|
| 210 |
if(taskId==='routing') { const error=RoutingUI.validate(); if(error){showError(error);return;} }
|
| 211 |
if(taskId==='clustering') { const error=ClusteringUI.validate(); if(error){showError(error);return;} }
|
| 212 |
+
const payload=taskId==='hallucination'?{task:taskId,text:input,...HallucinationUI.values()}:taskId==='clustering'?{task:taskId,texts:ClusteringUI.texts()}:{task:taskId,text:input,candidates:['embedding','reranker'].includes(taskId)?candidates:[]};
|
| 213 |
if(taskId==='routing')payload.categories=RoutingUI.categories();
|
| 214 |
$('error').hidden=true; lastResult=null; edited=false; activeJob=null; cancelRequested=false;
|
| 215 |
const controller=new AbortController(); runController=controller;
|
|
|
|
| 241 |
let pollFailures=0;
|
| 242 |
while(job && !cancelRequested && isCurrent() && Date.now()<deadline) {
|
| 243 |
if(job.status==='completed') {
|
| 244 |
+
if(taskId==='hallucination' && (job.result?.result?.kind!=='hallucination' || job.result.result.answer!==payload.answer)) throw new Error('The result did not match this answer. Please run it again.');
|
| 245 |
lastResult=job.result; lastInput=input; view='highlight'; clearInterval(runTimer);
|
| 246 |
renderResult(); setModelState(lastResult.cache_hit?'Example result reused':'Ready on CPU'); break;
|
| 247 |
}
|
|
|
|
| 286 |
}
|
| 287 |
function updateExampleNote() {
|
| 288 |
const candidates=$('candidates').value.split('\n').map(text=>text.trim()).filter(Boolean);
|
| 289 |
+
const match=catalog[selected]?.examples?.findIndex(example=>selected==='hallucination'?HallucinationUI.matchesExample(example):selected==='clustering'?JSON.stringify(example.texts || [])===JSON.stringify(ClusteringUI.texts()):example.text===$('input').value && (selected==='embedding'?(example.candidates || []).join('\n')===$('candidates').value:JSON.stringify(example.candidates || [])===JSON.stringify(candidates)));
|
| 290 |
+
if(['embedding','clustering','hallucination'].includes(selected))$('examples').querySelectorAll('.example-button').forEach(button=>button.setAttribute('aria-pressed',String(Number(button.dataset.exampleIndex)===match)));
|
| 291 |
const message=catalog[selected]?.example_notes?.[String(match)] || '';
|
| 292 |
const note=$('example-note'); note.hidden=!message; note.textContent=message;
|
| 293 |
}
|
|
|
|
| 302 |
else if (['ranking','similarity'].includes(result.kind)) renderRanking(result);
|
| 303 |
else if (result.kind==='routing') RoutingUI.renderResult(result);
|
| 304 |
else if (result.kind==='clustering') ClusteringUI.renderResult(result);
|
| 305 |
+
else if (result.kind==='hallucination') HallucinationUI.renderResult(result);
|
| 306 |
else { showError('This result format could not be displayed. Please try again.'); return; }
|
| 307 |
if(!['routing','clustering'].includes(result.kind)) {
|
| 308 |
updateResultConnections(result);
|
|
|
|
| 315 |
}
|
| 316 |
function updateResultConnections(result) {
|
| 317 |
if(lastResult?.result!==result || selected==='routing') return;
|
| 318 |
+
if(result.kind==='hallucination') {
|
| 319 |
+
FlowUI.setTargets([{element:$('output').querySelector('.hallucination-answer'),winner:!edited}]);
|
| 320 |
+
return;
|
| 321 |
+
}
|
| 322 |
if(result.kind==='pii') {
|
| 323 |
const text=$('output').querySelector('.pii-text,.redacted') || $('output');
|
| 324 |
FlowUI.setTargets([{element:text,winner:!edited}]);
|
|
|
|
| 373 |
if(result.kind==='similarity') $('output').append(textNode('div',`${result.dimensions} dimensions · normalized embeddings`,'result-caption'));
|
| 374 |
}
|
| 375 |
async function copyText(value,button) { try {await navigator.clipboard.writeText(value); const before=button.textContent;button.textContent='Copied';setTimeout(()=>button.textContent=before,1500);}catch{showError('Clipboard access is unavailable. Select the result text to copy it.');} }
|
| 376 |
+
$('run-button').addEventListener('click',analyze); $('clear-button').addEventListener('click',()=>{$('input').value='';$('candidates').value='';HallucinationUI.clear();inputChanged();$('input').focus();});
|
| 377 |
$('input').addEventListener('input',inputChanged);$('candidates').addEventListener('input',inputChanged);
|
| 378 |
+
$('hallucination-context').addEventListener('input',inputChanged);$('hallucination-answer').addEventListener('input',inputChanged);
|
| 379 |
$('similarity-group').addEventListener('change',selectSimilarityGroup);
|
| 380 |
$('highlight-button').addEventListener('click',()=>{view='highlight';renderResult();});$('redact-button').addEventListener('click',()=>{view='redact';renderResult();});
|
| 381 |
$('copy-button').addEventListener('click',()=>{if(lastResult)copyText(JSON.stringify(lastResult.result.kind==='clustering'?ClusteringUI.exportResult():lastResult.result,null,2),$('copy-button'));});
|
static/flow.js
CHANGED
|
@@ -1,4 +1,4 @@
|
|
| 1 |
-
/* One shared canvas and model node for all
|
| 2 |
'use strict';
|
| 3 |
window.FlowUI = (() => {
|
| 4 |
const el = id => document.getElementById(id);
|
|
@@ -85,7 +85,7 @@ window.FlowUI = (() => {
|
|
| 85 |
const body = node('div', '', 'flow-model-body');
|
| 86 |
el('model-link').textContent = `Vela ${name} ↗`;
|
| 87 |
const metadata = node('div', '', 'flow-model-metadata');
|
| 88 |
-
const variant = ['embedding', 'routing', 'clustering'].includes(id) ? '768D' : {classification:'Classifier', pii:'Token tags', hazard:'Multi-label', ranking:'Cross-encoder'}[task?.kind] || 'Encoder';
|
| 89 |
metadata.append(node('span', '307M'), node('i', '·'), node('span', variant));
|
| 90 |
body.append(el('model-link'), metadata, el('model-state'));
|
| 91 |
modelCard.append(art, body);
|
|
|
|
| 1 |
+
/* One shared canvas and model node for all Vela demos. */
|
| 2 |
'use strict';
|
| 3 |
window.FlowUI = (() => {
|
| 4 |
const el = id => document.getElementById(id);
|
|
|
|
| 85 |
const body = node('div', '', 'flow-model-body');
|
| 86 |
el('model-link').textContent = `Vela ${name} ↗`;
|
| 87 |
const metadata = node('div', '', 'flow-model-metadata');
|
| 88 |
+
const variant = ['embedding', 'routing', 'clustering'].includes(id) ? '768D' : {classification:'Classifier', pii:'Token tags', hazard:'Multi-label', hallucination:'Grounding', ranking:'Cross-encoder'}[task?.kind] || 'Encoder';
|
| 89 |
metadata.append(node('span', '307M'), node('i', '·'), node('span', variant));
|
| 90 |
body.append(el('model-link'), metadata, el('model-state'));
|
| 91 |
modelCard.append(art, body);
|
static/hallucination.css
ADDED
|
@@ -0,0 +1,75 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/* Three source fields and quiet copper annotations share the Vela field guide. */
|
| 2 |
+
.hallucination-mode #input {
|
| 3 |
+
min-height: 82px;
|
| 4 |
+
max-height: 170px;
|
| 5 |
+
font-size: 13px;
|
| 6 |
+
}
|
| 7 |
+
.hallucination-fields label {
|
| 8 |
+
display: block;
|
| 9 |
+
border-top: 1px solid #e2e7dc;
|
| 10 |
+
padding: 12px 16px 0;
|
| 11 |
+
color: #61725a;
|
| 12 |
+
font: 8px var(--mono);
|
| 13 |
+
letter-spacing: 1px;
|
| 14 |
+
}
|
| 15 |
+
.hallucination-mode .hallucination-fields textarea {
|
| 16 |
+
display: block;
|
| 17 |
+
min-height: 110px;
|
| 18 |
+
max-height: 240px;
|
| 19 |
+
width: 100%;
|
| 20 |
+
padding: 10px 16px 14px;
|
| 21 |
+
font-size: 12px;
|
| 22 |
+
line-height: 1.8;
|
| 23 |
+
resize: vertical;
|
| 24 |
+
}
|
| 25 |
+
.hallucination-mode .editor-foot {
|
| 26 |
+
align-items: flex-start;
|
| 27 |
+
gap: 12px;
|
| 28 |
+
line-height: 1.65;
|
| 29 |
+
}
|
| 30 |
+
.hallucination-mode #char-count { white-space: nowrap; }
|
| 31 |
+
.hallucination-mode .hallucination-result-title {
|
| 32 |
+
font-size: 25px;
|
| 33 |
+
line-height: 1.3;
|
| 34 |
+
}
|
| 35 |
+
.hallucination-answer {
|
| 36 |
+
border: 1px solid #d9dfcf;
|
| 37 |
+
border-radius: 5px;
|
| 38 |
+
padding: 18px;
|
| 39 |
+
background: #fcfaf5;
|
| 40 |
+
color: var(--ink);
|
| 41 |
+
font-size: 13px;
|
| 42 |
+
line-height: 2;
|
| 43 |
+
white-space: pre-wrap;
|
| 44 |
+
overflow-wrap: anywhere;
|
| 45 |
+
max-height: 520px;
|
| 46 |
+
overflow: auto;
|
| 47 |
+
}
|
| 48 |
+
.hallucination-span {
|
| 49 |
+
background: #ead5b8;
|
| 50 |
+
color: #5b402b;
|
| 51 |
+
border-bottom: 1px solid #b78053;
|
| 52 |
+
border-radius: 2px;
|
| 53 |
+
box-decoration-break: clone;
|
| 54 |
+
-webkit-box-decoration-break: clone;
|
| 55 |
+
}
|
| 56 |
+
.hallucination-legend {
|
| 57 |
+
margin-top: 12px;
|
| 58 |
+
padding-left: 11px;
|
| 59 |
+
border-left: 3px solid #bb895d;
|
| 60 |
+
color: #795437;
|
| 61 |
+
font-size: 9px;
|
| 62 |
+
line-height: 1.7;
|
| 63 |
+
}
|
| 64 |
+
.hallucination-result-note {
|
| 65 |
+
margin: 13px 0 0;
|
| 66 |
+
color: #66735d;
|
| 67 |
+
font-size: 10px;
|
| 68 |
+
line-height: 1.8;
|
| 69 |
+
}
|
| 70 |
+
@media (max-width: 700px) {
|
| 71 |
+
.hallucination-mode #input { min-height: 90px; font-size: 14px; }
|
| 72 |
+
.hallucination-mode .hallucination-fields textarea { font-size: 14px; min-height: 115px; }
|
| 73 |
+
.hallucination-answer { padding: 15px; }
|
| 74 |
+
.hallucination-mode .editor-foot { flex-direction: column; gap: 3px; }
|
| 75 |
+
}
|
static/hallucination.js
ADDED
|
@@ -0,0 +1,94 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/* Grounding highlights use the answer's Unicode code-point offsets. */
|
| 2 |
+
'use strict';
|
| 3 |
+
window.HallucinationUI = (() => {
|
| 4 |
+
const el = id => document.getElementById(id);
|
| 5 |
+
const node = (tag, text = '', className = '') => {
|
| 6 |
+
const element = document.createElement(tag);
|
| 7 |
+
element.textContent = text;
|
| 8 |
+
if (className) element.className = className;
|
| 9 |
+
return element;
|
| 10 |
+
};
|
| 11 |
+
function enter() {
|
| 12 |
+
document.querySelector('.task-main').classList.add('hallucination-mode');
|
| 13 |
+
el('hallucination-fields').hidden = false;
|
| 14 |
+
}
|
| 15 |
+
function leave() {
|
| 16 |
+
document.querySelector('.task-main').classList.remove('hallucination-mode');
|
| 17 |
+
el('hallucination-fields').hidden = true;
|
| 18 |
+
}
|
| 19 |
+
function values() {
|
| 20 |
+
return {context: el('hallucination-context').value, answer: el('hallucination-answer').value};
|
| 21 |
+
}
|
| 22 |
+
function fillExample(example) {
|
| 23 |
+
el('hallucination-context').value = example.context || '';
|
| 24 |
+
el('hallucination-answer').value = example.answer || '';
|
| 25 |
+
}
|
| 26 |
+
function clear() { fillExample({}); }
|
| 27 |
+
function matchesExample(example) {
|
| 28 |
+
const current = values();
|
| 29 |
+
return example.text === el('input').value && example.context === current.context && example.answer === current.answer;
|
| 30 |
+
}
|
| 31 |
+
function validate() {
|
| 32 |
+
for (const [id, message] of [
|
| 33 |
+
['input', 'Add the question or choose an example first.'],
|
| 34 |
+
['hallucination-context', 'Add the evidence the answer should be based on.'],
|
| 35 |
+
['hallucination-answer', 'Add the answer you want to check.'],
|
| 36 |
+
]) {
|
| 37 |
+
if (!el(id).value.trim()) { el(id).focus(); return message; }
|
| 38 |
+
}
|
| 39 |
+
return null;
|
| 40 |
+
}
|
| 41 |
+
function characterCount() {
|
| 42 |
+
const current = values();
|
| 43 |
+
return [el('input').value, current.context, current.answer].reduce((total, value) => total + Array.from(value).length, 0);
|
| 44 |
+
}
|
| 45 |
+
function renderEmpty() {
|
| 46 |
+
const empty = node('div', '', 'empty-state hallucination-empty');
|
| 47 |
+
empty.append(node('span', 'AN ANSWER, IN THE LIGHT OF EVIDENCE', 'flow-empty-label'));
|
| 48 |
+
empty.append(node('strong', 'See what the evidence supports.'));
|
| 49 |
+
empty.append(node('p', 'Add a question, its evidence, and an answer. Potentially unsupported passages will appear here.'));
|
| 50 |
+
el('output').replaceChildren(empty);
|
| 51 |
+
el('result-meta').textContent = '';
|
| 52 |
+
el('output-tabs').hidden = true;
|
| 53 |
+
el('copy-button').hidden = true;
|
| 54 |
+
el('output').setAttribute('aria-live', 'polite');
|
| 55 |
+
}
|
| 56 |
+
function segments(answer, spans) {
|
| 57 |
+
if (typeof answer !== 'string' || !answer.trim() || !Array.isArray(spans)) throw new Error('The answer highlights could not be read. Please try again.');
|
| 58 |
+
const chars = Array.from(answer), result = [];
|
| 59 |
+
let previous = 0;
|
| 60 |
+
for (const span of [...spans].sort((a, b) => a.start - b.start)) {
|
| 61 |
+
if (!Number.isInteger(span.start) || !Number.isInteger(span.end) || span.start < previous || span.end <= span.start || span.end > chars.length || !Number.isFinite(span.score) || span.score < 0 || span.score > 1) {
|
| 62 |
+
throw new Error('The answer highlights did not align with the text. Please try again.');
|
| 63 |
+
}
|
| 64 |
+
const text = chars.slice(span.start, span.end).join('');
|
| 65 |
+
if (span.text !== undefined && span.text !== text) throw new Error('The answer highlights did not match the text. Please try again.');
|
| 66 |
+
if (span.start > previous) result.push({text: chars.slice(previous, span.start).join(''), highlighted: false});
|
| 67 |
+
result.push({text, highlighted: true, score: span.score});
|
| 68 |
+
previous = span.end;
|
| 69 |
+
}
|
| 70 |
+
if (previous < chars.length) result.push({text: chars.slice(previous).join(''), highlighted: false});
|
| 71 |
+
return result;
|
| 72 |
+
}
|
| 73 |
+
function renderResult(result) {
|
| 74 |
+
// Validate the entire response before creating a reassuring empty-span state.
|
| 75 |
+
const parts = segments(result.answer, result.spans);
|
| 76 |
+
const count = result.spans.length;
|
| 77 |
+
const title = count ? `${count} potentially unsupported ${count === 1 ? 'span' : 'spans'}` : 'No unsupported spans detected';
|
| 78 |
+
const heading = node('div', title, 'result-title hallucination-result-title');
|
| 79 |
+
const answer = node('div', '', 'hallucination-answer');
|
| 80 |
+
answer.setAttribute('aria-label', 'Answer with potentially unsupported spans highlighted');
|
| 81 |
+
for (const part of parts) {
|
| 82 |
+
if (!part.highlighted) { answer.append(document.createTextNode(part.text)); continue; }
|
| 83 |
+
const mark = node('mark', part.text, 'hallucination-span');
|
| 84 |
+
mark.title = `Potentially unsupported · model score ${(part.score * 100).toFixed(1)}%`;
|
| 85 |
+
answer.append(mark);
|
| 86 |
+
}
|
| 87 |
+
const note = node('p', 'Assessed against the evidence you provided. This does not guarantee factual accuracy beyond that evidence.', 'hallucination-result-note');
|
| 88 |
+
const children = [heading, answer];
|
| 89 |
+
if (count) children.push(node('div', 'Highlighted · potentially unsupported by the evidence', 'hallucination-legend'));
|
| 90 |
+
children.push(note);
|
| 91 |
+
el('output').replaceChildren(...children);
|
| 92 |
+
}
|
| 93 |
+
return {enter, leave, values, fillExample, clear, matchesExample, validate, characterCount, renderEmpty, renderResult, segments};
|
| 94 |
+
})();
|
static/index.html
CHANGED
|
@@ -4,7 +4,7 @@
|
|
| 4 |
<meta charset="utf-8">
|
| 5 |
<meta name="viewport" content="width=device-width, initial-scale=1">
|
| 6 |
<meta name="theme-color" content="#f5f2e9">
|
| 7 |
-
<meta name="description" content="Explore Vela 1.0:
|
| 8 |
<title>Vela Studio — A better direction.</title>
|
| 9 |
<link rel="icon" href="/static/favicon.svg" type="image/svg+xml">
|
| 10 |
<link rel="preconnect" href="https://fonts.googleapis.com">
|
|
@@ -14,10 +14,12 @@
|
|
| 14 |
<link rel="stylesheet" href="/static/flow.css?v=e7df90552db0">
|
| 15 |
<link rel="stylesheet" href="/static/routing.css?v=d2bea5f416">
|
| 16 |
<link rel="stylesheet" href="/static/clustering.css?v=1a774c530cb1">
|
| 17 |
-
<
|
|
|
|
|
|
|
| 18 |
<script src="/static/routing.js?v=c3f8df099d" defer></script>
|
| 19 |
<script src="/static/clustering.js?v=531c6c50bc4f" defer></script>
|
| 20 |
-
<script src="/static/app.js?v=
|
| 21 |
</head>
|
| 22 |
<body>
|
| 23 |
<a class="skip-link" href="#studio">Skip to playground</a>
|
|
@@ -41,7 +43,7 @@
|
|
| 41 |
<span class="art-caption">FIG. 01 — A LITTLE INTELLIGENCE GOES A LONG WAY.</span>
|
| 42 |
</div>
|
| 43 |
</section>
|
| 44 |
-
<div class="specimen-strip"><span><b>307M</b> parameters per model</span><span><b>
|
| 45 |
<main id="studio" tabindex="-1">
|
| 46 |
<div class="section-title"><div><span class="eyebrow">THE INTERACTIVE FIELD GUIDE</span><h2>The signal room<span>.</span></h2></div><div class="session-badge"><i class="status-dot"></i><span>Vela 1.0</span><span class="badge-divider">/</span> CPU</div></div>
|
| 47 |
<div class="workbench">
|
|
@@ -51,7 +53,7 @@
|
|
| 51 |
<div class="model-strip"><a id="model-link" target="_blank" rel="noopener">Vela-1.0 · Embedding <span>↗</span></a><span id="model-state" role="status" aria-live="polite"><span class="hollow-dot"></span> Loads on first run</span></div>
|
| 52 |
<section id="routing-scenarios" class="routing-scenarios" aria-label="Routing scenario" hidden></section>
|
| 53 |
<div class="panels">
|
| 54 |
-
<section class="input-panel"><div class="panel-heading"><label id="input-label" for="input">YOUR PROMPT</label><button class="text-button" id="clear-button">Clear <span>×</span></button></div><textarea id="input" maxlength="12000" spellcheck="false" placeholder="Write a little. Discover a lot." aria-describedby="input-hint"></textarea><div id="candidates-wrap" hidden><label for="candidates"><span id="candidates-label">PASSAGES</span> <span>One per line · up to 6</span></label><textarea id="candidates" maxlength="36000" rows="5" spellcheck="false"></textarea></div><div class="editor-foot"><span id="input-hint">Up to 512 tokens</span><span id="char-count">0 characters</span></div></section>
|
| 55 |
<section class="output-panel"><div class="panel-heading"><span>THE SIGNAL</span><div id="output-tabs" class="output-tabs" aria-label="PII output view" hidden><button id="highlight-button" aria-pressed="true">Highlight</button><button id="redact-button" aria-pressed="false">Redact</button></div><button class="text-button" id="copy-button" hidden>Copy</button></div><div id="output" class="output" aria-live="polite" aria-atomic="true"></div><div id="result-meta" class="result-meta"></div></section>
|
| 56 |
</div>
|
| 57 |
<div id="example-note" class="example-note" role="note" hidden></div>
|
|
|
|
| 4 |
<meta charset="utf-8">
|
| 5 |
<meta name="viewport" content="width=device-width, initial-scale=1">
|
| 6 |
<meta name="theme-color" content="#f5f2e9">
|
| 7 |
+
<meta name="description" content="Explore Vela 1.0: eleven specialized models for understanding, protecting, grounding, and retrieving language.">
|
| 8 |
<title>Vela Studio — A better direction.</title>
|
| 9 |
<link rel="icon" href="/static/favicon.svg" type="image/svg+xml">
|
| 10 |
<link rel="preconnect" href="https://fonts.googleapis.com">
|
|
|
|
| 14 |
<link rel="stylesheet" href="/static/flow.css?v=e7df90552db0">
|
| 15 |
<link rel="stylesheet" href="/static/routing.css?v=d2bea5f416">
|
| 16 |
<link rel="stylesheet" href="/static/clustering.css?v=1a774c530cb1">
|
| 17 |
+
<link rel="stylesheet" href="/static/hallucination.css?v=ddd852a221a6">
|
| 18 |
+
<script src="/static/hallucination.js?v=76e36d418177" defer></script>
|
| 19 |
+
<script src="/static/flow.js?v=0d15a0d480c5" defer></script>
|
| 20 |
<script src="/static/routing.js?v=c3f8df099d" defer></script>
|
| 21 |
<script src="/static/clustering.js?v=531c6c50bc4f" defer></script>
|
| 22 |
+
<script src="/static/app.js?v=8eeeb11f8795" defer></script>
|
| 23 |
</head>
|
| 24 |
<body>
|
| 25 |
<a class="skip-link" href="#studio">Skip to playground</a>
|
|
|
|
| 43 |
<span class="art-caption">FIG. 01 — A LITTLE INTELLIGENCE GOES A LONG WAY.</span>
|
| 44 |
</div>
|
| 45 |
</section>
|
| 46 |
+
<div class="specimen-strip"><span><b>307M</b> parameters per model</span><span><b>11</b> specialized models</span><span><i class="status-dot"></i> Open weights, open possibilities</span><a href="https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M" target="_blank" rel="noopener">Meet the encoder <span>↗</span></a></div>
|
| 47 |
<main id="studio" tabindex="-1">
|
| 48 |
<div class="section-title"><div><span class="eyebrow">THE INTERACTIVE FIELD GUIDE</span><h2>The signal room<span>.</span></h2></div><div class="session-badge"><i class="status-dot"></i><span>Vela 1.0</span><span class="badge-divider">/</span> CPU</div></div>
|
| 49 |
<div class="workbench">
|
|
|
|
| 53 |
<div class="model-strip"><a id="model-link" target="_blank" rel="noopener">Vela-1.0 · Embedding <span>↗</span></a><span id="model-state" role="status" aria-live="polite"><span class="hollow-dot"></span> Loads on first run</span></div>
|
| 54 |
<section id="routing-scenarios" class="routing-scenarios" aria-label="Routing scenario" hidden></section>
|
| 55 |
<div class="panels">
|
| 56 |
+
<section class="input-panel"><div class="panel-heading"><label id="input-label" for="input">YOUR PROMPT</label><button class="text-button" id="clear-button">Clear <span>×</span></button></div><textarea id="input" maxlength="12000" spellcheck="false" placeholder="Write a little. Discover a lot." aria-describedby="input-hint"></textarea><div id="hallucination-fields" class="hallucination-fields" hidden><label for="hallucination-context">EVIDENCE</label><textarea id="hallucination-context" rows="4" spellcheck="false" placeholder="Paste the source material the answer should be based on." aria-describedby="input-hint"></textarea><label for="hallucination-answer">ANSWER TO CHECK</label><textarea id="hallucination-answer" rows="4" spellcheck="false" placeholder="Paste an answer to compare with the evidence." aria-describedby="input-hint"></textarea></div><div id="candidates-wrap" hidden><label for="candidates"><span id="candidates-label">PASSAGES</span> <span>One per line · up to 6</span></label><textarea id="candidates" maxlength="36000" rows="5" spellcheck="false"></textarea></div><div class="editor-foot"><span id="input-hint">Up to 512 tokens</span><span id="char-count">0 characters</span></div></section>
|
| 57 |
<section class="output-panel"><div class="panel-heading"><span>THE SIGNAL</span><div id="output-tabs" class="output-tabs" aria-label="PII output view" hidden><button id="highlight-button" aria-pressed="true">Highlight</button><button id="redact-button" aria-pressed="false">Redact</button></div><button class="text-button" id="copy-button" hidden>Copy</button></div><div id="output" class="output" aria-live="polite" aria-atomic="true"></div><div id="result-meta" class="result-meta"></div></section>
|
| 58 |
</div>
|
| 59 |
<div id="example-note" class="example-note" role="note" hidden></div>
|
tests/test_hallucination.py
ADDED
|
@@ -0,0 +1,238 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Evidence-pair, Unicode span, queue cache, and API contracts without weights."""
|
| 2 |
+
|
| 3 |
+
from copy import deepcopy
|
| 4 |
+
import math
|
| 5 |
+
from types import SimpleNamespace
|
| 6 |
+
import unittest
|
| 7 |
+
from unittest.mock import Mock, patch
|
| 8 |
+
|
| 9 |
+
from fastapi.testclient import TestClient
|
| 10 |
+
|
| 11 |
+
import app as server
|
| 12 |
+
from catalog import TASKS
|
| 13 |
+
from hallucination import evidence_prompt, hallucination_result, validate_hallucination_model
|
| 14 |
+
from hallucination_catalog import HALLUCINATION_EXAMPLES
|
| 15 |
+
from inference import InferenceEngine, InputLimitError
|
| 16 |
+
from job_queue import InferenceQueue
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
def public_request(index=0):
|
| 20 |
+
example = HALLUCINATION_EXAMPLES[index]
|
| 21 |
+
return {"task": "hallucination", **{key: example[key] for key in ("text", "context", "answer")}}
|
| 22 |
+
|
| 23 |
+
|
| 24 |
+
class HallucinationSpanTests(unittest.TestCase):
|
| 25 |
+
def test_only_answer_tokens_contribute_and_unicode_offsets_preserve_original(self):
|
| 26 |
+
answer = "😀 李雷在上海。"
|
| 27 |
+
result = hallucination_result(answer,
|
| 28 |
+
[None, 0, None, 1, 1, 1, 1, 1, 1, None],
|
| 29 |
+
[[0, 0], [0, 100], [0, 0], [0, 1], [2, 3], [3, 4], [4, 5], [5, 7], [7, 8], [0, 0]],
|
| 30 |
+
[1, .99, 1, .1, .8, .9, .5, .7, .2, 1])
|
| 31 |
+
self.assertEqual(result["kind"], "hallucination")
|
| 32 |
+
self.assertEqual(result["answer"], answer)
|
| 33 |
+
self.assertEqual(result["spans"], [
|
| 34 |
+
{"start": 2, "end": 4, "text": "李雷", "score": .9},
|
| 35 |
+
{"start": 5, "end": 7, "text": "上海", "score": .7},
|
| 36 |
+
])
|
| 37 |
+
self.assertEqual(result["max_hallucination_score"], .9)
|
| 38 |
+
|
| 39 |
+
def test_strict_threshold_and_supported_token_separate_spans(self):
|
| 40 |
+
result = hallucination_result("ABC", [1, 1, 1], [[0, 1], [1, 2], [2, 3]], [.8, .5, .9])
|
| 41 |
+
self.assertEqual([span["text"] for span in result["spans"]], ["A", "C"])
|
| 42 |
+
self.assertEqual(hallucination_result("A", [1], [[0, 1]], [.5])["spans"], [])
|
| 43 |
+
|
| 44 |
+
def test_adjacent_positive_tokens_keep_internal_space_and_byte_fallback(self):
|
| 45 |
+
result = hallucination_result("𐍈 B", [1] * 5,
|
| 46 |
+
[[0, 1]] * 4 + [[2, 3]], [.7, .8, .9, .8, .6])
|
| 47 |
+
self.assertEqual(result["spans"], [{"start": 0, "end": 3, "text": "𐍈 B", "score": .9}])
|
| 48 |
+
|
| 49 |
+
def test_alternating_byte_fallback_scores_union_only_overlapping_characters(self):
|
| 50 |
+
result = hallucination_result("𐍈", [1] * 4, [[0, 1]] * 4, [.8, .1, .9, .5])
|
| 51 |
+
self.assertEqual(result["spans"], [{"start": 0, "end": 1, "text": "𐍈", "score": .9}])
|
| 52 |
+
# A supported byte token closes its span. A later positive token over
|
| 53 |
+
# another character remains distinct, even when the ranges just touch.
|
| 54 |
+
separate = hallucination_result("𐍈B", [1] * 3, [[0, 1], [0, 1], [1, 2]], [.8, .5, .9])
|
| 55 |
+
self.assertEqual(separate["spans"], [
|
| 56 |
+
{"start": 0, "end": 1, "text": "𐍈", "score": .8},
|
| 57 |
+
{"start": 1, "end": 2, "text": "B", "score": .9},
|
| 58 |
+
])
|
| 59 |
+
|
| 60 |
+
def test_invalid_scores_offsets_or_empty_answer_mapping_fail_closed(self):
|
| 61 |
+
cases = [([1], [[0, 2]], [.7]), ([1], [[-1, 1]], [.7]),
|
| 62 |
+
([1], [[0, 1]], [math.nan]), ([1], [[0, 1]], [1.1]),
|
| 63 |
+
([0], [[0, 1]], [.9]), ([1, 1], [[0, 1]], [.7])]
|
| 64 |
+
for sequences, offsets, scores in cases:
|
| 65 |
+
with self.subTest(case=(sequences, offsets, scores)), self.assertRaises(ValueError):
|
| 66 |
+
hallucination_result("A", sequences, offsets, scores)
|
| 67 |
+
|
| 68 |
+
|
| 69 |
+
class HallucinationAPITests(unittest.TestCase):
|
| 70 |
+
def setUp(self):
|
| 71 |
+
self.engine_patch = patch.object(server, "engine", InferenceEngine())
|
| 72 |
+
self.engine_patch.start()
|
| 73 |
+
self.addCleanup(self.engine_patch.stop)
|
| 74 |
+
self.queue = InferenceQueue(server.run_inference, tasks=TASKS)
|
| 75 |
+
self.queue_patch = patch.object(server, "jobs", self.queue)
|
| 76 |
+
self.queue_patch.start()
|
| 77 |
+
self.addCleanup(self.queue_patch.stop)
|
| 78 |
+
self.client = TestClient(server.app).__enter__()
|
| 79 |
+
self.addCleanup(self.client.__exit__, None, None, None)
|
| 80 |
+
|
| 81 |
+
def test_all_three_original_fields_reach_both_endpoints_unchanged(self):
|
| 82 |
+
payload = {"task": "hallucination", "text": " Question?\n", "context": " Evidence. ",
|
| 83 |
+
"answer": " 😀 Answer.\n"}
|
| 84 |
+
with patch.object(server.engine, "analyze_locked", return_value={"result": {"kind": "hallucination"}}) as analyze:
|
| 85 |
+
created = self.client.post("/api/jobs", json=payload)
|
| 86 |
+
self.assertEqual(created.status_code, 202)
|
| 87 |
+
self.assertEqual(self.queue.wait(created.json()["id"])["status"], "completed")
|
| 88 |
+
self.assertEqual(analyze.call_args.args, ("hallucination", payload["text"], []))
|
| 89 |
+
self.assertEqual(analyze.call_args.kwargs["context"], payload["context"])
|
| 90 |
+
self.assertEqual(analyze.call_args.kwargs["answer"], payload["answer"])
|
| 91 |
+
self.assertEqual(self.client.post("/api/analyze", json=payload).status_code, 200)
|
| 92 |
+
|
| 93 |
+
def test_required_fields_char_limits_and_unrelated_fields_reject_without_echo(self):
|
| 94 |
+
invalid = []
|
| 95 |
+
for field in ("text", "context", "answer"):
|
| 96 |
+
for value in ("", " \n", "x" * 16001, 123, None):
|
| 97 |
+
invalid.append(dict(public_request(), **{field: value}))
|
| 98 |
+
missing = public_request()
|
| 99 |
+
del missing[field]
|
| 100 |
+
invalid.append(missing)
|
| 101 |
+
invalid += [dict(public_request(), candidates=["passage"]),
|
| 102 |
+
dict(public_request(), categories=[{"id": "a", "name": "A"}]),
|
| 103 |
+
dict(public_request(), texts=["first", "second", "third"])]
|
| 104 |
+
with patch.object(server.engine, "analyze_locked") as analyze:
|
| 105 |
+
for payload in invalid:
|
| 106 |
+
for path in ("/api/jobs", "/api/analyze"):
|
| 107 |
+
response = self.client.post(path, json=payload)
|
| 108 |
+
self.assertEqual(response.status_code, 422)
|
| 109 |
+
self.assertNotIn("The museum opens", response.text)
|
| 110 |
+
analyze.assert_not_called()
|
| 111 |
+
|
| 112 |
+
def test_existing_tasks_reject_nonempty_evidence_answer_and_keep_empty_defaults(self):
|
| 113 |
+
with patch.object(server.engine, "analyze_locked", return_value={"result": {}}) as analyze:
|
| 114 |
+
for task, spec in TASKS.items():
|
| 115 |
+
if task == "hallucination":
|
| 116 |
+
continue
|
| 117 |
+
payload = {key: value for key, value in spec["examples"][0].items()
|
| 118 |
+
if key in {"text", "candidates", "categories", "texts"}}
|
| 119 |
+
payload.update(task=task, context="", answer="")
|
| 120 |
+
self.assertEqual(self.client.post("/api/analyze", json=payload).status_code, 200, task)
|
| 121 |
+
for field in ("context", "answer"):
|
| 122 |
+
count = analyze.call_count
|
| 123 |
+
response = self.client.post("/api/analyze", json=dict(payload, **{field: "private@example.com"}))
|
| 124 |
+
self.assertEqual(response.status_code, 422, task)
|
| 125 |
+
self.assertNotIn("private@example.com", response.text)
|
| 126 |
+
self.assertEqual(analyze.call_count, count)
|
| 127 |
+
|
| 128 |
+
def test_combined_token_error_is_a_422_and_releases_worker_slot(self):
|
| 129 |
+
message = "Question + evidence + answer has 513 tokens. This CPU demo accepts at most 512."
|
| 130 |
+
with patch.object(server.engine, "analyze_locked", side_effect=InputLimitError(message)):
|
| 131 |
+
response = self.client.post("/api/analyze", json=public_request())
|
| 132 |
+
self.assertEqual(response.status_code, 422)
|
| 133 |
+
self.assertEqual(response.json()["detail"], message)
|
| 134 |
+
self.assertFalse(server.engine.slot.locked())
|
| 135 |
+
|
| 136 |
+
def test_catalog_has_halu_pin_examples_and_demo_limits_without_loading(self):
|
| 137 |
+
catalog = self.client.get("/api/catalog").json()
|
| 138 |
+
spec = next(task for task in catalog["tasks"] if task["id"] == "hallucination")
|
| 139 |
+
self.assertEqual(spec["model"], "llm-semantic-router/Vela-1.0-Encoder-307M-Halu")
|
| 140 |
+
self.assertEqual(spec["revision"], "521cd05d15e1959e120d002663c3b775649ae4dd")
|
| 141 |
+
self.assertEqual(spec["kind"], "hallucination")
|
| 142 |
+
self.assertEqual(len(spec["examples"]), 5)
|
| 143 |
+
self.assertTrue(all({"name", "text", "context", "answer"} <= example.keys()
|
| 144 |
+
for example in spec["examples"]))
|
| 145 |
+
self.assertEqual(catalog["limits"]["max_tokens"], 512)
|
| 146 |
+
self.assertEqual(catalog["limits"]["max_hallucination_field_chars"], 16000)
|
| 147 |
+
self.assertIsNone(self.client.get("/api/status").json()["resident_task"])
|
| 148 |
+
|
| 149 |
+
|
| 150 |
+
class HallucinationInferenceTests(unittest.TestCase):
|
| 151 |
+
def test_512_token_pair_budget_preserves_exact_prompt_and_rejects_before_load(self):
|
| 152 |
+
observed = []
|
| 153 |
+
|
| 154 |
+
def tokenizer(text, text_pair=None, **kwargs):
|
| 155 |
+
observed.append((text, text_pair, kwargs))
|
| 156 |
+
return {"input_ids": [0] * (len(text.split()) + len(text_pair.split()) + 3)}
|
| 157 |
+
|
| 158 |
+
engine = InferenceEngine()
|
| 159 |
+
question, context, answer = " Q? ", " evidence " * 250, " answer " * 256
|
| 160 |
+
engine._check_tokens(tokenizer, "hallucination", question, [], context=context, answer=answer)
|
| 161 |
+
self.assertEqual(observed[-1][0], f"User request: {question}\n\n{context}")
|
| 162 |
+
self.assertEqual(observed[-1][1], answer)
|
| 163 |
+
self.assertIs(observed[-1][2]["truncation"], False)
|
| 164 |
+
with patch.object(engine, "_get_tokenizer", return_value=tokenizer), patch.object(engine, "_ensure_model") as load:
|
| 165 |
+
with self.assertRaisesRegex(InputLimitError, r"Question \+ evidence \+ answer has 513 tokens.*at most 512"):
|
| 166 |
+
engine.analyze_locked("hallucination", question, [], context=context, answer=answer + " extra")
|
| 167 |
+
load.assert_not_called()
|
| 168 |
+
|
| 169 |
+
def test_halu_loads_native_token_classification_cpu_contract_and_reuses_lru(self):
|
| 170 |
+
model = SimpleNamespace(eval=Mock())
|
| 171 |
+
token_loader = SimpleNamespace(from_pretrained=Mock(return_value=model))
|
| 172 |
+
wrong_loader = SimpleNamespace(from_pretrained=Mock(side_effect=AssertionError("wrong model class")))
|
| 173 |
+
transformers = SimpleNamespace(AutoModel=wrong_loader, AutoModelForSequenceClassification=wrong_loader,
|
| 174 |
+
AutoModelForTokenClassification=token_loader)
|
| 175 |
+
torch = SimpleNamespace(float32="float32", set_num_threads=lambda count: None)
|
| 176 |
+
engine, tokenizer = InferenceEngine(), object()
|
| 177 |
+
with patch.dict("sys.modules", {"torch": torch, "transformers": transformers}), \
|
| 178 |
+
patch("inference.validate_hallucination_model") as validate:
|
| 179 |
+
engine._ensure_model("hallucination", tokenizer)
|
| 180 |
+
engine._ensure_model("hallucination", tokenizer)
|
| 181 |
+
validate.assert_called_once_with(model)
|
| 182 |
+
token_loader.from_pretrained.assert_called_once_with(TASKS["hallucination"]["model"],
|
| 183 |
+
revision=TASKS["hallucination"]["revision"], torch_dtype="float32",
|
| 184 |
+
attn_implementation="sdpa", trust_remote_code=False, reference_compile=False)
|
| 185 |
+
self.assertEqual(engine.status()["cached_tasks"], ["hallucination"])
|
| 186 |
+
self.assertEqual(engine.model_cache_size, 2)
|
| 187 |
+
|
| 188 |
+
def test_wrong_labels_fail_before_a_model_can_enter_the_cache(self):
|
| 189 |
+
config = SimpleNamespace(model_type="modernbert", id2label={0: "hallucinated", 1: "supported"},
|
| 190 |
+
label2id={"supported": 1, "hallucinated": 0}, num_labels=2,
|
| 191 |
+
reference_compile=False, _attn_implementation="sdpa")
|
| 192 |
+
model = type("ModernBertForTokenClassification", (), {"config": config})()
|
| 193 |
+
with self.assertRaisesRegex(ValueError, "native inference contract"):
|
| 194 |
+
validate_hallucination_model(model)
|
| 195 |
+
|
| 196 |
+
|
| 197 |
+
class HallucinationCacheTests(unittest.TestCase):
|
| 198 |
+
def test_only_full_exact_public_question_evidence_answer_can_hit_cache(self):
|
| 199 |
+
spec = deepcopy(TASKS["hallucination"])
|
| 200 |
+
calls = []
|
| 201 |
+
|
| 202 |
+
def runner(payload, stage):
|
| 203 |
+
calls.append(deepcopy(payload))
|
| 204 |
+
return {"task": "hallucination", "model": spec["model"], "revision": spec["revision"],
|
| 205 |
+
"result": {"kind": "hallucination", "answer": payload["answer"], "spans": []}}
|
| 206 |
+
|
| 207 |
+
now = [100.0]
|
| 208 |
+
queue = InferenceQueue(runner, tasks={"hallucination": spec}, clock=lambda: now[0])
|
| 209 |
+
queue.start()
|
| 210 |
+
self.addCleanup(queue.stop)
|
| 211 |
+
|
| 212 |
+
def run(payload):
|
| 213 |
+
created = queue.submit(payload)
|
| 214 |
+
self.assertTrue(queue._jobs[created["id"]].done.wait(2))
|
| 215 |
+
return queue.get(created["id"])
|
| 216 |
+
|
| 217 |
+
self.assertFalse(run(public_request())["result"]["cache_hit"])
|
| 218 |
+
self.assertTrue(run(public_request())["result"]["cache_hit"])
|
| 219 |
+
# Same question/evidence with a different public answer is a separate key.
|
| 220 |
+
self.assertFalse(run(public_request(1))["result"]["cache_hit"])
|
| 221 |
+
self.assertTrue(run(public_request(1))["result"]["cache_hit"])
|
| 222 |
+
self.assertEqual(len(calls), 2)
|
| 223 |
+
for field in ("text", "context", "answer"):
|
| 224 |
+
custom = dict(public_request(), **{field: public_request()[field] + " private@example.com"})
|
| 225 |
+
for _ in range(2):
|
| 226 |
+
self.assertFalse(run(custom)["result"]["cache_hit"])
|
| 227 |
+
self.assertEqual(len(calls), 8)
|
| 228 |
+
self.assertEqual(queue.status()["example_cache_entries"], 2)
|
| 229 |
+
self.assertNotIn("private@example.com", str(queue._examples))
|
| 230 |
+
self.assertTrue(all(job.payload is None for job in queue._jobs.values()))
|
| 231 |
+
now[0] += 301
|
| 232 |
+
self.assertTrue(run(public_request())["result"]["cache_hit"])
|
| 233 |
+
spec["revision"] = "new-pinned-revision"
|
| 234 |
+
self.assertFalse(run(public_request())["result"]["cache_hit"])
|
| 235 |
+
|
| 236 |
+
|
| 237 |
+
if __name__ == "__main__":
|
| 238 |
+
unittest.main()
|