FIELD NOTES / SEPTEMBER 2026
Laya Email Triage: A Reproducible Test Kit
Sources checked
Evaluate whether a decision model can send an incoming email to the right queue. This kit contains synthetic examples, expected labels and a script for recording actual model responses.
Define the queues first
- billing: invoices, charges and refunds.
- technical: application errors and failed functionality.
- sales: pre-purchase questions and quotes.
- other: everything else or insufficient information.
Download the kit
Download test script Download 8 example cases
Observed local run: September 29, 2026
All 8 synthetic examples matched their prewritten labels in one run. This small smoke test checks that the pipeline works; it does not estimate production accuracy.
Windows · Intel Core i5-1335U · CPU inference · Python 3.12.14 · Laya 0.3.20 · PyTorch 2.14.0+cpu. Model revision: 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851.
| Case | Expected | Observed | Seconds |
|---|---|---|---|
| 1 | billing | billing | 59.130 |
| 2 | technical | technical | 0.289 |
| 3 | sales | sales | 0.269 |
| 4 | other | other | 0.261 |
| 5 | billing | billing | 0.262 |
| 6 | technical | technical | 0.268 |
| 7 | sales | sales | 0.258 |
| 8 | other | other | 0.241 |
Case 1 includes first-use downloading/loading and is not an inference-speed measurement. The remaining seven calls took 0.241–0.289 seconds each in this run. We did not run repeated trials or compare devices.
The SDK emitted a temperature-clamping warning for its choice:11+ entry. This experiment used four choices; we did not assess probability calibration and do not recommend a confidence threshold from these results.
Download raw model results · Download dependency versions
Score the output fairly
- Read each example and confirm its expected label before running the model.
- Use the same checkpoint and question definition for every example.
- Inspect raw decisions, including failures. The script stops if the SDK call fails.
- Track billing mistakes separately: a high overall score can hide errors in a small queue.
- Add a separate held-out set of your own de-identified examples before evaluating a real workflow.
Handle uncertainty outside the model
Messages that combine a billing problem and a technical problem may need more than one queue. Decide the routing policy before evaluating them. A sensible first deployment suggests a queue to a reviewer and does not send replies or issue refunds automatically.
What this experiment cannot establish
Eight hand-written cases are a smoke test, not a benchmark. They cannot prove multilingual quality, production reliability or an optimal confidence threshold. Measure those on representative held-out data.
Sources: developer model card and developer repository. Source-based guidance; this site has not completed an independent model benchmark.