FIELD NOTES / SEPTEMBER 2026

Laya Email Triage: A Reproducible Test Kit

Evaluate whether a decision model can send an incoming email to the right queue. This kit contains synthetic examples, expected labels and a script for recording actual model responses.

Local inference completed on eight synthetic English examples. Expected labels were written before the run. Results and limitations are reported below.

Define the queues first

Download the kit

Download test script Download 8 example cases

Observed local run: September 29, 2026

All 8 synthetic examples matched their prewritten labels in one run. This small smoke test checks that the pipeline works; it does not estimate production accuracy.

Windows · Intel Core i5-1335U · CPU inference · Python 3.12.14 · Laya 0.3.20 · PyTorch 2.14.0+cpu. Model revision: 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851.

CaseExpectedObservedSeconds
1billingbilling59.130
2technicaltechnical0.289
3salessales0.269
4otherother0.261
5billingbilling0.262
6technicaltechnical0.268
7salessales0.258
8otherother0.241

Case 1 includes first-use downloading/loading and is not an inference-speed measurement. The remaining seven calls took 0.241–0.289 seconds each in this run. We did not run repeated trials or compare devices.

The SDK emitted a temperature-clamping warning for its choice:11+ entry. This experiment used four choices; we did not assess probability calibration and do not recommend a confidence threshold from these results.

Download raw model results · Download dependency versions

Score the output fairly

  1. Read each example and confirm its expected label before running the model.
  2. Use the same checkpoint and question definition for every example.
  3. Inspect raw decisions, including failures. The script stops if the SDK call fails.
  4. Track billing mistakes separately: a high overall score can hide errors in a small queue.
  5. Add a separate held-out set of your own de-identified examples before evaluating a real workflow.

Handle uncertainty outside the model

Messages that combine a billing problem and a technical problem may need more than one queue. Decide the routing policy before evaluating them. A sensible first deployment suggests a queue to a reviewer and does not send replies or issue refunds automatically.

What this experiment cannot establish

Eight hand-written cases are a smoke test, not a benchmark. They cannot prove multilingual quality, production reliability or an optimal confidence threshold. Measure those on representative held-out data.

Sources: developer model card and developer repository. Source-based guidance; this site has not completed an independent model benchmark.