用 SGLang 扩展 JEV-like 决策模型

用 SGLang 扩展 JEV-like 决策模型

引言

当客户询问订单是否已发货时,智能体已掌握订单 ID,并面临三种可能行动:查询订单状态服务、搜索配送政策文档,或询问客户订单 ID。在执行前,它需要先做出选择。JEV-like 决策模型可直接返回类别或分数,而非长篇解释,应用代码据此快速行动。

决策 LLM 的本质

决策过程本质上是分类或评分任务。模型在答案边界位置通过下一个 token 的分数给出判断,无需生成解释文本。

两种提问方式

点式提示为每个候选单独评分,集合式提示则让模型一次性看到所有选项后再给出判断。

Pointwise prompts isolate each candidate; a setwise prompt places all options before one answer boundary.

为什么需要专用评分接口

SGLang 的 /v1/score API 允许显式请求特定标签 token 的分数,避免生成式 top-k 遗漏关键标签。同时支持单项评分(SIS)与多项评分(MIS),后者可复用共享查询计算。

Grouped bars compare p95 decision latency for 2, 5, 9, and 16 candidates on Qwen3-0.6B, Qwen3-8B, and Qwen3.5-4B.

基准测试:点式决策

使用 Open-Jev 数据集,在单张 NVIDIA H200 GPU 上测试 Qwen3-0.6B、Qwen3-8B 和 Qwen3.5-4B 模型。

Panels for Qwen3-0.6B, Qwen3-8B, and Qwen3.5-4B compare Generate, SIS, and MIS p95 latency as the target question rate increases, using a logarithmic latency axis.

结果显示,MIS 在高负载下优势明显,延迟增长缓慢。

The Score API explicitly returns requested labels, while MIS separately enables shared-query execution for independent pointwise candidates.

集合式与融合选择对比

进一步测试 Fused-Choice 与 Setwise SIS 在不同 QPS 下的表现。

Three panels compare Fused-Choice and Setwise SIS p95 latency versus offered QPS on Qwen3-0.6B, Qwen3-8B, and Qwen3.5-4B, with a shared logarithmic latency scale.Grouped bars show Fused-Choice and Setwise SIS p95 latency for 2, 5, 9, and 16 candidates on the three models, using a common zero-based 0 to 60 millisecond scale.

结论:MIS 能有效摊薄共享查询开销,适合高候选数量场景。