工具概述
在AI Agent系统中,大量"系统一"决策任务——意图分类、工具路由、优先级排序、合规检查——并不真正需要大语言模型的生成能力。这些任务本质上是多分类或二分类问题,用传统LLM处理既昂贵又缓慢。Ollaya正是为此而生:它是一个开源(Apache-2.0协议)的本地运行时,专为Jev风格的决策模型设计,以单前向传播替代逐token生成,在消费级GPU上实现毫秒级响应。
与Ollama服务Llama、Qwen等生成式模型不同,Ollaya服务的模型不生成文本,而是直接输出带置信度概率的结构化答案。一个典型的五问题请求,在RTX 4090上端到端只需约8.9毫秒,相比TypeSafe官方托管API的236-276毫秒,性能提升超过26倍。更关键的是,数据完全保留在本地,对于客服工单、用户消息等敏感信息,这是不可替代的优势。
环境准备
系统要求:
- 操作系统:Windows 10/11、macOS 12+、Linux(Ubuntu 20.04+)
- 内存:最低8GB,推荐16GB
- GPU(可选但强烈推荐):NVIDIA显卡,需驱动R580+,CUDA支持;Apple Silicon和Intel/AMD CPU也可运行,仅速度较慢
- 网络:首次启动需下载模型权重(约200MB-1GB)
前置依赖:
- Python 3.10+(用于SDK集成)或任意HTTP客户端(用于API调用)
- Docker(服务器部署可选)
安装部署
Windows
从GitHub Releases下载最新安装包,安装完成后运行:
# 命令行安装(推荐) ollaya run laya # 或启动桌面应用(双击图标) ollaya-desktop
安装程序会自动检测NVIDIA GPU并安装CUDA库(如有),无GPU时自动使用CPU模式。
macOS
# Homebrew安装 brew install ollaya/tap/ollaya # 或使用直接下载 curl -L -o /usr/local/bin/ollaya \ "https://github.com/convai-labs/ollaya/releases/latest/download/ollaya-macos-arm64" chmod +x /usr/local/bin/ollaya # 运行 ollaya run laya
Linux(Docker方式)
# 拉取镜像 docker pull ghcr.io/convai-labs/ollaya:latest # 启动服务(CPU模式) docker run -p 11435:11435 ghcr.io/convai-labs/ollaya:latest # 启动服务(GPU模式,需nvidia-container-toolkit) docker run --gpus all -p 11435:11435 ghcr.io/convai-labs/ollaya:latest
Linux(原生安装)
curl -L -o /tmp/ollaya-linux-amd64.tar.gz \ "https://github.com/convai-labs/ollaya/releases/latest/download/ollaya-linux-amd64.tar.gz" tar -xzf /tmp/ollaya-linux-amd64.tar.gz -C /usr/local/bin/ ollaya run laya
快速上手
API兼容:TypeSafe SDK无缝对接
Ollaya完全兼容TypeSafe的API格式,官方Python SDK无需修改即可切换至本地服务:
import os
from typesafe import TypeSafeClient
# 指向本地Ollaya服务
os.environ["TYPESAFE_BASE_URL"] = "http://localhost:11435"
os.environ["TYPESAFE_API_KEY"] = "local" # 任意值即可
os.environ["TYPESAFE_DEFAULT_MODEL"] = "laya"
client = TypeSafeClient()
response = client.systemone.create(
model="laya",
state="Can I get an invoice for last month's order?",
questions={
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": {
"invoice": "Needs an invoice or receipt",
"refund": "Wants money back",
"other": "Anything else"
}
}
}
)
print(response.answers["intent"]["choice"]) # → "invoice"
print(response.answers["intent"]["confidence"]) # → 0.9547
print(response.answers["intent"]["probabilities"])
# → {"invoice": 0.9698, "refund": 0.0172, "other": 0.013}直接HTTP调用
curl -X POST http://localhost:11435/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"model": "laya:en",
"state": "User reports login failed with error code 5001",
"questions": {
"category": {
"type": "choice",
"criteria": {
"auth_error": "Authentication or authorization issue",
"server_error": "Backend system malfunction",
"network_error": "Connection or timeout problem"
}
},
"severity": {
"type": "choice",
"criteria": {
"critical": "System-wide outage",
"high": "Multiple users affected",
"medium": "Single user impact",
"low": "Non-blocking issue"
}
}
}
}'批量多问题推理
# 单次请求同时回答多个问题,批量效率极高
response = client.systemone.create(
model="laya",
state=context_text,
questions={
"intent": {"type": "choice", "criteria": {...}},
"sentiment": {"type": "choice", "criteria": {...}},
"priority": {"type": "rank", "criteria": {...}},
"routing": {"type": "choice", "criteria": {...}},
"escalate": {"type": "boolean"}
}
)
# 五问端到端延迟约 8-10ms(RTX 4090,fp16)进阶用法
模型路由:自动选最合适的模型
Ollaya内置路由器模型,可自动判断输入并选择最匹配的底层模型:
# 使用默认路由器(laya:router) ollaya run laya:router # 路由器会根据问题复杂度自动分流: # - 简单分类 → laya:en(最快) # - 多语言 → laya:multilingual # - 复杂逻辑 → 专用fine-tuned模型
Agent系统集成:Jev + LangGraph
Harrison Chase将Jev与LangGraph结合,用Jev处理Agent运行中的高频非生成决策,Opus只处理高复杂度推理:
from langgraph.graph import StateGraph, END
from typing import TypedDict
class AgentState(TypedDict):
task: str
needs_routing: bool
routing_decision: str
confidence: float
# Jev只负责:是否需要人工审批、下一步路由、任务是否完成
# 每个小决策 ~8ms,比调LLM快26倍
def decide_next_step(state: AgentState) -> AgentState:
resp = client.systemone.create(
model="laya",
state=state["task"],
questions={"next_action": {"type": "choice", "criteria": {...}}}
)
state["routing_decision"] = resp.answers["next_action"]["choice"]
state["confidence"] = resp.answers["next_action"]["confidence"]
return state成本对比实测
| 方案 | 1000次决策成本 | 平均延迟 |
|---|---|---|
| TypeSafe托管API | $42.00 | 236-276ms |
| Ollaya本地(RTX 4090) | $0.00(电费忽略) | 8.9ms |
| Ollaya本地(批量13问/请求) | $0.00 | ~0.7ms/问 |
批量请求可将单问成本进一步降至约0.7毫秒,适合高频Agent循环。
常见问题与排查
问题1:GPU未识别,退回CPU模式
现象:启动后日志显示"Running on CPU",延迟从8ms变为200ms+。
排查:
# 检查NVIDIA驱动版本 nvidia-smi --query-gpu=driver_version --format=csv,noheader # 确保驱动 >= 580.x # Windows: 控制面板 → NVIDIA控制面板 → 帮助 → 系统信息 # Linux: apt install nvidia-driver-580
问题2:CUDA库下载失败
现象:Windows安装时报"Failed to download CUDA libraries"。
解决:手动下载CUDA Runtime并设置环境变量:
# 设置CUDA路径 export CUDA_HOME=/usr/local/cuda export PATH=$CUDA_HOME/bin:$PATH export LD_LIBRARY_PATH=$CUDA_HOME/lib64:$LD_LIBRARY_PATH
问题3:多问题请求返回空结果
现象:questions字段结构正确,但answers为空或所有字段缺失。
排查:确保每个question的type字段是合法的枚举值(choice/rank/boolean),criteria的值只能是字符串键,不能是嵌套对象:
# ✅ 正确
"criteria": {"invoice": " Needs invoice", "refund": "Wants refund"}
# ❌ 错误
"criteria": {"invoice": {"label": "Needs invoice", "weight": 1.0}}问题4:Apple Silicon上延迟异常偏高
现象:Mac M系列芯片上laya模型延迟30-50ms,远超预期。
原因:默认fp32精度。可通过指定fp16优化:
# 启动时启用fp16 ollaya run laya --precision fp16 # 或通过环境变量 export OLLAYA_PRECISION=fp16 ollaya run laya
问题5:SDK版本不兼容
现象:使用TypeSafe Python SDK时报"Endpoint not found"。
解决:确保SDK版本 >= 0.7.1:
pip install --upgrade typesafe-sdk
总结与选型建议
Ollaya填补了AI Agent系统中一个长期被忽视的性能空白:高频、低延迟、低成本的结构化决策。它的核心定位不是替代大语言模型,而是承担Agent运行中那些"不需要生成、只需要判断"的任务——意图分类、工具路由、优先级排序、合规检查。
选型建议:
- 如果你的Agent每秒需要执行10次以上的分类/路由决策,Ollaya可以显著降低延迟和成本
- 如果数据涉及隐私(客服工单、用户消息),本地部署是唯一合规方案
- 如果已有TypeSafe托管API依赖,Ollaya提供零代码改造的迁移路径
- 如果硬件受限(无GPU),CPU模式仍可接受(~50ms/问),适合低频场景
与Ollama的组合使用也越来越流行:Ollama服务生成式LLM(System Two),Ollaya服务决策模型(System One),两者通过统一HTTP接口对外暴露,形成完整的本地AI推理栈。