🤖 ตั้งค่า LLM Models
LiteLLM Aggregator · จัดการ Open-Source + Commercial LLM ทั้งหมด
🦙 Open Source LLMs (ในระบบ Sovereign)
Llama 3.3 70B Instruct
Meta AI · Open Source · 128K context
⭐ Default
🇹🇭 Thai-tuned
GPU x4
Endpoint:https://llm.nbtc.local:8001/v1
Quantization:FP8 (vLLM)
GPU:NVIDIA H100 × 4 · 67% utilized
Latency p50:1.2s
Throughput:847 tok/s
Typhoon-2 70B (Thai)
SCB10X · Open Source · 32K context
🇹🇭 Native Thai
Active
Endpoint:https://llm.nbtc.local:8002/v1
ใช้สำหรับ:เอกสารภาษาไทย เฉพาะทาง
Latency p50:1.4s
🌐 Commercial LLMs (External — ผ่าน LiteLLM Aggregator)
GPT-4o
OpenAI · 128K · $5/1M tok
Quota:234K / 250K (94%)
ใช้เดือนนี้:$1,170
Claude 3.7 Sonnet
Anthropic · 200K · $3/1M tok
Quota:156K / 200K (78%)
ใช้เดือนนี้:$468
Gemini 2.0 Pro
Google · 1M ctx · $1.25/1M tok
Quota:97K / 100K (97%)
ใช้เดือนนี้:$121
🚦 Routing Rules (LiteLLM)
| ลำดับ | เงื่อนไข | Route ไป | เหตุผล | สถานะ | |
|---|---|---|---|---|---|
| 1 | เอกสารระดับ Confidential | 🦙 Llama 3.3 (in-house) | Privacy — ห้ามส่งออกระบบ | Active | |
| 2 | คำถามภาษาไทย + เอกสาร TH | 🌪️ Typhoon-2 | คุณภาพภาษาไทยดีที่สุด | Active | |
| 3 | Context > 128K tokens | ✨ Gemini 2.0 Pro | รองรับ 1M tokens | Active | |
| 4 | Code/SQL Generation | 🌟 Claude 3.7 | เก่ง code สูงสุด | Active | |
| 5 | Default (อื่น ๆ) | 🦙 Llama 3.3 | Fallback · ฟรี · ในประเทศ | Active |