🤖 ตั้งค่า LLM Models

LiteLLM Aggregator · จัดการ Open-Source + Commercial LLM ทั้งหมด

🦙 Open Source LLMs (ในระบบ Sovereign)

🦙

Llama 3.3 70B Instruct

Meta AI · Open Source · 128K context

⭐ Default 🇹🇭 Thai-tuned GPU x4
Endpoint:https://llm.nbtc.local:8001/v1
Quantization:FP8 (vLLM)
GPU:NVIDIA H100 × 4 · 67% utilized
Latency p50:1.2s
Throughput:847 tok/s
🌪️

Typhoon-2 70B (Thai)

SCB10X · Open Source · 32K context

🇹🇭 Native Thai Active
Endpoint:https://llm.nbtc.local:8002/v1
ใช้สำหรับ:เอกสารภาษาไทย เฉพาะทาง
Latency p50:1.4s

🌐 Commercial LLMs (External — ผ่าน LiteLLM Aggregator)

🧬

GPT-4o

OpenAI · 128K · $5/1M tok

Quota:234K / 250K (94%)
ใช้เดือนนี้:$1,170
🌟

Claude 3.7 Sonnet

Anthropic · 200K · $3/1M tok

Quota:156K / 200K (78%)
ใช้เดือนนี้:$468
✨

Gemini 2.0 Pro

Google · 1M ctx · $1.25/1M tok

Quota:97K / 100K (97%)
ใช้เดือนนี้:$121

🚦 Routing Rules (LiteLLM)

ลำดับเงื่อนไขRoute ไปเหตุผลสถานะ
1 เอกสารระดับ Confidential 🦙 Llama 3.3 (in-house) Privacy — ห้ามส่งออกระบบ Active
2 คำถามภาษาไทย + เอกสาร TH 🌪️ Typhoon-2 คุณภาพภาษาไทยดีที่สุด Active
3 Context > 128K tokens ✨ Gemini 2.0 Pro รองรับ 1M tokens Active
4 Code/SQL Generation 🌟 Claude 3.7 เก่ง code สูงสุด Active
5 Default (อื่น ๆ) 🦙 Llama 3.3 Fallback · ฟรี · ในประเทศ Active