What are models in CREAO?
CREAO gives you access to multiple AI models from different providers, so you can choose the one that best fits the task in front of you, whether you need maximum intelligence, the largest context window, or the most cost-efficient option. The model is the reasoning engine behind every chat and agent.
Crucially, the choice is never a trade-off in capability. All models have full access to the same tools: code execution, web search, image generation, file handling, and every connected skill and integration. Swapping models changes how the agent thinks and what it costs, not what it can do.
Model comparison
CREAO offers frontier models across six providers. Every model shares the same toolset; they differ in reasoning depth, context window, caching support, and cost tier.
Model
Provider
Context
Cache
Tier
Best for
Claude Opus 4.8
Anthropic
1M
Yes
Premium
Frontier reasoning and coding with a 1M-token context window
Claude Opus 4.7
Anthropic
1M
Yes
Premium
Hardest software engineering and long-horizon agent work
Claude Opus 4.6
Anthropic
1M
Yes
Premium
Complex reasoning at the same price as Opus 4.7
Fugu Ultra
Sakana AI
1M
Yes
Premium
Multi-agent reasoning for hard, long-horizon work
Claude Sonnet 4.6
Anthropic
1M
Yes
Standard
Best balance of speed and intelligence
Claude Haiku 4.5
Anthropic
200K
Yes
Economy
Fast, cost-efficient tasks
Gemini 3.1 Pro
Google
1M
Yes
Standard
Advanced reasoning with multimodal input
Gemini 3.5 Flash
Google
1M
Yes
Economy
Fast, agentic Gemini for high-throughput workflows
GPT-5.5
OpenAI
1M
Yes
Standard
Smartest GPT model for frontier agentic coding
GLM 5.2
Z.ai
1M
No
Economy
Strong coding at very low cost, great for iterative sessions
MiniMax M3
MiniMax
1M
Yes
Economy
Sparse-attention reasoning, fast at very long context
Fugu Ultra delivers output in larger batches rather than token-by-token. It thinks first, then sends a complete response. The chat header shows a notice when Fugu Ultra is selected.
Choosing a model
There is no single right answer: the best model depends on what you are optimizing for. Use this as a quick guide.
Best quality, cost no object
Claude Opus 4.8 for frontier reasoning and the largest context window. Fugu Ultra (Sakana) handles hard, long-horizon multi-agent work on paid plans. Opus 4.7 and 4.6 remain available for tuned prompts.
Balanced speed, quality, and cost
Claude Sonnet 4.6 is the default and recommended for most users. It handles coding, analysis, and creative tasks well at moderate cost.
Minimize credit usage
GLM 5.2 and MiniMax M3 use very few credits per message with strong coding and 1M-token context. GLM 5.2 shines at reasoning; MiniMax M3 excels at very long context.
Fast responses for simple tasks
Claude Haiku 4.5 is the fastest model. Use it for quick questions, formatting, or lightweight code where speed matters more than depth.
Switching models
Select a model from the model dropdown at the top of the chat interface. Your choice persists per thread, so you can use different models for different conversations.
I
Open the model dropdown
Find it at the top of the chat interface, above the conversation.
II
Pick a model
Choose based on the intelligence, context window, or cost you need for the task.
III
Keep working
Your choice persists for that thread, so the model stays selected as you continue.
IV
Switch any time
Use different models in different threads: premium reasoning here, an economy model there.
Prompt caching
Most models support prompt caching, which reduces cost and latency on follow-up messages in the same thread. When caching is active, repeated parts of the conversation (the system prompt and earlier messages) are served from cache at a reduced rate.
Caching happens automatically. You don't need to configure anything.
Credit costs
Credits are deducted based on actual token usage. The cost tier determines how many credits each message uses.
Cost tier
Credits per message
Models
Economy
~0.1 to 1 credits
GLM 5.2, MiniMax M3, Claude Haiku 4.5, Gemini 3.5 Flash
Standard
~1 to 6 credits
Claude Sonnet 4.6, GPT-5.5, Gemini 3.1 Pro
Premium
~5 to 20 credits
Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Fugu Ultra
Exact costs depend on the length of your message, the conversation history, and how much the model outputs. You can see your remaining credits in the sidebar.
Use cases
Hardest engineering
Reach for Claude Opus 4.8 or 4.7 on long-horizon agent work and the toughest software problems.
Everyday building
Use Claude Sonnet 4.6 for coding, analysis, and creative work at moderate cost.
High-volume iteration
Switch to GLM 5.2 or MiniMax M3 to send many messages cheaply during long sessions.
Quick tasks
Pick Claude Haiku 4.5 or Gemini 3.5 Flash for fast, lightweight responses.
Best practices
Start with Sonnet 4.6. It is the default for a reason: strong across coding, analysis, and writing at moderate cost. Move up or down only when a task calls for it.
Reach for premium on the hard problems. Use Opus 4.8 or Fugu Ultra for long-horizon agent work and the toughest engineering, where deeper reasoning pays for itself.
Drop to economy for iteration. GLM 5.2 and MiniMax M3 keep credit usage low across the many messages of a long session, with 1M-token context.
Let caching work for you. Keep related work in the same thread so follow-up messages are served from cache at a reduced rate.
Frequently asked questions
How many models does CREAO offer?+
CREAO offers frontier models across six providers (Anthropic, Google, OpenAI, Sakana AI, Z.ai, and MiniMax), and new models are added regularly.
Which model should I use?+
Claude Sonnet 4.6 is the default and best for most work. Use Claude Opus 4.8 or Fugu Ultra for the hardest problems, GLM 5.2 or MiniMax M3 to save credits, and Claude Haiku 4.5 for the fastest responses.
Do all models have the same tools?+
Yes. Every model has full access to the same tools: code execution, web search, image generation, file handling, and all connected skills and integrations.
How do I switch models?+
Select a model from the dropdown at the top of the chat interface. Your choice persists per thread, so you can use different models in different conversations.
What is prompt caching?+
Most models support prompt caching, which serves repeated parts of a conversation (the system prompt and earlier messages) from cache at a reduced rate, cutting cost and latency on follow-ups. It happens automatically.
How are credits charged?+
Credits are deducted based on actual token usage. The cost tier sets the rate, and the exact amount depends on your message length, conversation history, and how much the model outputs.
What is Fugu Ultra?+
Fugu Ultra (Sakana AI) is a multi-agent reasoning model with a 1M-token context window, available on paid plans. It thinks first and delivers output in larger batches rather than token-by-token, so the chat header shows a notice when it is selected.
Pick the right model for the job
Frontier reasoning when it matters, economy models when it doesn't, same tools every time.