Test in Playground
Test an agent’s saved configuration with temporary model and prompt changes, then save the change that works.
Open Playground from the agent’s sidebar. Playground tests the agent’s saved knowledge in an authenticated workspace session. The normal view starts with the agent’s saved model and system prompt. Playground is the safest place to test a new configuration before you use it in a public widget or channel.
Test without saving
Choose a Model, or edit the System Prompt. Then select Apply & Test. This remounts the test chat with the current changes. It does not save the model or the prompt. The Unsaved changes badge shows that the editor differs from the saved agent. If the candidate configuration passes your acceptance criteria, select Save to Agent. Select Reset to restore the last saved model and prompt.
If you plan to save the prompt, keep it within 8,192 characters. The chat endpoint accepts a longer test prompt. The saved System Instruction has the lower limit. When retrieval produces sources, a normal reply can show expandable source chips after streaming. These chips are useful evidence. They do not appear for every answer. Do not treat them as guaranteed citations.
Compare models
Select Compare Models to begin with two panels: the saved primary model and Claude Sonnet 4.6. You can use two to four panels. A shared input sends the same message to every panel that is not paused. It shows the total credit cost for the message first.
Each panel can use a different model, temperature, and prompt. The browser saves each panel’s configuration locally, for that agent only. The server does not store this configuration. After a reply, the panel shows timing, token, and credit statistics. Before you save a panel, confirm the change. Saving overwrites the agent’s model and system prompt. The temperature setting stays in Playground only. It does not save.
All comparison panels share the agent’s saved sources and retrieval configuration. When you need source chips to diagnose retrieval, use normal mode. Use the response-quality guide to grade results consistently.