This page documents known limitations, current issues, and expected behaviors of the L&S AI Inference platform during early access. Check back as the service evolves.
Last updated: September 18, 2026
| Issue | Status | Workaround |
|---|---|---|
| Occasional high latency during peak usage | Ongoing | Try off-peak hours; use gemma-4-26b-a4b-it for faster responses |
| Web search may return no results for niche queries | Ongoing | Disable web search and rely on model knowledge; rephrase the query |
| Very long responses may be cut off near token limits | Ongoing | Ask for shorter responses or request output in sections |
Issues will be added and resolved here as the service evolves.
The models on this service are open-weight large language models hosted by UCSB Letters & Science IT (CIT). They include safety behavior built in by their developers, and we test them before deployment. However, no safeguard is complete. Models can still produce inaccurate, biased, offensive, or harmful output, and their behavior may change as models are updated or replaced. This service is provided as-is.
You are responsible for how you use this service and anything it generates. Verify output before relying on it. Do not use it as a substitute for medical, legal, or financial advice. Do not use it to produce content or take actions that would violate law or University policy. As the UC AI Council states, if an action wasn't permissible before AI, it isn't permissible with AI.
Use of this service is governed by the UC Electronic Communications Policy, the UC Electronic Information Security Policy (IS-3), the UC Responsible AI Principles, and UCSB's Guidelines for AI Use.
To report problematic output, contact help@cit.ucsb.edu.
Context limits vary by model: gemma-4-26b-a4b-it and qwen3.8-27b support up to 256k tokens, and gpt-oss-120b supports up to 128k tokens. Longer contexts come with tradeoffs:
See the Context Windows Guide for strategies.
If you encounter a problem not listed here, please contact help@cit.ucsb.edu with: