Latency: stream tokens, show progress for multi-step work ('searching 3 sources…'), and use optimistic UI where possible. Uncertainty: show sources and citations, signal low confidence, and make 'I couldn't find that' a first-class answer.
Control: let users edit, regenerate, give feedback (thumbs plus a reason) and undo. Keep the human in charge of consequential actions.
Failure states: timeouts, refusals, rate limits and bad outputs all need designed responses, not raw errors. Feedback signals double as eval data.
Going deeper
Set expectations at the point of use (what the feature can and can't do), make errors recoverable (edit and regenerate), and collect feedback with a reason, not just thumbs. Google's People + AI Guidebook has worked patterns for each.