10 Commits
Author SHA1 Message Date
owen b4ff4f9f14 Prepare Cassady v0.3.5
CI / Build (push) Waiting to run
CI / Test (push) Waiting to run
2026-06-26 11:53:01 -05:00
owen ac86ee3933 Updating README 2026-06-25 22:03:46 -05:00
IrrelevantandGitHub da577aed1c Update ROADMAP with completion status for versions 2026-06-25 21:58:39 -05:00
owen ab0c45aff4 Correct roadmap: real v0.3.4 release notes, rename planned Tool Output Context Reliability to v0.3.5
The v0.3.4 tag is the released 'Reasoning Off Handling' release (commit
b35c4fb), distinct from the planned Tool Output Context Reliability work.
Add proper v0.3.4 release notes based on the shipped commit, move the Tool
Output Context Reliability entry to the top as v0.3.5, and rename its plan
file to match.
2026-06-25 21:56:44 -05:00
owen b35c4fb7e1 Send reasoning effort 'none' when reasoning is off
CI / Build (push) Waiting to run
CI / Test (push) Waiting to run
When a reasoning-capable model has reasoning effort set to off, send
'none' rather than omitting the field or sending 'off'. This applies
to both the ChatGPT Codex Responses API and OpenAI-compatible chat
completions (both reasoning_effort and reasoning object formats).
Non-reasoning models continue to send no reasoning field at all.

Bumps version to 0.3.4.
2026-06-25 20:08:15 -05:00
owen d05bb87a23 Prepare v0.3.3 release
CI / Build (push) Waiting to run
CI / Test (push) Waiting to run
2026-06-25 19:48:48 -05:00
owen 52eca6477e Bump version to 0.3.2
CI / Build (push) Waiting to run
CI / Test (push) Waiting to run
2026-06-25 19:22:56 -05:00
IrrelevantandGitHub 8b87e11e53 Merge pull request #13 from owenqwenstarsky/feat/fast-mode
Implement fast mode preference
2026-06-25 19:20:05 -05:00
owen 62e1d53f60 Implement fast mode preference 2026-06-25 19:17:44 -05:00
owen 39e0c14eec Plan tool output context reliability 2026-06-25 19:04:49 -05:00
29 changed files with 1659 additions and 83 deletions
Generated
+1 -1
View File
@@ -226,7 +226,7 @@ checksum = "8ae3f5d315924270530207e2a68396c3cc547f6dca3fbdca317cfb1a51edb593"
[[package]]
name = "cassady"
version = "0.3.1"
version = "0.3.5"
dependencies = [
"anyhow",
"async-trait",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "cassady"
version = "0.3.1"
version = "0.3.5"
edition = "2021"
description = "Cassady/Cass minimal terminal coding agent"
license = "MIT"
+9 -2
View File
@@ -8,7 +8,7 @@ The project installs two equivalent commands, `cass` and `cassady`; examples use
- Provider support includes OpenAI-compatible chat/completions APIs plus the `ChatGPT Codex` preset for users already signed in to Codex.
- The primary interface is an interactive terminal UI.
- v0.2.6 adds an experimental Rust embedding API for headless sessions; it is useful for early integrations but not yet a stable long-term library contract.
- An experimental Rust embedding API for headless sessions is available (added in v0.2.6); it is useful for early integrations but not yet a stable long-term library contract.
- Config and conversation state live under `~/.cass`.
- Windows binaries are built for releases, but deeper Windows terminal, path, shell, and filesystem polish is planned for a later release.
- `cass update` can update release-archive installs from official GitHub releases; external package managers should still be updated through their own tools.
@@ -80,10 +80,11 @@ Common in-chat commands:
- `/branch` or `/restore`: open the branch/restore menu.
- `/login`: configure or update provider login settings.
- `/logout`: remove saved provider config and associated model entries.
- `/fast`, `/fast on`, `/fast off`, `/fast status`: prefer faster inference when the active provider/model supports it. ChatGPT Codex models, including `gpt-5.5`, are treated as fast-capable.
- `/model <model>`: switch to a model from `~/.cass/models.json`.
- `/new`: create a new chat for the current directory.
- `/resume <chat>`: resume a saved chat for the current directory.
- `/status`: show chat id, model, mode, cwd, record count, and current status.
- `/status`: show chat id, model, fast-mode state, mode, cwd, record count, and current status.
Helpful keys:
@@ -107,6 +108,12 @@ Cassady exposes tools according to the active access mode:
Use `--readonly`, `--workspace-edit`, or `--full-access` to choose a mode at launch, or press `Shift-Tab` while idle.
## Tool output and context recovery
Cassady keeps tool calls reviewable while fitting provider context windows. Very large tool results may be sent to the model as compacted or truncated head/tail excerpts with a note that names the tool, retained excerpt shape, file ranges or command provenance when available, and suggested follow-up reads/searches. Treat those notes as incomplete evidence: ask Cass to re-read a narrower line range, run a focused `grep`, or rerun a narrower shell command before editing from omitted details.
Conversation files keep the recorded tool result content; model-facing compaction happens when preparing provider messages. The compact/full tool-output UI toggle only changes display.
## Branch and restore
Press `Esc` twice while idle, or type `/branch`, to browse the current conversation's branch family. Selecting an earlier user message, assistant message, tool call, or tool result creates a new branch conversation instead of truncating the original chat. The menu also lets you switch back to related branches later.
+65 -9
View File
@@ -1,6 +1,69 @@
# Cassady (Cass) Roadmap
## v0.3.2 — Provider Fast Mode
## v0.3.5 — Tool Output Context Reliability
This release focuses on making large tool outputs easier for the assistant to recover from when model-context compaction or truncation hides important details. Cassady should guide the assistant toward smaller, targeted reads and searches, preserve enough provenance for follow-up inspection, and add regression coverage for broad-output workflows that previously stalled safe edits. See `plans/V0_3_5_TOOL_OUTPUT_CONTEXT_RELIABILITY_PLAN.md`.
### Model Context Recovery
- [x] **Improve compacted tool-output guidance.** Replace generic head/tail compaction notices with actionable guidance that tells the assistant what was omitted and how to inspect it again safely.
- Include tool name, output size, retained excerpt shape, and suggested narrower follow-up reads or searches when available.
- Keep model-facing guidance concise enough that it does not worsen context pressure.
- [x] **Preserve targeted reinspection metadata.** Track enough structured context for large reads and command output so the assistant can recover omitted details without repeating broad requests.
- For file reads, preserve path and line-range coverage even after compaction.
- For shell and search output, prefer guidance toward narrower commands or `grep`/`read` follow-ups rather than blindly rerunning the same broad command.
### Tool Behavior and Prompting
- [x] **Bias tool use toward smaller inspections.** Update tool descriptions, prompt guidance, and result messages so broad reads become a fallback rather than the default.
- Encourage search-first workflows for large files and unknown locations.
- Mention result limits before or at truncation points so the assistant knows when context may be incomplete.
- [x] **Make truncation and compaction visible across layers.** Align model-facing messages, stored conversation records, and UI summaries so users and the assistant can tell when output was incomplete.
- Do not let UI-only collapsed output change what is stored or sent to the model.
- Keep existing conversation files readable and resumable.
### Validation
- [x] **Add regression coverage for broad-output recovery.** Test workflows where an early broad read or command output is compacted before the assistant needs exact context for an edit.
- Cover superseded reads, compacted non-newest tool outputs, provider-message validity, and suggested follow-up guidance.
- Verify `cargo fmt` and `cargo test --locked --all-targets` pass before handoff.
## v0.3.4 — Reasoning Off Handling ✅ Completed
This release focuses on sending the correct reasoning effort value when a reasoning-capable model has reasoning turned off. Cassady now sends `none` rather than omitting the field or sending `off` to both the ChatGPT Codex Responses API and OpenAI-compatible chat completions, while non-reasoning models continue to send no reasoning field at all.
### Reasoning Effort Off
- [x] **Send `none` when reasoning is off for supported models.** For reasoning-capable models with reasoning effort set to off, send `none` as the effort value instead of omitting the field or sending `off`.
- Apply to the ChatGPT Codex Responses API (`reasoning.effort`) and OpenAI-compatible chat completions in both `reasoning_effort` and `reasoning` object request formats.
- Keep non-reasoning models sending no reasoning field.
- [x] **Gate reasoning fields on model capability.** Track whether the active model supports reasoning from its metadata so reasoning request fields are only added for reasoning-capable models.
### Validation
- [x] **Add regression coverage for reasoning-off requests.** Test that supported models send `none` when reasoning is off across both provider kinds and both request formats, and that unsupported models send no reasoning field.
- Verify `cargo fmt` and `cargo test --locked --all-targets` pass before handoff.
## v0.3.3 — Codex Fast-Mode Compatibility ✅ Completed
This release focuses on keeping fast mode available for ChatGPT Codex users when local model metadata predates the fast-mode capability flag. Cassady should treat active `chatgpt-codex` provider models, including `gpt-5.5`, as fast-capable while leaving OpenAI-compatible and custom providers capability-gated by model metadata.
### Fast Mode Compatibility
- [x] **Treat ChatGPT Codex models as fast-capable at runtime.** Make `/fast` active for any active `chatgpt-codex` provider model even when legacy `models.json` metadata says `fast_mode.supported` is false.
- Keep non-Codex providers governed by their model metadata.
- Preserve the saved fast-mode preference behavior and status reporting.
- [x] **Document the Codex capability fallback.** Update README and bundled docs so users understand that ChatGPT Codex models, including `gpt-5.5`, can honor fast mode without refreshed metadata.
- Keep docs clear that provider-specific fast-mode request shaping remains Codex-only.
- [x] **Add regression coverage for legacy metadata.** Test that a ChatGPT Codex model with older `fast_mode.supported: false` metadata still reports fast mode as supported and active when preferred.
- Verify `cargo fmt` and `cargo test --locked --all-targets` pass before handoff.
## v0.3.2 — Provider Fast Mode ✅ Completed
This release focuses on adding a `/fast` command that lets users prefer faster inference when the active provider/model supports it. The first supported provider is `ChatGPT Codex`; other providers can add their own fast-mode request behavior later without changing the user-facing command. See `plans/V0_3_2_FAST_MODE_PLAN.md`.
@@ -32,7 +95,7 @@ This release focuses on adding a `/fast` command that lets users prefer faster i
- [ ] **Test fast-mode preference, switching, and provider requests.** Cover command parsing, persistence, status rendering, model/provider switching, and Codex request body behavior.
- Verify `cargo fmt` and `cargo test --locked --all-targets` pass before handoff.
## v0.3.1 — Transcript Scroll Stability
## v0.3.1 — Transcript Scroll Stability ✅ Completed
This release focuses on keeping the live transcript anchored correctly above the input and footer during long sessions with blank reasoning or tool-output lines.
@@ -451,13 +514,6 @@ This release focuses on making Cass easier to interrupt, easier to audit, and sa
These sections describe work Cassady intends to complete before or as part of the next major release, but which has not yet been assigned to a specific version. Scope, order, and version numbers may change.
### Tool Output Context Reliability
- [ ] **Reduce tool-output compaction stalls.** Make large file reads and command output easier to recover from when output is compacted or truncated, so the assistant can quickly switch to targeted inspection instead of getting stuck.
- Prefer smaller, focused file ranges and search-first workflows when large outputs are likely.
- Surface clearer guidance when tool results are compacted, including suggested narrower follow-up reads.
- Add regression coverage or dogfood checks for workflows where broad reads previously obscured the context needed for safe edits.
### Windows CLI Usability
This work focuses on making Cassady feel reliable and native when the CLI is run on Windows. It covers runtime usability after `cass` or `cassady` is already available on the machine; installers, package managers, PATH setup, code signing, and update delivery are intentionally out of scope.
+1 -1
View File
@@ -8,7 +8,7 @@ Cassady tools may list, search, and read this directory. Mutating tools are bloc
- [Commands](commands.md): CLI forms, global flags, `cass update`, in-chat commands, and keys.
- [Configuration](configuration.md): `~/.cass` files, setup, precedence, schema examples, and validation.
- [Providers and models](providers.md): built-in provider presets, custom OpenAI-compatible endpoints, ChatGPT Codex auth, model discovery, and reasoning metadata.
- [Providers and models](providers.md): built-in provider presets, custom OpenAI-compatible endpoints, ChatGPT Codex auth, model discovery, reasoning metadata, and fast-mode support.
- [Access modes and tool safety](access-modes.md): what tools can read, write, edit, and run in each mode.
- [Experimental Rust embedding API](embedding.md): import Cassady from Rust, start headless sessions, stream events, and handle approvals.
- [Workflows](workflows.md): common ways to inspect code, apply edits, run checks, switch models, and resume chats.
+2 -1
View File
@@ -107,12 +107,13 @@ The updater does not invoke `sudo` or administrator prompts. If the install dire
Type `/` to open command autocomplete.
- `/branch` or `/restore`: open the branch/restore menu for the current conversation family.
- `/fast`, `/fast on`, `/fast off`, `/fast status`: toggle or inspect a persisted fast-mode preference. Fast mode is active only when the current provider/model supports it; ChatGPT Codex models, including `gpt-5.5`, are treated as fast-capable.
- `/login`: configure or update provider login settings, then reload active provider/model config.
- `/logout`: remove saved providers and their associated models, then reload active provider/model config when any remain.
- `/model <model>`: switch the model for future turns. Autocomplete lists models from `~/.cass/models.json`.
- `/new`: create a new chat for the current directory.
- `/resume <chat>`: resume a saved chat from the current directory. Autocomplete lists matching chats.
- `/status`: show chat id, state, model, access mode, cwd, record count, and current status.
- `/status`: show chat id, state, model, fast-mode state, access mode, cwd, record count, and current status.
Local commands can be used only when the agent is idle.
+10 -1
View File
@@ -42,6 +42,7 @@ Example:
"default_provider": "openai",
"default_model": "gpt-4.1",
"default_reasoning_effort": "medium",
"default_fast_mode": false,
"default_access_mode": "read-only",
"context_message_limit": 80,
"model_tool_result_limit": 24000,
@@ -56,9 +57,10 @@ Fields:
- `default_provider`: optional provider id from `providers.json`. If omitted, Cassady infers the provider from `default_model` when possible.
- `default_model`: optional model id to use by default.
- `default_reasoning_effort`: optional `off`, `low`, `medium`, or `high`, clamped to model metadata.
- `default_fast_mode`: optional boolean, defaults to `false`. When `true`, Cassady requests faster inference only for provider/model combinations that advertise fast-mode support.
- `default_access_mode`: `"read-only"`, `"workspace-edit"`, or `"full-access"`.
- `context_message_limit`: optional legacy upper bound for recent non-system messages. Cassady primarily budgets context from model metadata and trims along valid tool-call boundaries.
- `model_tool_result_limit`: optional max bytes of tool output sent back to the model.
- `model_tool_result_limit`: optional approximate max characters of each tool output sent back to the model. Larger results are model-facing head/tail excerpts with recovery guidance; conversation records keep the tool result content.
- `ui_tool_result_limit`: optional max bytes of tool output shown in the UI unless full output is toggled.
- `show_reasoning`: optional boolean, defaults to `false`. Shows provider-streamed reasoning in the transcript.
- `confirm_destructive_operations`: optional compatibility preference currently stored in config.
@@ -136,6 +138,9 @@ Example:
"required": false,
"default_effort": "medium",
"request_format": "reasoning_effort"
},
"fast_mode": {
"supported": false
}
}
]
@@ -156,9 +161,13 @@ Fields:
- `required`: optional boolean, defaults to `false`.
- `default_effort`: optional `off`, `low`, `medium`, or `high`; defaults to `medium`. Cannot effectively be `off` when `required` is `true`.
- `request_format`: optional `reasoning_effort` or `reasoning_object`; defaults to `reasoning_effort`.
- `fast_mode`: optional object. Defaults to unsupported.
- `supported`: optional boolean, defaults to `false`. Cassady treats active `chatgpt-codex` provider models as fast-capable even if older metadata says otherwise; custom and OpenAI-compatible model entries default to unsupported.
Reasoning effort is a runtime per-turn setting. Press `Tab` to cycle it while idle. Provider-streamed reasoning is persisted and sent back in future model context using the provider's reasoning field, such as `reasoning_content` or `reasoning`.
Fast mode is a persisted preference, not a guarantee. Use `/fast` to toggle it while idle. `/status` shows `enabled` only when the preference is on and the current provider/model can honor it; otherwise it reports `off` or `preferred, unavailable ...`. Fast-mode request shaping is implemented only for `chatgpt-codex`.
## Precedence
- CLI access-mode flags override `default_access_mode` for the current session.
+7 -1
View File
@@ -12,11 +12,15 @@
**Config root**: The `~/.cass` directory containing config, conversations, global instructions, and installed docs.
**Compacted tool output**: A model-facing replacement for a large tool result that keeps a head/tail excerpt plus provenance and recovery guidance so the assistant can re-read or re-search narrowly before relying on omitted details.
**Exact edit**: An `edit` tool replacement where each `old_text` must match exactly once in the original file before anything is written.
**Fast mode**: A saved preference enabled with `/fast`. It is active only when the current provider/model advertises fast-mode support; otherwise Cassady keeps the preference but reports it as unavailable.
**Global instructions**: Optional text in `~/.cass/global.md` included in new chat system prompts. Cassady follows these instructions when they fit the active request, but they cannot override runtime safety constraints such as access modes, tool denials, approvals, or workspace boundaries.
**Model metadata**: The `models.json` entry describing a model id, owning provider, display name, context limits, tool/streaming support, and reasoning behavior.
**Model metadata**: The `models.json` entry describing a model id, owning provider, display name, context limits, tool/streaming support, reasoning behavior, and fast-mode support.
**OpenAI-compatible provider**: A provider exposing an API compatible with the OpenAI-style chat/completions behavior Cassady uses.
@@ -26,4 +30,6 @@
**Tool call**: A model-requested operation such as `ls`, `read`, `grep`, `write`, `edit`, or `shell`.
**Truncated tool output**: A model-facing shortened tool result produced when output exceeds `model_tool_result_limit`. Cassady tells the model that output was incomplete and suggests narrower follow-up inspection.
**Workspace**: The launch cwd, either the current directory or the path passed with `--cwd`. In workspace-edit mode, writes must stay inside this root.
+12 -2
View File
@@ -74,7 +74,8 @@ Provider protocols that are not OpenAI-compatible are supported only when Cassad
- display name;
- context length and max output tokens;
- tool and streaming support;
- reasoning support and request format.
- reasoning support and request format;
- fast-mode support.
`config.json` selects active defaults, such as `default_provider`, `default_model`, and `default_access_mode`.
@@ -90,6 +91,15 @@ Reasoning metadata controls how the runtime reasoning effort behaves:
Reasoning display is separate. `show_reasoning` controls whether provider-streamed reasoning is visible in the transcript; press `Ctrl-Shift-R` or `Ctrl-R` to toggle it at runtime.
## Fast-mode metadata
Fast mode has two parts:
- `default_fast_mode` in `config.json`: the user's saved preference.
- `fast_mode.supported` in `models.json`: whether non-Codex provider/model metadata can honor that preference.
Cassady sends fast-mode requests only for `ChatGPT Codex`. Any active `chatgpt-codex` provider model, including `gpt-5.5`, is treated as fast-capable so older model metadata does not block the feature. OpenAI-compatible and custom model entries default to unsupported, so `/fast` can remember the preference without sending provider-specific fields.
## Switching models
Use one of these approaches:
@@ -104,7 +114,7 @@ or inside a chat:
/model MODEL
```
The in-chat model autocomplete lists entries from `~/.cass/models.json`. Switching the model also updates the default model and reasoning effort in `config.json` for future sessions.
The in-chat model autocomplete lists entries from `~/.cass/models.json`. Switching the model also updates the default provider, default model, and reasoning effort in `config.json` for future sessions. If fast mode is preferred, Cassady recomputes whether it is active after the switch.
## Health checks
+8
View File
@@ -136,6 +136,14 @@ Likely cause: the command itself failed, the working directory is wrong, depende
Fix: inspect stdout/stderr, verify cwd in `/status`, and ask Cassady to rerun the smallest relevant command.
## Tool output was compacted or truncated
Symptom: a tool result note says Cassady compacted or truncated output, shows retained head/tail excerpts, or says a search stopped after a match limit.
Likely cause: the raw tool output exceeded the model-facing result limit or the active model context budget.
Fix: treat omitted output as incomplete. Ask Cassady to re-read the exact file line range named in the note, run a narrower `grep` query/path, lower `max_matches`, or rerun a shell command with a more focused flag/filter before making edits based on omitted lines.
## Exact-text edit failed
Symptom: edit reports `old_text not found`, `old_text is not unique`, or overlapping edits.
+32 -2
View File
@@ -29,7 +29,25 @@ Start in read-only mode or press `Shift-Tab` until the status shows `read-only`.
Find where configuration is loaded and summarize the precedence rules.
```
Cassady can use `ls`, `read`, and `grep` to inspect the workspace and bundled docs.
Cassady can use `ls`, `read`, and `grep` to inspect the workspace and bundled docs. For large files or unknown locations, prefer a search-first flow: `grep` for a symbol or phrase, then `read` a small line range around the relevant match.
## Recover from compacted or truncated output
When a tool result is too large for the model context, Cassady sends the model a head/tail excerpt with a recovery note. The note includes the tool name, retained excerpt shape, and file range or shell-command provenance when available.
If Cassady reports compacted or truncated output, do not rely on omitted lines for edits. Ask Cass to narrow the inspection instead:
```text
Re-read src/app.rs lines 220-280 before editing that function.
```
```text
Search only src/ for "load_config" and then read around the matching lines.
```
```text
Rerun the test command with a focused package/filter, or pipe the noisy output through grep/head/tail.
```
## Apply a focused edit
@@ -95,7 +113,7 @@ Inside a chat:
/model MODEL_ID
```
Autocomplete lists models from `~/.cass/models.json`. Switching models is allowed only when idle. Cassady persists the last used model and reasoning effort into `config.json`.
Autocomplete lists models from `~/.cass/models.json`. Switching models is allowed only when idle. Cassady persists the last used provider, model, and reasoning effort into `config.json`.
You can also launch with a model override:
@@ -103,6 +121,18 @@ You can also launch with a model override:
cass --model MODEL_ID
```
## Prefer fast mode
Inside a chat:
```text
/fast
```
Use `/fast on`, `/fast off`, or `/fast status` when you want an explicit action. Cassady saves the preference in `config.json`, but fast mode is active only when the current provider/model supports it. ChatGPT Codex models, including `gpt-5.5`, are treated as fast-capable.
Switching to an unsupported provider/model keeps the preference but makes `/status` show fast mode as unavailable. Switching back to ChatGPT Codex enables it again.
## Resume a chat
List chats for the current directory:
@@ -0,0 +1,175 @@
# v0.3.5 Tool Output Context Reliability Implementation Plan
## Goal
v0.3.5 makes Cassady more reliable after broad tool output has been truncated, compacted, or superseded in the model context. The assistant should be able to tell when details are missing, understand which file range or command produced them, and quickly recover by using narrower reads or searches instead of stalling or making unsafe edits from incomplete context.
Success statement:
> After a large read or command output is compacted out of the request context, the assistant receives concise recovery guidance with enough provenance to inspect the exact missing area again before editing.
## Scope
### In scope
- Improve model-facing compaction notes for large tool outputs.
- Preserve useful provenance for compacted `read`, `grep`, and `shell` outputs.
- Add focused guidance that nudges the assistant toward smaller line ranges and search-first workflows.
- Keep existing provider message structure valid when tool outputs are compacted or earlier records are omitted.
- Align UI summaries, stored records, and model-facing transformed output so truncation/compaction is understandable without changing the full conversation history.
- Add regression tests for broad-output recovery and context-budget trimming.
- Update README and bundled docs where they describe context management, tool output limits, and recommended inspection workflows.
### Out of scope
- Implementing semantic summarization with an additional model call.
- Replacing Cassady's approximate token estimator with provider-specific tokenizers.
- Adding a full retrieval index over prior tool outputs or repository contents.
- Changing the JSONL conversation storage format in a way that makes existing chats unreadable.
- Changing UI collapsed-tool behavior except where labels or summaries need to expose truncation/compaction status.
- Automatically editing files based on compacted output without reinspection.
## Context or Current State
Relevant current behavior:
- `src/agent.rs` converts conversation records into provider messages, supersedes older repeated read outputs, compacts older tool outputs with `compact_tool_outputs`, and trims records with `trim_to_context_budget` and `trim_to_message_limit`.
- `compact_text` currently emits a generic head/tail note: `Cass compacted this tool output from ... chars to fit the model context`.
- `superseded_read_note` already preserves file path and line range when a later read covers an earlier read range.
- `src/tools/read.rs` returns headers like `--- path lines start-end ---`, followed by numbered lines. This is good provenance, but compaction can obscure the most useful middle section.
- `src/tools/grep.rs` already recommends narrowing when `max_matches` is reached.
- `src/tools/shell.rs` returns complete stdout/stderr/exit-code text to the agent loop; if the output is large, current compaction does not know the original command or suggest a narrower command.
- `src/ui/render.rs` has separate collapsed/full tool-output presentation. That display choice must remain UI-only and must not affect the stored record or model-facing context.
The main reliability gap is that once a broad output has been compacted, the assistant may see only a generic excerpt and lose the clue needed to make the next targeted tool call.
## Design Principles
1. **Never hide incompleteness.** If Cassady compacts or truncates output before sending it to the model, the transformed content must say so plainly.
2. **Recovery beats summarization.** Prefer actionable provenance and follow-up instructions over trying to summarize omitted content heuristically.
3. **Keep guidance compact.** The fix must not consume enough context to make context pressure worse.
4. **Preserve valid provider conversations.** Tool result messages must still match their assistant tool calls after compaction and trimming.
5. **Do not mutate history for UI convenience.** JSONL records should retain original tool outputs unless a future storage migration explicitly changes that contract.
## Design
### Model-facing compaction notes
Replace the generic `compact_text(content, target_chars)` path with a metadata-aware formatter, for example:
```rust
struct ToolOutputCompactionHint {
tool_name: Option<String>,
original_chars: usize,
retained_head_chars: usize,
retained_tail_chars: usize,
provenance: ToolOutputProvenance,
}
enum ToolOutputProvenance {
Read { sections: Vec<ReadOutputSection> },
Grep { stopped_after: Option<usize> },
Shell { command: Option<String> },
Unknown,
}
```
The first implementation can infer provenance from the tool result text and nearby conversation/tool-call data rather than changing stored record schemas.
Example compacted read result:
```text
[Cass compacted this read output from 48,212 chars to fit the model context. The omitted content came from src/app.rs lines 1-1820. Use read with a narrower line range, or grep for a symbol before reading, before relying on omitted details.]
--- retained head excerpt ---
...
--- omitted middle ---
--- retained tail excerpt ---
...
```
Example compacted shell result:
```text
[Cass compacted this shell output from 81,004 chars to fit the model context. Rerun a narrower command, pipe through grep/head/tail, or inspect the specific files named in the excerpt before making edits based on omitted lines.]
```
Keep notes deterministic and short; avoid per-line summaries of omitted content.
### Read-output provenance
Reuse and extend the existing `ReadOutputSection` parsing in `src/agent.rs`:
- Detect every `--- path lines start-end ---` section before compaction.
- Preserve observed line coverage from numbered lines when available.
- Include one compact range summary in compaction notes:
- Single section: `path lines 35-220`.
- Multiple sections: `3 read sections including path_a lines 1-120 and path_b lines 40-90`.
- When a later read supersedes an earlier range, continue using the current superseded-read note and ensure tests cover interaction with compaction.
### Grep and shell guidance
For `grep` output:
- Preserve existing `… stopped after N matches` text.
- If compacted, add a note suggesting a narrower query, smaller path scope, lower `max_matches`, or a focused `read` around matching lines.
For `shell` output:
- If the tool-call arguments are available in the message conversion path, include a sanitized command preview in the note when reasonably short.
- Suggest command narrowing patterns without prescribing platform-specific syntax unless the command itself is already shell-specific, e.g. `grep`, `head`, `tail`, or a more targeted subcommand.
- Do not rerun shell commands automatically.
### Prompt and tool descriptions
Update the base prompt and tool descriptions only enough to reinforce reliable behavior:
- Prefer `grep` before broad `read` when the target location is unknown.
- Read smaller line ranges when files are large or when previous output says it was compacted.
- Treat compacted/truncated output as incomplete evidence; reinspect before editing.
Avoid bloating `src/prompt.rs`; keep additions short and test expected key phrases rather than full prompt text.
### UI and storage alignment
- Stored JSONL should keep the original tool result content.
- Model-facing transformed messages may contain compacted/superseded notes.
- UI collapsed mode should keep using summaries, but summaries should not imply the model saw the full output when it did not.
- If practical, make collapsed tool summaries include a compact `compacted` or `truncated` marker only when the stored/result text itself says that Cassady truncated or stopped output.
## Implementation Steps
1. Inspect provider-message conversion in `src/agent.rs` and identify where tool-call names/arguments are still available when compacting tool results.
2. Refactor `compact_text` into metadata-aware helpers that can produce deterministic compaction notes for read, grep, shell, and unknown outputs.
3. Reuse existing read-section parsing to build concise read range summaries for compacted read outputs.
4. Add shell and grep-specific recovery guidance based on tool name and output markers.
5. Preserve the newest tool result behavior unless tests show that the newest result can still exceed practical context limits; if changed, document the tradeoff explicitly.
6. Update `src/tools/read.rs`, `src/tools/grep.rs`, and `src/prompt.rs` descriptions with concise search-first and narrow-range guidance.
7. Add or update UI summary helpers in `src/ui/render.rs` only if needed to expose stored truncation/compaction markers consistently.
8. Update README and bundled docs for context reliability, broad-output recovery, and recommended inspection workflow.
9. Run `cargo fmt` and `cargo test --locked --all-targets`.
## Tests
- Large read output compacts to a note that includes original size, path, line range, and a narrower-read/search suggestion.
- Multi-file read output compacts to a concise multi-section provenance summary.
- Superseded read output still produces the superseded note and does not lose provider-message validity after context trimming.
- Large grep output compaction preserves or adds narrowing guidance.
- Large shell output compaction suggests rerunning a narrower command and does not include unsafe automatic actions.
- Context-budget trimming does not leave orphaned tool results or assistant tool calls.
- Stored conversation records retain original tool output while model-facing messages can be compacted.
- Prompt/tool spec tests verify the presence of concise search-first and reinspection guidance.
## Documentation
- Update `README.md` where tool output/context behavior is described.
- Update `docs/workflows.md` with recommended search-first and narrow-read workflows.
- Update `docs/troubleshooting.md` with recovery steps for compacted or truncated output.
- Update `docs/glossary.md` if terms such as compacted output, superseded read, or model-facing context need clarification.
## Acceptance Criteria
- Compacted tool outputs include actionable provenance and recovery guidance.
- Broad read and command-output workflows have regression coverage demonstrating safe reinspection before edits.
- Existing conversations remain loadable and resumable.
- Provider message conversion remains valid for tool-call/tool-result pairs after compaction and trimming.
- `cargo fmt` and `cargo test --locked --all-targets` pass.
+475 -23
View File
@@ -3,7 +3,7 @@ use crate::config::{Config, ReasoningEffort};
use crate::conversation::{now_ts, Conversation, Record, StoredToolCall};
use crate::prompt;
use crate::providers::types::ModelMessage;
use crate::providers::ProviderClient;
use crate::providers::{ProviderClient, ProviderRuntimeOptions};
use crate::security::PolicyDecision;
use crate::tools::{self, ToolContext, ToolRuntimeEvent};
use anyhow::Result;
@@ -88,7 +88,13 @@ pub async fn run_turn_with_commands(
let reasoning_effort = settings
.reasoning_effort
.clamp_for_model(settings.config.model_metadata.as_ref());
let provider = match ProviderClient::from_config(&settings.config, reasoning_effort) {
let provider = match ProviderClient::from_config(
&settings.config,
ProviderRuntimeOptions {
reasoning_effort,
fast_mode: settings.config.fast_mode_state().active,
},
) {
Ok(provider) => provider,
Err(err) => {
append_visible_assistant(
@@ -107,7 +113,9 @@ pub async fn run_turn_with_commands(
cwd: settings.cwd.clone(),
read_roots: vec![settings.cwd.clone(), docs_dir.clone()],
blocked_write_roots: vec![docs_dir.clone()],
model_result_limit: settings.config.model_tool_result_limit,
// Keep stored/UI tool results intact; build_messages applies the
// model-facing result limit when preparing provider messages.
model_result_limit: usize::MAX,
runtime_tx: None,
};
@@ -470,6 +478,7 @@ fn build_messages(records: &[Record], system: String, config: &Config) -> Vec<Mo
messages.extend(records.iter().filter_map(record_to_model_message));
messages = sanitize_tool_message_structure(messages);
supersede_old_read_outputs(&mut messages);
apply_model_tool_result_limits(&mut messages, config.model_tool_result_limit);
let budget = context_budget_tokens(config);
if estimate_messages_tokens(&messages) > budget {
@@ -512,7 +521,7 @@ fn record_to_model_message(record: &Record) -> Option<ModelMessage> {
}
fn supersede_old_read_outputs(messages: &mut [ModelMessage]) {
let read_calls = read_tool_calls_by_id(messages);
let read_calls = tool_calls_by_id(messages);
let mut read_outputs = Vec::new();
for (message_idx, message) in messages.iter().enumerate() {
let ModelMessage::Tool {
@@ -573,16 +582,14 @@ fn tool_content(messages: &[ModelMessage], idx: usize) -> &str {
}
}
fn read_tool_calls_by_id(messages: &[ModelMessage]) -> BTreeMap<String, StoredToolCall> {
fn tool_calls_by_id(messages: &[ModelMessage]) -> BTreeMap<String, StoredToolCall> {
let mut calls = BTreeMap::new();
for message in messages {
let ModelMessage::Assistant { tool_calls, .. } = message else {
continue;
};
for call in tool_calls {
if call.name == "read" {
calls.insert(call.id.clone(), call.clone());
}
calls.insert(call.id.clone(), call.clone());
}
}
calls
@@ -824,43 +831,134 @@ fn context_budget_tokens(config: &Config) -> usize {
.max(MIN_INPUT_BUDGET_TOKENS)
}
#[derive(Debug, Clone, Copy)]
enum ToolOutputTransform {
Compaction,
Truncation,
}
impl ToolOutputTransform {
fn verb(self) -> &'static str {
match self {
ToolOutputTransform::Compaction => "compacted",
ToolOutputTransform::Truncation => "truncated",
}
}
fn reason(self) -> &'static str {
match self {
ToolOutputTransform::Compaction => "to fit the model context",
ToolOutputTransform::Truncation => {
"before sending it to the model because it exceeded the configured model_tool_result_limit"
}
}
}
}
fn apply_model_tool_result_limits(messages: &mut [ModelMessage], limit_chars: usize) {
let tool_calls = tool_calls_by_id(messages);
let target_chars = limit_chars.max(1);
for idx in 1..messages.len() {
transform_tool_output_at(
messages,
idx,
target_chars,
ToolOutputTransform::Truncation,
&tool_calls,
);
}
}
fn compact_tool_outputs(messages: &mut [ModelMessage], budget: usize) {
let newest_tool_idx = messages
.iter()
.rposition(|message| matches!(message, ModelMessage::Tool { .. }));
let tool_calls = tool_calls_by_id(messages);
for idx in 1..messages.len() {
if Some(idx) == newest_tool_idx {
continue;
}
compact_tool_output_at(messages, idx, TOOL_OUTPUT_COMPACT_CHARS);
transform_tool_output_at(
messages,
idx,
TOOL_OUTPUT_COMPACT_CHARS,
ToolOutputTransform::Compaction,
&tool_calls,
);
if estimate_messages_tokens(messages) <= budget {
return;
}
}
for idx in 1..messages.len() {
compact_tool_output_at(messages, idx, TOOL_OUTPUT_TINY_CHARS);
transform_tool_output_at(
messages,
idx,
TOOL_OUTPUT_TINY_CHARS,
ToolOutputTransform::Compaction,
&tool_calls,
);
if estimate_messages_tokens(messages) <= budget {
return;
}
}
}
fn compact_tool_output_at(messages: &mut [ModelMessage], idx: usize, target_chars: usize) {
let Some(ModelMessage::Tool { content, .. }) = messages.get_mut(idx) else {
return;
fn transform_tool_output_at(
messages: &mut [ModelMessage],
idx: usize,
target_chars: usize,
transform: ToolOutputTransform,
tool_calls: &BTreeMap<String, StoredToolCall>,
) {
let replacement = {
let Some(ModelMessage::Tool {
tool_call_id,
name,
content,
}) = messages.get(idx)
else {
return;
};
if content.chars().count() <= target_chars.max(1) {
return;
}
format_tool_output_transform(
name,
tool_calls.get(tool_call_id),
content,
target_chars,
transform,
)
};
if content.chars().count() <= target_chars {
return;
if let Some(ModelMessage::Tool { content, .. }) = messages.get_mut(idx) {
*content = replacement;
}
*content = compact_text(content, target_chars);
}
fn compact_text(content: &str, target_chars: usize) -> String {
fn format_tool_output_transform(
tool_name: &str,
call: Option<&StoredToolCall>,
content: &str,
target_chars: usize,
transform: ToolOutputTransform,
) -> String {
let original_chars = content.chars().count();
let head_chars = (target_chars * 2 / 3).max(1);
let tail_chars = target_chars.saturating_sub(head_chars).max(1);
if original_chars <= target_chars.max(1) {
return content.to_string();
}
let target_chars = target_chars
.max(1)
.min(original_chars.saturating_sub(1).max(1));
let head_chars = if target_chars <= 1 {
1
} else {
(target_chars * 2 / 3).clamp(1, target_chars - 1)
};
let tail_chars = target_chars.saturating_sub(head_chars);
let head: String = content.chars().take(head_chars).collect();
let tail: String = content
.chars()
@@ -870,11 +968,208 @@ fn compact_text(content: &str, target_chars: usize) -> String {
.into_iter()
.rev()
.collect();
let notice = tool_output_transform_notice(
tool_name,
call,
content,
original_chars,
head_chars,
tail_chars,
transform,
);
let mut out = String::new();
out.push_str(&notice);
out.push_str("\n--- retained head excerpt ---\n");
out.push_str(&head);
out.push_str("\n--- omitted middle ---\n");
if tail_chars > 0 {
out.push_str("--- retained tail excerpt ---\n");
out.push_str(&tail);
}
out
}
fn tool_output_transform_notice(
tool_name: &str,
call: Option<&StoredToolCall>,
content: &str,
original_chars: usize,
head_chars: usize,
tail_chars: usize,
transform: ToolOutputTransform,
) -> String {
let retained_shape = if tail_chars > 0 {
format!("{head_chars} chars from the start and {tail_chars} chars from the end")
} else {
format!("{head_chars} chars from the start")
};
let mut sentences = vec![format!(
"Cass {} this tool output from {original_chars} chars {}. Tool: `{tool_name}`. Retained excerpt shape: {retained_shape}.",
transform.verb(),
transform.reason(),
)];
if let Some(provenance) = tool_output_provenance_sentence(tool_name, call, content) {
sentences.push(provenance);
}
sentences.push(tool_output_recovery_guidance(tool_name).to_string());
format!("[{}]", sentences.join(" "))
}
fn tool_output_provenance_sentence(
tool_name: &str,
call: Option<&StoredToolCall>,
content: &str,
) -> Option<String> {
match tool_name {
"read" => read_output_provenance_sentence(call, content),
"grep" => grep_output_provenance_sentence(content),
"shell" => shell_output_provenance_sentence(call),
_ => None,
}
}
fn read_output_provenance_sentence(call: Option<&StoredToolCall>, content: &str) -> Option<String> {
let request_specs = call
.filter(|call| call.name == "read")
.map(|call| read_request_specs(&call.arguments))
.unwrap_or_default();
let sections = parse_read_output_sections(content, &request_specs);
if !sections.is_empty() {
return Some(format!(
"Omitted content came from {}.",
read_sections_summary(&sections)
));
}
read_request_summary(call).map(|summary| format!("The read request targeted {summary}."))
}
fn read_sections_summary(sections: &[ReadOutputSection]) -> String {
let labels: Vec<String> = sections.iter().take(2).map(read_section_label).collect();
match sections.len() {
0 => "no read sections".into(),
1 => labels[0].clone(),
2 => format!("2 read sections: {} and {}", labels[0], labels[1]),
count => format!(
"{count} read sections including {} and {}",
labels[0], labels[1]
),
}
}
fn read_section_label(section: &ReadOutputSection) -> String {
format!(
"[Cass compacted this tool output from {original_chars} chars to fit the model context. Head/tail excerpt follows.]\n{head}\n… omitted …\n{tail}"
"{} lines {}-{}",
section.path, section.start_line, section.end_line
)
}
fn read_request_summary(call: Option<&StoredToolCall>) -> Option<String> {
let call = call.filter(|call| call.name == "read")?;
let mut labels = Vec::new();
if let Some(files) = call
.arguments
.get("files")
.and_then(|files| files.as_array())
{
for file in files.iter().take(2) {
let Some(path) = file.get("path").and_then(|path| path.as_str()) else {
continue;
};
labels.push(read_request_label(
path,
file.get("lines").and_then(|lines| lines.as_str()),
));
}
return match labels.len() {
0 => None,
1 => labels.first().cloned(),
2 if files.len() == 2 => Some(format!("2 files: {} and {}", labels[0], labels[1])),
_ => Some(format!(
"{} files including {} and {}",
files.len(),
labels[0],
labels[1]
)),
};
}
call.arguments
.get("path")
.and_then(|path| path.as_str())
.map(|path| {
read_request_label(
path,
call.arguments.get("lines").and_then(|lines| lines.as_str()),
)
})
}
fn read_request_label(path: &str, lines: Option<&str>) -> String {
match lines.map(str::trim).filter(|lines| !lines.is_empty()) {
Some(lines) => format!("{path} lines {lines}"),
None => path.to_string(),
}
}
fn grep_output_provenance_sentence(content: &str) -> Option<String> {
grep_stopped_after(content)
.map(|count| format!("The grep output reported it stopped after {count} matches."))
}
fn grep_stopped_after(content: &str) -> Option<usize> {
content.lines().find_map(|line| {
let rest = line.trim().strip_prefix("… stopped after ")?;
rest.split_whitespace().next()?.parse().ok()
})
}
fn shell_output_provenance_sentence(call: Option<&StoredToolCall>) -> Option<String> {
let call = call.filter(|call| call.name == "shell")?;
let command = call.arguments.get("command")?.as_str()?.trim();
if command.is_empty() {
return None;
}
let preview = command.split_whitespace().collect::<Vec<_>>().join(" ");
if command_preview_may_contain_sensitive_text(&preview) {
return Some(
"Shell command preview omitted because it may contain sensitive text; inspect the preceding tool-call arguments before rerunning."
.into(),
);
}
let chars = preview.chars().count();
if chars > 160 {
return Some(format!(
"Shell command was {chars} chars; inspect the preceding tool-call arguments before rerunning."
));
}
Some(format!("Shell command: `{preview}`."))
}
fn command_preview_may_contain_sensitive_text(command: &str) -> bool {
let lowered = command.to_ascii_lowercase();
[
"api_key",
"apikey",
"authorization",
"bearer",
"password",
"secret",
"token",
]
.iter()
.any(|needle| lowered.contains(needle))
}
fn tool_output_recovery_guidance(tool_name: &str) -> &'static str {
match tool_name {
"read" => "Recovery: use `read` with a narrower line range, or `grep` for a symbol before reading, before relying on omitted details.",
"grep" => "Recovery: rerun `grep` with a narrower query/path, lower `max_matches`, or `read` around specific matching lines before relying on omitted details.",
"shell" => "Recovery: rerun a narrower command, filter output with grep/head/tail, or inspect specific files named in the excerpt before making edits from omitted lines.",
_ => "Recovery: rerun a narrower tool request or inspect the specific file/range from the excerpt before relying on omitted details.",
}
}
fn trim_to_context_budget(mut messages: Vec<ModelMessage>, budget: usize) -> Vec<ModelMessage> {
let mut omitted = false;
while estimate_messages_tokens(&messages) > budget && messages.len() > 1 {
@@ -888,7 +1183,7 @@ fn trim_to_context_budget(mut messages: Vec<ModelMessage>, budget: usize) -> Vec
if omitted {
let note = ModelMessage::System {
content: "Cass omitted earlier conversation messages to fit the model context budget. Included tool results still follow their matching assistant tool calls; older large tool outputs may be compacted.".to_string(),
content: "Cass omitted earlier conversation messages to fit the model context budget. Included tool results still follow their matching assistant tool calls; older large tool outputs may be compacted with recovery guidance.".to_string(),
};
messages.insert(1, note);
while estimate_messages_tokens(&messages) > budget && messages.len() > 2 {
@@ -1066,7 +1361,7 @@ fn _calls(_calls: Vec<StoredToolCall>) {}
mod tests {
use super::*;
use crate::config::default_model_definition;
use serde_json::json;
use serde_json::{json, Value};
fn small_config(context_length: u64, max_output_tokens: u64) -> Config {
let mut config = Config::default();
@@ -1093,6 +1388,22 @@ mod tests {
}
}
fn named_call(id: &str, name: &str, arguments: Value) -> StoredToolCall {
StoredToolCall {
id: id.to_string(),
name: name.to_string(),
arguments,
}
}
fn numbered_read_output(path: &str, start: usize, end: usize) -> String {
let mut out = format!("--- {path} lines {start}-{end} ---\n");
for line in start..=end {
out.push_str(&format!("{line:>6} | line {line}\n"));
}
out
}
fn assert_valid_tool_structure(messages: &[ModelMessage]) {
let mut idx = 1;
while idx < messages.len() {
@@ -1147,7 +1458,9 @@ mod tests {
ts: now_ts(),
},
];
let messages = build_messages(&records, "system".into(), &small_config(8_000, 512));
let mut config = small_config(8_000, 512);
config.model_tool_result_limit = 100_000;
let messages = build_messages(&records, "system".into(), &config);
assert_valid_tool_structure(&messages);
assert!(messages.iter().any(|message| matches!(
@@ -1156,6 +1469,145 @@ mod tests {
)));
}
#[test]
fn compacted_read_output_includes_range_and_reinspection_guidance() {
let call = call_with_lines("call_1", "src/app.rs", "1-200");
let content = numbered_read_output("src/app.rs", 1, 200);
let compacted = format_tool_output_transform(
"read",
Some(&call),
&content,
320,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("Cass compacted this tool output"));
assert!(compacted.contains("Tool: `read`"));
assert!(compacted.contains("Retained excerpt shape"));
assert!(compacted.contains("src/app.rs lines 1-200"));
assert!(compacted.contains("narrower line range"));
assert!(compacted.contains("grep"));
assert!(compacted.contains("--- retained head excerpt ---"));
assert!(compacted.contains("--- retained tail excerpt ---"));
}
#[test]
fn compacted_multi_file_read_output_summarizes_sections() {
let call = named_call(
"call_1",
"read",
json!({"files":[
{"path":"src/a.rs","lines":"1-80"},
{"path":"src/b.rs","lines":"20-90"}
]}),
);
let content = format!(
"{}{}",
numbered_read_output("src/a.rs", 1, 80),
numbered_read_output("src/b.rs", 20, 90)
);
let compacted = format_tool_output_transform(
"read",
Some(&call),
&content,
360,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("2 read sections"));
assert!(compacted.contains("src/a.rs lines 1-80"));
assert!(compacted.contains("src/b.rs lines 20-90"));
}
#[test]
fn compacted_grep_output_adds_narrowing_guidance() {
let call = named_call(
"call_1",
"grep",
json!({"query":"needle","paths":["src"],"max_matches":300}),
);
let mut content = String::new();
for line in 1..=300 {
content.push_str(&format!("src/lib.rs:{line}: needle {line}\n"));
}
content.push_str("… stopped after 300 matches. Narrow the query or raise max_matches.\n");
let compacted = format_tool_output_transform(
"grep",
Some(&call),
&content,
240,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("Tool: `grep`"));
assert!(compacted.contains("stopped after 300 matches"));
assert!(compacted.contains("narrower query/path"));
assert!(compacted.contains("read` around specific matching lines"));
}
#[test]
fn compacted_shell_output_mentions_command_and_narrowing() {
let call = named_call(
"call_1",
"shell",
json!({"command":"cargo test --locked --all-targets"}),
);
let content = format!("stdout:\n{}\nexit code: 0\n", "test output\n".repeat(200));
let compacted = format_tool_output_transform(
"shell",
Some(&call),
&content,
240,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("Tool: `shell`"));
assert!(compacted.contains("Shell command: `cargo test --locked --all-targets`"));
assert!(compacted.contains("filter output with grep/head/tail"));
assert!(compacted.contains("before making edits from omitted lines"));
}
#[test]
fn model_tool_result_limit_is_model_facing_only() {
let original = numbered_read_output("src/main.rs", 1, 120);
let records = vec![
Record::Assistant {
content: String::new(),
reasoning: String::new(),
reasoning_field: None,
tool_calls: vec![call_with_lines("call_1", "src/main.rs", "1-120")],
ts: now_ts(),
},
Record::Tool {
tool_call_id: "call_1".into(),
name: "read".into(),
ok: true,
content: original.clone(),
ts: now_ts(),
},
];
let mut config = Config::default();
config.model_tool_result_limit = 180;
let messages = build_messages(&records, "system".into(), &config);
assert_valid_tool_structure(&messages);
assert!(messages.iter().any(|message| matches!(
message,
ModelMessage::Tool { content, .. }
if content.contains("Cass truncated this tool output")
&& content.contains("src/main.rs lines 1-120")
)));
assert!(matches!(
&records[1],
Record::Tool { content, .. } if content == &original
));
}
#[test]
fn context_budget_trimming_does_not_leave_orphaned_tool_results() {
let records = vec![
+225 -13
View File
@@ -1,6 +1,6 @@
use crate::agent::{self, AgentCommand, AgentEvent, AgentSettings};
use crate::cli::{self, Cli, Command};
use crate::config::{self, Config, ModelDefinition, ReasoningEffort};
use crate::config::{self, Config, FastModeState, ModelDefinition, ReasoningEffort};
use crate::conversation::{self, Conversation, Record};
use crate::prompt;
use crate::ui::autofill::{AutoFillItem, AutoFillMenu};
@@ -397,6 +397,7 @@ async fn run_tui(
show_full_tools,
show_reasoning,
reasoning_effort,
fast_mode_active: config.fast_mode_state().active,
scroll,
autofill: if branch_menu.is_some() {
None
@@ -862,10 +863,46 @@ async fn run_tui(
}
}
}
Ok(LocalCommand::Fast(command)) => {
if busy {
status = "fast mode can be changed when idle".into();
} else {
input.clear();
autofill_selected = 0;
match apply_fast_mode_command(&mut config, command) {
Ok(message) => {
transcript.push(TranscriptBlock {
kind: TranscriptKind::Status,
title: "fast".into(),
content: message.clone(),
});
status = message;
}
Err(err) => {
status =
format!("fast mode update failed: {err}");
transcript.push(TranscriptBlock {
kind: TranscriptKind::Error,
title: "fast".into(),
content: err.to_string(),
});
}
}
if stick_to_bottom {
scroll = bottom_scroll(
&terminal,
&input,
&transcript,
show_full_tools,
show_reasoning,
)?;
}
}
}
Ok(LocalCommand::Status) => {
let content = chat_status(
&chat_id,
&config.model,
&config,
mode,
&cwd,
busy,
@@ -894,14 +931,13 @@ async fn run_tui(
if busy {
status = "model can be changed when idle".into();
} else {
config.model = model.clone();
config.model_metadata =
model_metadata_for(&config, &model)?;
apply_model_selection(&mut config, &model)?;
reasoning_effort = ReasoningEffort::default_for_model(
config.model_metadata.as_ref(),
);
let _ = crate::config::save_last_used(
let _ = crate::config::save_last_used_provider(
&config.root,
&config.provider_id,
&config.model,
reasoning_effort,
);
@@ -910,9 +946,9 @@ async fn run_tui(
transcript.push(TranscriptBlock {
kind: TranscriptKind::Status,
title: "model".into(),
content: format!("model changed to {model}"),
content: model_status_message(&config),
});
status = format!("model: {model}");
status = model_status_message(&config);
if stick_to_bottom {
scroll = bottom_scroll(
&terminal,
@@ -1885,6 +1921,7 @@ fn assistant_content_matches(a: &str, b: &str) -> bool {
#[derive(Debug, Clone, PartialEq, Eq)]
enum LocalCommand {
Branch,
Fast(FastModeCommand),
Login,
Logout,
Model(String),
@@ -1893,6 +1930,14 @@ enum LocalCommand {
Status,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum FastModeCommand {
Toggle,
On,
Off,
Status,
}
struct CommandSpec {
name: &'static str,
usage: &'static str,
@@ -1907,6 +1952,12 @@ const COMMANDS: &[CommandSpec] = &[
description: "open branch/restore menu",
takes_value: false,
},
CommandSpec {
name: "fast",
usage: "/fast [on|off|status]",
description: "toggle faster Codex inference when supported",
takes_value: false,
},
CommandSpec {
name: "login",
usage: "/login",
@@ -2038,14 +2089,38 @@ fn model_autofill(input: &str, selected: usize, config: &Config) -> Result<Optio
}
}
fn model_metadata_for(config: &Config, model_id: &str) -> Result<Option<ModelDefinition>> {
fn apply_model_selection(config: &mut Config, model_id: &str) -> Result<()> {
let models = crate::config::load_or_create_default_model_registry(&config.root)?;
Ok(models
let metadata = models
.models
.iter()
.find(|model| model.id == model_id && model.provider == config.provider_id)
.cloned()
.or_else(|| models.models.into_iter().find(|model| model.id == model_id)))
.or_else(|| {
models
.models
.iter()
.find(|model| model.id == model_id)
.cloned()
});
if let Some(model) = &metadata {
if model.provider != config.provider_id {
let providers = crate::config::load_or_create_default_provider_registry(&config.root)?;
if let Some(provider) = providers
.providers
.iter()
.find(|provider| provider.id == model.provider)
{
config.provider_id = provider.id.clone();
config.active_provider = provider.to_resolved();
}
}
}
config.model = model_id.to_string();
config.model_metadata = metadata;
Ok(())
}
fn model_matches(model: &ModelDefinition, query: &str) -> bool {
@@ -2175,6 +2250,19 @@ fn parse_local_command(input: &str) -> std::result::Result<LocalCommand, String>
}
Ok(LocalCommand::Branch)
}
"/fast" => {
let command = match parts.next() {
None => FastModeCommand::Toggle,
Some("on") => FastModeCommand::On,
Some("off") => FastModeCommand::Off,
Some("status") => FastModeCommand::Status,
Some(_) => return Err("usage: /fast [on|off|status]".into()),
};
if parts.next().is_some() {
return Err("usage: /fast [on|off|status]".into());
}
Ok(LocalCommand::Fast(command))
}
"/login" => {
if parts.next().is_some() {
return Err("usage: /login".into());
@@ -2223,7 +2311,7 @@ fn parse_local_command(input: &str) -> std::result::Result<LocalCommand, String>
fn chat_status(
chat_id: &str,
model: &str,
config: &Config,
mode: crate::access::AccessMode,
cwd: &Path,
busy: bool,
@@ -2231,13 +2319,76 @@ fn chat_status(
record_count: usize,
) -> String {
format!(
"chat: {chat_id}\nstate: {}\nmodel: {model}\nmode: {mode}\ncwd: {}\nrecords: {record_count}\nstatus: {}",
"chat: {chat_id}\nstate: {}\nmodel: {}\nfast: {}\nmode: {mode}\ncwd: {}\nrecords: {record_count}\nstatus: {}",
if busy { "running" } else { "idle" },
config.model,
fast_mode_status(&config.fast_mode_state()),
cwd.display(),
if status.is_empty() { "idle" } else { status }
)
}
fn apply_fast_mode_command(config: &mut Config, command: FastModeCommand) -> Result<String> {
match command {
FastModeCommand::Status => Ok(fast_mode_status(&config.fast_mode_state())),
FastModeCommand::Toggle | FastModeCommand::On | FastModeCommand::Off => {
let enabled = match command {
FastModeCommand::Toggle => !config.default_fast_mode,
FastModeCommand::On => true,
FastModeCommand::Off => false,
FastModeCommand::Status => unreachable!(),
};
crate::config::save_fast_mode_preference(&config.root, enabled)?;
config.default_fast_mode = enabled;
Ok(fast_mode_change_message(&config.fast_mode_state()))
}
}
}
fn fast_mode_change_message(state: &FastModeState) -> String {
if state.active {
"fast mode enabled".into()
} else if state.preferred {
format!(
"fast mode preference on; unavailable for {}",
state
.unavailable_reason
.as_deref()
.unwrap_or("this provider/model")
)
} else {
"fast mode off".into()
}
}
fn fast_mode_status(state: &FastModeState) -> String {
if state.active {
"enabled".into()
} else if state.preferred {
format!(
"preferred, unavailable for {}",
state
.unavailable_reason
.as_deref()
.unwrap_or("this provider/model")
)
} else {
"off".into()
}
}
fn model_status_message(config: &Config) -> String {
let model = &config.model;
let state = config.fast_mode_state();
if state.active {
format!("model: {model} · fast enabled")
} else if state.preferred {
format!("model: {model} · fast unavailable")
} else {
format!("model: {model}")
}
}
fn transcript_from_loaded(
conversation: &Conversation,
warning: Option<String>,
@@ -2667,6 +2818,67 @@ mod tests {
);
}
#[test]
fn parse_local_command_accepts_fast_forms() {
assert_eq!(
parse_local_command("/fast").unwrap(),
LocalCommand::Fast(FastModeCommand::Toggle)
);
assert_eq!(
parse_local_command("/fast on").unwrap(),
LocalCommand::Fast(FastModeCommand::On)
);
assert_eq!(
parse_local_command("/fast off").unwrap(),
LocalCommand::Fast(FastModeCommand::Off)
);
assert_eq!(
parse_local_command("/fast status").unwrap(),
LocalCommand::Fast(FastModeCommand::Status)
);
assert_eq!(
parse_local_command("/fast maybe"),
Err("usage: /fast [on|off|status]".into())
);
}
#[test]
fn fast_mode_status_distinguishes_preference_and_activation() {
let mut config = Config {
default_fast_mode: true,
..Config::default()
};
assert_eq!(
fast_mode_status(&config.fast_mode_state()),
"preferred, unavailable for provider fireworks"
);
config.provider_id = config::CHATGPT_CODEX_PROVIDER_ID.into();
config.active_provider.kind = config::CHATGPT_CODEX_PROVIDER_KIND.into();
config.model = config::CHATGPT_CODEX_DEFAULT_MODEL.into();
config.model_metadata = Some(config::ModelDefinition {
id: config::CHATGPT_CODEX_DEFAULT_MODEL.into(),
provider: config::CHATGPT_CODEX_PROVIDER_ID.into(),
display_name: None,
context_length: None,
max_output_tokens: None,
supports_tools: true,
supports_streaming: true,
reasoning: Default::default(),
fast_mode: config::FastModeMetadata { supported: true },
});
assert_eq!(fast_mode_status(&config.fast_mode_state()), "enabled");
assert_eq!(
model_status_message(&config),
format!(
"model: {} · fast enabled",
config::CHATGPT_CODEX_DEFAULT_MODEL
)
);
}
#[test]
fn cancelled_turn_repairs_missing_tool_results() {
let root = tempdir().unwrap();
+88
View File
@@ -27,6 +27,8 @@ pub struct ConfigFile {
pub default_model: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub default_reasoning_effort: Option<ReasoningEffort>,
#[serde(skip_serializing_if = "Option::is_none")]
pub default_fast_mode: Option<bool>,
// Deprecated compatibility fields accepted from older config.json files.
#[serde(skip_serializing_if = "Option::is_none")]
@@ -97,6 +99,8 @@ pub struct ModelDefinition {
pub supports_streaming: bool,
#[serde(default)]
pub reasoning: ReasoningMetadata,
#[serde(default)]
pub fast_mode: FastModeMetadata,
}
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -112,6 +116,13 @@ pub struct ReasoningMetadata {
pub request_format: ReasoningRequestFormat,
}
#[derive(Debug, Clone, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct FastModeMetadata {
#[serde(default)]
pub supported: bool,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum ReasoningEffort {
@@ -145,6 +156,7 @@ pub struct Config {
pub provider_id: String,
pub model: String,
pub reasoning_effort: ReasoningEffort,
pub default_fast_mode: bool,
pub active_provider: ResolvedProviderConfig,
pub model_metadata: Option<ModelDefinition>,
pub default_access_mode: AccessMode,
@@ -157,6 +169,14 @@ pub struct Config {
pub docs_dir: PathBuf,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct FastModeState {
pub preferred: bool,
pub supported: bool,
pub active: bool,
pub unavailable_reason: Option<String>,
}
#[derive(Debug, Clone, Default)]
pub struct ConfigOverrides {
pub model: Option<String>,
@@ -202,6 +222,12 @@ impl Default for ReasoningMetadata {
}
}
impl Default for FastModeMetadata {
fn default() -> Self {
Self { supported: false }
}
}
impl Default for ReasoningRequestFormat {
fn default() -> Self {
Self::ReasoningEffort
@@ -287,6 +313,7 @@ impl Default for Config {
provider_id: DEFAULT_PROVIDER_ID.to_string(),
model: DEFAULT_MODEL.to_string(),
reasoning_effort: ReasoningEffort::Medium,
default_fast_mode: false,
active_provider,
model_metadata: Some(default_model_definition()),
default_access_mode: AccessMode::ReadOnly,
@@ -374,6 +401,9 @@ impl Config {
if let Some(v) = file.confirm_destructive_operations {
cfg.confirm_destructive_operations = v;
}
if let Some(v) = file.default_fast_mode {
cfg.default_fast_mode = v;
}
}
if let Some(access_mode) = overrides.access_mode {
@@ -443,6 +473,41 @@ impl Config {
kind => bail!("unsupported provider kind `{kind}`"),
}
}
pub fn fast_mode_state(&self) -> FastModeState {
let preferred = self.default_fast_mode;
let supported = self.fast_mode_supported();
let active = preferred && supported;
let unavailable_reason = if preferred && !supported {
Some(self.fast_mode_unavailable_reason())
} else {
None
};
FastModeState {
preferred,
supported,
active,
unavailable_reason,
}
}
fn fast_mode_supported(&self) -> bool {
if self.active_provider.kind == CHATGPT_CODEX_PROVIDER_KIND {
return true;
}
self.model_metadata
.as_ref()
.is_some_and(|model| model.fast_mode.supported)
}
fn fast_mode_unavailable_reason(&self) -> String {
if self.active_provider.kind != CHATGPT_CODEX_PROVIDER_KIND {
format!("provider {}", self.provider_id)
} else {
format!("model {}", self.model)
}
}
}
impl ProviderDefinition {
@@ -483,6 +548,28 @@ pub fn save_last_used(root: &Path, model: &str, reasoning_effort: ReasoningEffor
file.default_reasoning_effort = Some(reasoning_effort);
write_json_pretty(&path, &file)
}
pub fn save_last_used_provider(
root: &Path,
provider_id: &str,
model: &str,
reasoning_effort: ReasoningEffort,
) -> Result<()> {
let path = config_path(root);
let mut file = load_config_file(root)?.unwrap_or_default();
file.default_provider = Some(provider_id.to_string());
file.default_model = Some(model.to_string());
file.default_reasoning_effort = Some(reasoning_effort);
write_json_pretty(&path, &file)
}
pub fn save_fast_mode_preference(root: &Path, enabled: bool) -> Result<()> {
let path = config_path(root);
let mut file = load_config_file(root)?.unwrap_or_default();
file.default_fast_mode = Some(enabled);
write_json_pretty(&path, &file)
}
pub fn load_or_create_default_provider_registry(root: &Path) -> Result<ProvidersFile> {
fs::create_dir_all(root).with_context(|| format!("creating {}", root.display()))?;
let path = providers_path(root);
@@ -533,6 +620,7 @@ pub fn default_model_definition() -> ModelDefinition {
supports_tools: true,
supports_streaming: true,
reasoning: ReasoningMetadata::default(),
fast_mode: FastModeMetadata::default(),
}
}
+1 -1
View File
@@ -24,7 +24,7 @@ Make the smallest useful plan, then act. Prefer current project evidence over gu
"## Transcript and tools\n\
Assistant text is streamed to the user. Tool calls, tool results, edit diffs, denials, and approval prompts are visible in the transcript. Request tools directly when they are the right next step; Cassady enforces access policy and shows approval UI separately. Do not ask for chat permission before every tool call, and do not say a tool succeeded before its result arrives. If a tool fails or is denied, adapt instead of repeating the same request.\n\n\
## Tool use\n\
Use tools when the current filesystem or command result matters. Use `ls` for directory orientation, `grep` to locate definitions/usages or inspect large or unknown areas before opening files, `read` for relevant files or ranges, `edit` for focused changes to existing files, `write` for new files or intentional full rewrites, and `shell` for tests, builds, formatting, diagnostics, or project commands when allowed and useful. Prefer targeted inspection and related batched reads over broad exploration. Do not use `shell` for file inspection when `ls`, `grep`, or `read` is safer and sufficient.\n\n\
Use tools when the current filesystem or command result matters. Use `ls` for directory orientation, `grep` to locate definitions/usages or inspect large or unknown areas before opening files, `read` for relevant files or ranges, `edit` for focused changes to existing files, `write` for new files or intentional full rewrites, and `shell` for tests, builds, formatting, diagnostics, or project commands when allowed and useful. Prefer targeted inspection and related batched reads over broad exploration. Treat compacted or truncated tool output as incomplete: re-read a narrower range, search, or rerun a narrower command before editing from omitted details. Do not use `shell` for file inspection when `ls`, `grep`, or `read` is safer and sufficient.\n\n\
## Editing\n\
Inspect before editing. Prefer `edit` for small and medium modifications to existing files. For `edit`, each old text must match exactly and uniquely in the original file; keep replacements minimal, unique, and non-overlapping, and combine related replacements for the same file in one call when practical. Use `write` only for new files or full rewrites where that is safer and intentional. After meaningful code changes, run relevant tests or formatters when allowed, or tell the user what should be run.\n\n\
## Safety and final response\n\
+80 -3
View File
@@ -1,7 +1,7 @@
use super::types::{CompletionResult, ModelMessage};
use crate::agent::AgentEvent;
use crate::codex_auth::load_codex_access_token;
use crate::config::{ReasoningEffort, CHATGPT_CODEX_RESPONSES_URL};
use crate::config::{ReasoningEffort, CHATGPT_CODEX_DEFAULT_MODEL, CHATGPT_CODEX_RESPONSES_URL};
use crate::conversation::StoredToolCall;
use crate::tools::ToolSpec;
use anyhow::{bail, Result};
@@ -17,6 +17,7 @@ pub struct ChatGptCodexProvider {
model: String,
endpoint: String,
reasoning_effort: ReasoningEffort,
fast_mode: bool,
}
#[derive(Debug, Clone)]
@@ -24,6 +25,7 @@ pub struct ChatGptCodexSettings {
pub model: String,
pub endpoint: String,
pub reasoning_effort: ReasoningEffort,
pub fast_mode: bool,
}
#[derive(Debug, Default, Clone)]
@@ -40,6 +42,7 @@ impl ChatGptCodexProvider {
model: settings.model,
endpoint: normalize_endpoint(&settings.endpoint),
reasoning_effort: settings.reasoning_effort,
fast_mode: settings.fast_mode,
}
}
@@ -51,7 +54,13 @@ impl ChatGptCodexProvider {
) -> Result<CompletionResult> {
let token = load_codex_access_token()?;
let secret = token.as_secret().to_string();
let body = responses_body(&self.model, messages, tools, self.reasoning_effort);
let body = responses_body(
&self.model,
messages,
tools,
self.reasoning_effort,
self.fast_mode,
);
let resp = self
.client
.post(&self.endpoint)
@@ -128,6 +137,7 @@ fn responses_body(
messages: Vec<ModelMessage>,
tools: Vec<ToolSpec>,
reasoning_effort: ReasoningEffort,
fast_mode: bool,
) -> Value {
let mut instructions = Vec::new();
let mut input = Vec::new();
@@ -183,12 +193,24 @@ fn responses_body(
if !instructions.is_empty() {
body["instructions"] = Value::String(instructions.join("\n\n"));
}
if let Some(effort) = reasoning_effort.request_value() {
if fast_mode {
body["reasoning"] = json!({"effort": fast_mode_reasoning_effort(model), "summary": "auto"});
} else if let Some(effort) = reasoning_effort.request_value() {
body["reasoning"] = json!({"effort": effort, "summary": "auto"});
} else if reasoning_effort == ReasoningEffort::Off {
body["reasoning"] = json!({"effort": "none", "summary": "auto"});
}
body
}
fn fast_mode_reasoning_effort(model: &str) -> &'static str {
if model == CHATGPT_CODEX_DEFAULT_MODEL {
"low"
} else {
"minimal"
}
}
fn tools_to_responses(tools: Vec<ToolSpec>) -> Vec<Value> {
tools
.into_iter()
@@ -464,6 +486,7 @@ mod tests {
],
Vec::new(),
ReasoningEffort::Off,
false,
);
assert_eq!(body["model"], "gpt-test");
@@ -473,6 +496,60 @@ mod tests {
}));
}
#[test]
fn responses_body_uses_low_reasoning_for_gpt_5_5_fast_mode() {
let body = responses_body(
CHATGPT_CODEX_DEFAULT_MODEL,
vec![ModelMessage::User {
content: "hello".into(),
}],
Vec::new(),
ReasoningEffort::High,
true,
);
assert_eq!(
body["reasoning"],
json!({"effort": "low", "summary": "auto"})
);
}
#[test]
fn responses_body_keeps_minimal_reasoning_for_other_fast_mode_models() {
let body = responses_body(
"gpt-test",
vec![ModelMessage::User {
content: "hello".into(),
}],
Vec::new(),
ReasoningEffort::High,
true,
);
assert_eq!(
body["reasoning"],
json!({"effort": "minimal", "summary": "auto"})
);
}
#[test]
fn responses_body_sends_none_effort_when_reasoning_is_off() {
let body = responses_body(
"gpt-test",
vec![ModelMessage::User {
content: "hello".into(),
}],
Vec::new(),
ReasoningEffort::Off,
false,
);
assert_eq!(
body["reasoning"],
json!({"effort": "none", "summary": "auto"})
);
}
#[test]
fn stream_parser_collects_text_and_function_call() {
let (tx, _rx) = mpsc::unbounded_channel();
+16 -6
View File
@@ -17,23 +17,32 @@ pub enum ProviderClient {
ChatGptCodex(ChatGptCodexProvider),
}
#[derive(Debug, Clone, Copy)]
pub struct ProviderRuntimeOptions {
pub reasoning_effort: ReasoningEffort,
pub fast_mode: bool,
}
impl ProviderClient {
pub fn from_config(config: &Config, reasoning_effort: ReasoningEffort) -> Result<Self> {
pub fn from_config(config: &Config, options: ProviderRuntimeOptions) -> Result<Self> {
match config.active_provider.kind.as_str() {
DEFAULT_PROVIDER_KIND => {
let api_key = config.resolved_api_key()?;
let reasoning_request_format = config
.model_metadata
.as_ref()
let model_metadata = config.model_metadata.as_ref();
let reasoning_request_format = model_metadata
.map(|model| model.reasoning.request_format)
.unwrap_or_default();
let reasoning_supported = model_metadata
.map(|model| model.reasoning.supported)
.unwrap_or(false);
Ok(Self::OpenAiCompatible(OpenAiCompatibleProvider::new(
OpenAiCompatibleSettings {
model: config.model.clone(),
base_url: config.active_provider.base_url.clone(),
api_key,
reasoning_effort,
reasoning_effort: options.reasoning_effort,
reasoning_request_format,
reasoning_supported,
},
)))
}
@@ -41,7 +50,8 @@ impl ProviderClient {
ChatGptCodexSettings {
model: config.model.clone(),
endpoint: config.active_provider.base_url.clone(),
reasoning_effort,
reasoning_effort: options.reasoning_effort,
fast_mode: options.fast_mode,
},
))),
kind => bail!("unsupported provider kind `{kind}`"),
+69 -3
View File
@@ -18,6 +18,7 @@ pub struct OpenAiCompatibleProvider {
api_key: String,
reasoning_effort: ReasoningEffort,
reasoning_request_format: ReasoningRequestFormat,
reasoning_supported: bool,
}
#[derive(Debug, Clone)]
@@ -27,6 +28,7 @@ pub struct OpenAiCompatibleSettings {
pub api_key: String,
pub reasoning_effort: ReasoningEffort,
pub reasoning_request_format: ReasoningRequestFormat,
pub reasoning_supported: bool,
}
#[derive(Debug, Default)]
@@ -45,6 +47,7 @@ impl OpenAiCompatibleProvider {
api_key: settings.api_key,
reasoning_effort: settings.reasoning_effort,
reasoning_request_format: settings.reasoning_request_format,
reasoning_supported: settings.reasoning_supported,
}
}
@@ -65,6 +68,7 @@ impl OpenAiCompatibleProvider {
&mut body,
self.reasoning_effort,
self.reasoning_request_format,
self.reasoning_supported,
);
let resp = self
.client
@@ -219,9 +223,17 @@ fn apply_reasoning_request(
body: &mut Value,
effort: ReasoningEffort,
format: ReasoningRequestFormat,
supported: bool,
) {
let Some(effort) = effort.request_value() else {
if !supported {
return;
}
let effort_str = match effort {
ReasoningEffort::Off => "none",
_ => match effort.request_value() {
Some(value) => value,
None => return,
},
};
let Value::Object(obj) = body else {
return;
@@ -230,11 +242,11 @@ fn apply_reasoning_request(
ReasoningRequestFormat::ReasoningEffort => {
obj.insert(
"reasoning_effort".to_string(),
Value::String(effort.to_string()),
Value::String(effort_str.to_string()),
);
}
ReasoningRequestFormat::ReasoningObject => {
obj.insert("reasoning".to_string(), json!({ "effort": effort }));
obj.insert("reasoning".to_string(), json!({ "effort": effort_str }));
}
}
}
@@ -316,3 +328,57 @@ fn chat_url(base: &str) -> String {
format!("{}/chat/completions", base.trim_end_matches('/'))
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn reasoning_effort_format_sends_none_when_off_and_supported() {
let mut body = json!({"model": "test"});
apply_reasoning_request(
&mut body,
ReasoningEffort::Off,
ReasoningRequestFormat::ReasoningEffort,
true,
);
assert_eq!(body["reasoning_effort"], Value::String("none".to_string()));
}
#[test]
fn reasoning_object_format_sends_none_when_off_and_supported() {
let mut body = json!({"model": "test"});
apply_reasoning_request(
&mut body,
ReasoningEffort::Off,
ReasoningRequestFormat::ReasoningObject,
true,
);
assert_eq!(body["reasoning"], json!({ "effort": "none" }));
}
#[test]
fn reasoning_sends_nothing_when_unsupported_even_if_off() {
let mut body = json!({"model": "test"});
apply_reasoning_request(
&mut body,
ReasoningEffort::Off,
ReasoningRequestFormat::ReasoningEffort,
false,
);
assert!(body.get("reasoning_effort").is_none());
assert!(body.get("reasoning").is_none());
}
#[test]
fn reasoning_sends_nothing_when_unsupported_even_if_high() {
let mut body = json!({"model": "test"});
apply_reasoning_request(
&mut body,
ReasoningEffort::High,
ReasoningRequestFormat::ReasoningObject,
false,
);
assert!(body.get("reasoning").is_none());
}
}
+7 -4
View File
@@ -2,10 +2,10 @@ use crate::check;
use crate::cli::Cli;
use crate::codex_auth;
use crate::config::{
self, ConfigFile, ModelDefinition, ModelsFile, ProviderDefinition, ProvidersFile,
ReasoningEffort, ReasoningMetadata, ReasoningRequestFormat, CHATGPT_CODEX_DEFAULT_MODEL,
CHATGPT_CODEX_PROVIDER_ID, CHATGPT_CODEX_PROVIDER_KIND, CHATGPT_CODEX_PROVIDER_NAME,
CHATGPT_CODEX_RESPONSES_URL, DEFAULT_PROVIDER_KIND,
self, ConfigFile, FastModeMetadata, ModelDefinition, ModelsFile, ProviderDefinition,
ProvidersFile, ReasoningEffort, ReasoningMetadata, ReasoningRequestFormat,
CHATGPT_CODEX_DEFAULT_MODEL, CHATGPT_CODEX_PROVIDER_ID, CHATGPT_CODEX_PROVIDER_KIND,
CHATGPT_CODEX_PROVIDER_NAME, CHATGPT_CODEX_RESPONSES_URL, DEFAULT_PROVIDER_KIND,
};
use crate::menu::{Menu, MenuItem, TextPrompt};
use anyhow::{bail, Context, Result};
@@ -970,6 +970,9 @@ fn upsert_model(models: &mut ModelsFile, selection: &SetupSelection) {
},
request_format: ReasoningRequestFormat::ReasoningEffort,
},
fast_mode: FastModeMetadata {
supported: is_chatgpt_codex_provider(&selection.provider_id),
},
};
if let Some(existing) = models.models.iter_mut().find(|existing| {
+1 -1
View File
@@ -31,7 +31,7 @@ fn default_max() -> usize {
pub fn spec() -> ToolSpec {
ToolSpec {
name: "grep".into(),
description: "Search files or directories for literal text or regex matches. Use before read for large inputs. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
description: "Search files or directories for literal text or regex matches. Use before read for large or unknown inputs; keep queries and paths focused, then read around matching lines. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
parameters: schema::object(json!({
"query": {"type":"string"},
"paths": {"type":"array", "items":{"type":"string"}, "default":["."]},
+5 -1
View File
@@ -227,7 +227,11 @@ fn truncate_model(mut s: String, limit: usize) -> String {
if s.len() <= limit {
return s;
}
s.truncate(limit);
let mut end = limit.min(s.len());
while !s.is_char_boundary(end) {
end = end.saturating_sub(1);
}
s.truncate(end);
s.push_str("\n… truncated by Cass; use grep or narrower line ranges for more.");
s
}
+1 -1
View File
@@ -18,7 +18,7 @@ struct FileArg {
pub fn spec() -> ToolSpec {
ToolSpec {
name: "read".into(),
description: "Read one or more text files, optionally with 1-indexed line ranges like 35-60, 35-, or -60. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
description: "Read one or more text files, optionally with 1-indexed line ranges like 35-60, 35-, or -60. Prefer grep first for unknown locations or large files, and re-read narrower ranges if prior output was compacted or truncated. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
parameters: schema::object(json!({
"files": {
"type":"array",
+1 -1
View File
@@ -16,7 +16,7 @@ struct Args {
pub fn spec() -> ToolSpec {
ToolSpec {
name: "shell".into(),
description: "Run a shell command in the launch cwd. Request this tool directly when shell is useful; do not ask the user for permission in chat. Cass may show a separate approval UI before execution depending on the active access mode. Streams stdout/stderr while running, then returns stdout, stderr, and exit code. Use timeout (seconds) to limit runtime."
description: "Run a shell command in the launch cwd. Request this tool directly when shell is useful; do not ask the user for permission in chat. Cass may show a separate approval UI before execution depending on the active access mode. Streams stdout/stderr while running, then returns stdout, stderr, and exit code. Use timeout (seconds) to limit runtime. Prefer narrow commands or filtering (grep/head/tail) for broad output, and rerun narrower commands if prior output was compacted or truncated."
.into(),
parameters: schema::object(
json!({
+40 -5
View File
@@ -53,6 +53,7 @@ pub struct RenderState<'a> {
pub show_full_tools: bool,
pub show_reasoning: bool,
pub reasoning_effort: ReasoningEffort,
pub fast_mode_active: bool,
pub scroll: u16,
pub autofill: Option<&'a AutoFillMenu>,
pub overlay: Option<&'a OverlayView>,
@@ -673,11 +674,26 @@ fn collapsed_tool_summary(content: &str) -> String {
return "no output".into();
}
let lines = content.lines().count();
format!(
"{} · {} · tool output hidden",
pluralize(lines, "line"),
human_bytes(content.len())
)
let mut parts = vec![pluralize(lines, "line"), human_bytes(content.len())];
if let Some(marker) = tool_incompleteness_marker(content) {
parts.push(marker.into());
}
parts.push("tool output hidden".into());
parts.join(" · ")
}
fn tool_incompleteness_marker(content: &str) -> Option<&'static str> {
if content.contains("Cass compacted this tool output") {
Some("compacted")
} else if content.contains("truncated by Cass")
|| content.contains("Cass truncated this tool output")
{
Some("truncated")
} else if content.contains("… stopped after ") {
Some("stopped early")
} else {
None
}
}
fn pluralize(count: usize, unit: &str) -> String {
@@ -724,6 +740,9 @@ fn footer_text(state: &RenderState<'_>) -> String {
parts.push("tools:full".into());
}
parts.push(format!("reasoning:{}", state.reasoning_effort));
if state.fast_mode_active {
parts.push("fast".into());
}
if state.show_reasoning {
parts.push("reasoning:visible".into());
}
@@ -895,6 +914,22 @@ mod tests {
assert!(!text.contains("one\ntwo\nthree"));
}
#[test]
fn collapsed_tool_output_marks_incomplete_results() {
let transcript = vec![TranscriptBlock {
kind: TranscriptKind::Tool,
title: "grep ✓ (call_1)".into(),
content: "match\n… stopped after 1 matches. Narrow the query or raise max_matches."
.into(),
}];
let rendered = transcript_lines_from(&transcript, false, false);
let text = rendered_text(&rendered);
assert!(text.contains("stopped early"));
assert!(text.contains("tool output hidden"));
}
#[test]
fn successful_ls_shows_summary_when_tools_are_collapsed() {
let transcript = vec![TranscriptBlock {
+136
View File
@@ -123,6 +123,62 @@ async fn reasoning_effort_supports_reasoning_object_format() {
.unwrap();
}
#[tokio::test]
async fn fast_mode_preference_does_not_change_openai_compatible_request() {
let server = MockServer::start().await;
Mock::given(method("POST"))
.and(path("/chat/completions"))
.respond_with(sse(
"data: {\"choices\":[{\"index\":0,\"delta\":{\"content\":\"Done.\"}}]}\r\n\r\ndata: [DONE]\r\n\r\n",
))
.expect(1)
.mount(&server)
.await;
let root = tempdir().unwrap();
let cwd = tempdir().unwrap();
let docs = tempdir().unwrap();
let config = Config {
root: root.path().to_path_buf(),
docs_dir: docs.path().to_path_buf(),
model: "test-model".into(),
default_fast_mode: true,
active_provider: cassady::config::ResolvedProviderConfig {
base_url: server.uri(),
api_key: "test-key".into(),
..Config::default().active_provider
},
..Config::default()
};
let conversation = Conversation::create(
&config.conversations_dir(),
&config.model,
cwd.path(),
"base prompt".into(),
)
.unwrap();
let (tx, _rx) = mpsc::unbounded_channel::<AgentEvent>();
run_turn(
conversation,
"stay compatible".into(),
AgentSettings {
config,
cwd: cwd.path().to_path_buf(),
mode: AccessMode::ReadOnly,
reasoning_effort: ReasoningEffort::Off,
},
tx,
)
.await
.unwrap();
let requests = server.received_requests().await.unwrap();
let body = String::from_utf8_lossy(&requests[0].body);
assert!(!body.contains("\"effort\":\"minimal\""));
assert!(!body.contains("\"fast_mode\""));
}
#[tokio::test]
async fn reasoning_is_streamed_persisted_and_sent_back() {
let server = MockServer::start().await;
@@ -312,6 +368,86 @@ async fn empty_final_response_is_reprompted_and_persisted() {
));
}
#[tokio::test]
async fn tool_results_are_stored_full_but_sent_to_model_with_limit_guidance() {
let server = MockServer::start().await;
Mock::given(method("POST"))
.and(path("/chat/completions"))
.and(body_string_contains("Cass truncated this tool output"))
.and(body_string_contains("large.txt lines 1-200"))
.respond_with(sse(
"data: {\"choices\":[{\"index\":0,\"delta\":{\"content\":\"Done.\"}}]}\r\n\r\ndata: [DONE]\r\n\r\n",
))
.with_priority(1)
.expect(1)
.mount(&server)
.await;
Mock::given(method("POST"))
.and(path("/chat/completions"))
.respond_with(tool_call_sse(
"call_read",
"read",
r#"{"files":[{"path":"large.txt"}]}"#,
))
.with_priority(10)
.expect(1)
.mount(&server)
.await;
let root = tempdir().unwrap();
let cwd = tempdir().unwrap();
let docs = tempdir().unwrap();
let large = (1..=200)
.map(|line| format!("line {line}"))
.collect::<Vec<_>>()
.join("\n");
std::fs::write(cwd.path().join("large.txt"), large).unwrap();
let config = Config {
root: root.path().to_path_buf(),
docs_dir: docs.path().to_path_buf(),
model: "test-model".into(),
model_tool_result_limit: 180,
active_provider: cassady::config::ResolvedProviderConfig {
base_url: server.uri(),
api_key: "test-key".into(),
..Config::default().active_provider
},
..Config::default()
};
let conversation = Conversation::create(
&config.conversations_dir(),
&config.model,
cwd.path(),
"base prompt".into(),
)
.unwrap();
let (tx, _rx) = mpsc::unbounded_channel::<AgentEvent>();
let updated = run_turn(
conversation,
"read the large file".into(),
AgentSettings {
config,
cwd: cwd.path().to_path_buf(),
mode: AccessMode::ReadOnly,
reasoning_effort: ReasoningEffort::Off,
},
tx,
)
.await
.unwrap();
assert!(updated.records.iter().any(|record| matches!(
record,
Record::Tool { name, content, .. }
if name == "read"
&& content.contains("line 200")
&& !content.contains("Cass truncated this tool output")
)));
}
#[tokio::test]
async fn workspace_edit_shell_does_not_execute_until_approved() {
let server = MockServer::start().await;
+170
View File
@@ -47,6 +47,7 @@ fn default_provider_and_model_files_are_created() {
.unwrap();
assert_eq!(models.models[0].provider, "fireworks");
assert!(models.models[0].reasoning.supported);
assert!(!models.models[0].fast_mode.supported);
assert_eq!(
models.models[0].reasoning.default_effort,
ReasoningEffort::Medium
@@ -127,6 +128,174 @@ fn reasoning_defaults_to_supported_medium_for_model_metadata() {
);
}
#[test]
fn fast_mode_defaults_to_off_and_unsupported() {
let model: config::ModelDefinition = serde_json::from_str(
r#"{
"id": "test-model",
"provider": "test-provider"
}
"#,
)
.unwrap();
assert!(!model.fast_mode.supported);
let root = tempdir().unwrap();
let cfg = Config::load_from_root_with_docs(
root.path().to_path_buf(),
root.path().join("docs"),
&cli(),
)
.unwrap();
let state = cfg.fast_mode_state();
assert!(!state.preferred);
assert!(!state.supported);
assert!(!state.active);
}
#[test]
fn fast_mode_preference_persists_without_losing_config_fields() {
let root = tempdir().unwrap();
std::fs::write(
root.path().join("config.json"),
r#"{
"default_access_mode": "workspace-edit",
"show_reasoning": true
}
"#,
)
.unwrap();
config::save_fast_mode_preference(root.path(), true).unwrap();
let cfg = Config::load_from_root_with_docs(
root.path().to_path_buf(),
root.path().join("docs"),
&cli(),
)
.unwrap();
assert!(cfg.default_fast_mode);
assert!(cfg.show_reasoning);
assert_eq!(cfg.default_access_mode.to_string(), "workspace-edit");
}
#[test]
fn fast_mode_state_is_active_for_supported_codex_model() {
let root = tempdir().unwrap();
std::fs::write(
root.path().join("providers.json"),
r#"{
"providers": [
{
"id": "chatgpt-codex",
"name": "ChatGPT Codex",
"kind": "chatgpt-codex",
"base_url": "https://chatgpt.com/backend-api/codex/responses",
"api_key": "",
"default_model": "gpt-5.5",
"models": ["gpt-5.5"]
}
]
}
"#,
)
.unwrap();
std::fs::write(
root.path().join("models.json"),
r#"{
"models": [
{
"id": "gpt-5.5",
"provider": "chatgpt-codex",
"fast_mode": { "supported": true }
}
]
}
"#,
)
.unwrap();
std::fs::write(
root.path().join("config.json"),
r#"{
"default_provider": "chatgpt-codex",
"default_model": "gpt-5.5",
"default_fast_mode": true
}
"#,
)
.unwrap();
let cfg = Config::load_from_root_with_docs(
root.path().to_path_buf(),
root.path().join("docs"),
&cli(),
)
.unwrap();
let state = cfg.fast_mode_state();
assert!(state.preferred);
assert!(state.supported);
assert!(state.active);
}
#[test]
fn fast_mode_state_is_active_for_chatgpt_codex_even_with_legacy_metadata() {
let root = tempdir().unwrap();
std::fs::write(
root.path().join("providers.json"),
r#"{
"providers": [
{
"id": "chatgpt-codex",
"name": "ChatGPT Codex",
"kind": "chatgpt-codex",
"base_url": "https://chatgpt.com/backend-api/codex/responses",
"api_key": "",
"default_model": "gpt-5.5",
"models": ["gpt-5.5"]
}
]
}
"#,
)
.unwrap();
std::fs::write(
root.path().join("models.json"),
r#"{
"models": [
{
"id": "gpt-5.5",
"provider": "chatgpt-codex",
"fast_mode": { "supported": false }
}
]
}
"#,
)
.unwrap();
std::fs::write(
root.path().join("config.json"),
r#"{
"default_provider": "chatgpt-codex",
"default_model": "gpt-5.5",
"default_fast_mode": true
}
"#,
)
.unwrap();
let cfg = Config::load_from_root_with_docs(
root.path().to_path_buf(),
root.path().join("docs"),
&cli(),
)
.unwrap();
let state = cfg.fast_mode_state();
assert!(state.preferred);
assert!(state.supported);
assert!(state.active);
}
#[test]
fn validation_accepts_chatgpt_codex_without_api_key() {
let providers = ProvidersFile {
@@ -150,6 +319,7 @@ fn validation_accepts_chatgpt_codex_without_api_key() {
supports_tools: true,
supports_streaming: true,
reasoning: Default::default(),
fast_mode: Default::default(),
}],
};
+16
View File
@@ -1,5 +1,6 @@
use cassady::access::AccessMode;
use cassady::prompt::{build_base_system_prompt, build_effective_system_prompt};
use cassady::tools;
use std::path::Path;
fn approximate_token_count(s: &str) -> usize {
@@ -44,6 +45,8 @@ fn base_prompt_has_required_sections_without_runtime_context() {
"Tool calls, tool results, edit diffs, denials, and approval prompts are visible"
));
assert!(prompt.contains("each old text must match exactly and uniquely"));
assert!(prompt.contains("compacted or truncated tool output as incomplete"));
assert!(prompt.contains("re-read a narrower range"));
assert!(prompt.contains("End every turn with a concise user-facing response"));
assert!(!prompt.contains("## Runtime context"));
assert!(!prompt.contains("Model:"));
@@ -131,6 +134,19 @@ fn runtime_constraints_stay_after_global_instructions() {
assert!(access_index < authority_index);
}
#[test]
fn tool_specs_bias_toward_narrow_reinspection() {
let specs = tools::specs(AccessMode::FullAccess);
let read = specs.iter().find(|spec| spec.name == "read").unwrap();
let grep = specs.iter().find(|spec| spec.name == "grep").unwrap();
let shell = specs.iter().find(|spec| spec.name == "shell").unwrap();
assert!(read.description.contains("Prefer grep first"));
assert!(read.description.contains("compacted or truncated"));
assert!(grep.description.contains("large or unknown inputs"));
assert!(shell.description.contains("grep/head/tail"));
}
#[test]
fn effective_prompt_size_remains_intentional() {
for mode in [
+5
View File
@@ -174,6 +174,11 @@ fn apply_setup_writes_chatgpt_codex_without_api_key() {
assert_eq!(provider.id, config::CHATGPT_CODEX_PROVIDER_ID);
assert_eq!(provider.kind, config::CHATGPT_CODEX_PROVIDER_KIND);
assert!(provider.api_key.is_empty());
let models: ModelsFile =
serde_json::from_str(&std::fs::read_to_string(root.path().join("models.json")).unwrap())
.unwrap();
assert!(models.models[0].fast_mode.supported);
}
#[test]