Prepare Cassady v0.3.5
CI / Build (push) Waiting to run
CI / Test (push) Waiting to run

This commit is contained in:
2026-06-26 11:53:01 -05:00
parent ac86ee3933
commit b4ff4f9f14
19 changed files with 683 additions and 47 deletions
Generated
+1 -1
View File
@@ -226,7 +226,7 @@ checksum = "8ae3f5d315924270530207e2a68396c3cc547f6dca3fbdca317cfb1a51edb593"
[[package]]
name = "cassady"
version = "0.3.4"
version = "0.3.5"
dependencies = [
"anyhow",
"async-trait",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "cassady"
version = "0.3.4"
version = "0.3.5"
edition = "2021"
description = "Cassady/Cass minimal terminal coding agent"
license = "MIT"
+6
View File
@@ -108,6 +108,12 @@ Cassady exposes tools according to the active access mode:
Use `--readonly`, `--workspace-edit`, or `--full-access` to choose a mode at launch, or press `Shift-Tab` while idle.
## Tool output and context recovery
Cassady keeps tool calls reviewable while fitting provider context windows. Very large tool results may be sent to the model as compacted or truncated head/tail excerpts with a note that names the tool, retained excerpt shape, file ranges or command provenance when available, and suggested follow-up reads/searches. Treat those notes as incomplete evidence: ask Cass to re-read a narrower line range, run a focused `grep`, or rerun a narrower shell command before editing from omitted details.
Conversation files keep the recorded tool result content; model-facing compaction happens when preparing provider messages. The compact/full tool-output UI toggle only changes display.
## Branch and restore
Press `Esc` twice while idle, or type `/branch`, to browse the current conversation's branch family. Selecting an earlier user message, assistant message, tool call, or tool result creates a new branch conversation instead of truncating the original chat. The menu also lets you switch back to related branches later.
+5 -5
View File
@@ -6,27 +6,27 @@ This release focuses on making large tool outputs easier for the assistant to re
### Model Context Recovery
- [ ] **Improve compacted tool-output guidance.** Replace generic head/tail compaction notices with actionable guidance that tells the assistant what was omitted and how to inspect it again safely.
- [x] **Improve compacted tool-output guidance.** Replace generic head/tail compaction notices with actionable guidance that tells the assistant what was omitted and how to inspect it again safely.
- Include tool name, output size, retained excerpt shape, and suggested narrower follow-up reads or searches when available.
- Keep model-facing guidance concise enough that it does not worsen context pressure.
- [ ] **Preserve targeted reinspection metadata.** Track enough structured context for large reads and command output so the assistant can recover omitted details without repeating broad requests.
- [x] **Preserve targeted reinspection metadata.** Track enough structured context for large reads and command output so the assistant can recover omitted details without repeating broad requests.
- For file reads, preserve path and line-range coverage even after compaction.
- For shell and search output, prefer guidance toward narrower commands or `grep`/`read` follow-ups rather than blindly rerunning the same broad command.
### Tool Behavior and Prompting
- [ ] **Bias tool use toward smaller inspections.** Update tool descriptions, prompt guidance, and result messages so broad reads become a fallback rather than the default.
- [x] **Bias tool use toward smaller inspections.** Update tool descriptions, prompt guidance, and result messages so broad reads become a fallback rather than the default.
- Encourage search-first workflows for large files and unknown locations.
- Mention result limits before or at truncation points so the assistant knows when context may be incomplete.
- [ ] **Make truncation and compaction visible across layers.** Align model-facing messages, stored conversation records, and UI summaries so users and the assistant can tell when output was incomplete.
- [x] **Make truncation and compaction visible across layers.** Align model-facing messages, stored conversation records, and UI summaries so users and the assistant can tell when output was incomplete.
- Do not let UI-only collapsed output change what is stored or sent to the model.
- Keep existing conversation files readable and resumable.
### Validation
- [ ] **Add regression coverage for broad-output recovery.** Test workflows where an early broad read or command output is compacted before the assistant needs exact context for an edit.
- [x] **Add regression coverage for broad-output recovery.** Test workflows where an early broad read or command output is compacted before the assistant needs exact context for an edit.
- Cover superseded reads, compacted non-newest tool outputs, provider-message validity, and suggested follow-up guidance.
- Verify `cargo fmt` and `cargo test --locked --all-targets` pass before handoff.
+1 -1
View File
@@ -60,7 +60,7 @@ Fields:
- `default_fast_mode`: optional boolean, defaults to `false`. When `true`, Cassady requests faster inference only for provider/model combinations that advertise fast-mode support.
- `default_access_mode`: `"read-only"`, `"workspace-edit"`, or `"full-access"`.
- `context_message_limit`: optional legacy upper bound for recent non-system messages. Cassady primarily budgets context from model metadata and trims along valid tool-call boundaries.
- `model_tool_result_limit`: optional max bytes of tool output sent back to the model.
- `model_tool_result_limit`: optional approximate max characters of each tool output sent back to the model. Larger results are model-facing head/tail excerpts with recovery guidance; conversation records keep the tool result content.
- `ui_tool_result_limit`: optional max bytes of tool output shown in the UI unless full output is toggled.
- `show_reasoning`: optional boolean, defaults to `false`. Shows provider-streamed reasoning in the transcript.
- `confirm_destructive_operations`: optional compatibility preference currently stored in config.
+4
View File
@@ -12,6 +12,8 @@
**Config root**: The `~/.cass` directory containing config, conversations, global instructions, and installed docs.
**Compacted tool output**: A model-facing replacement for a large tool result that keeps a head/tail excerpt plus provenance and recovery guidance so the assistant can re-read or re-search narrowly before relying on omitted details.
**Exact edit**: An `edit` tool replacement where each `old_text` must match exactly once in the original file before anything is written.
**Fast mode**: A saved preference enabled with `/fast`. It is active only when the current provider/model advertises fast-mode support; otherwise Cassady keeps the preference but reports it as unavailable.
@@ -28,4 +30,6 @@
**Tool call**: A model-requested operation such as `ls`, `read`, `grep`, `write`, `edit`, or `shell`.
**Truncated tool output**: A model-facing shortened tool result produced when output exceeds `model_tool_result_limit`. Cassady tells the model that output was incomplete and suggests narrower follow-up inspection.
**Workspace**: The launch cwd, either the current directory or the path passed with `--cwd`. In workspace-edit mode, writes must stay inside this root.
+8
View File
@@ -136,6 +136,14 @@ Likely cause: the command itself failed, the working directory is wrong, depende
Fix: inspect stdout/stderr, verify cwd in `/status`, and ask Cassady to rerun the smallest relevant command.
## Tool output was compacted or truncated
Symptom: a tool result note says Cassady compacted or truncated output, shows retained head/tail excerpts, or says a search stopped after a match limit.
Likely cause: the raw tool output exceeded the model-facing result limit or the active model context budget.
Fix: treat omitted output as incomplete. Ask Cassady to re-read the exact file line range named in the note, run a narrower `grep` query/path, lower `max_matches`, or rerun a shell command with a more focused flag/filter before making edits based on omitted lines.
## Exact-text edit failed
Symptom: edit reports `old_text not found`, `old_text is not unique`, or overlapping edits.
+19 -1
View File
@@ -29,7 +29,25 @@ Start in read-only mode or press `Shift-Tab` until the status shows `read-only`.
Find where configuration is loaded and summarize the precedence rules.
```
Cassady can use `ls`, `read`, and `grep` to inspect the workspace and bundled docs.
Cassady can use `ls`, `read`, and `grep` to inspect the workspace and bundled docs. For large files or unknown locations, prefer a search-first flow: `grep` for a symbol or phrase, then `read` a small line range around the relevant match.
## Recover from compacted or truncated output
When a tool result is too large for the model context, Cassady sends the model a head/tail excerpt with a recovery note. The note includes the tool name, retained excerpt shape, and file range or shell-command provenance when available.
If Cassady reports compacted or truncated output, do not rely on omitted lines for edits. Ask Cass to narrow the inspection instead:
```text
Re-read src/app.rs lines 220-280 before editing that function.
```
```text
Search only src/ for "load_config" and then read around the matching lines.
```
```text
Rerun the test command with a focused package/filter, or pipe the noisy output through grep/head/tail.
```
## Apply a focused edit
+467 -21
View File
@@ -113,7 +113,9 @@ pub async fn run_turn_with_commands(
cwd: settings.cwd.clone(),
read_roots: vec![settings.cwd.clone(), docs_dir.clone()],
blocked_write_roots: vec![docs_dir.clone()],
model_result_limit: settings.config.model_tool_result_limit,
// Keep stored/UI tool results intact; build_messages applies the
// model-facing result limit when preparing provider messages.
model_result_limit: usize::MAX,
runtime_tx: None,
};
@@ -476,6 +478,7 @@ fn build_messages(records: &[Record], system: String, config: &Config) -> Vec<Mo
messages.extend(records.iter().filter_map(record_to_model_message));
messages = sanitize_tool_message_structure(messages);
supersede_old_read_outputs(&mut messages);
apply_model_tool_result_limits(&mut messages, config.model_tool_result_limit);
let budget = context_budget_tokens(config);
if estimate_messages_tokens(&messages) > budget {
@@ -518,7 +521,7 @@ fn record_to_model_message(record: &Record) -> Option<ModelMessage> {
}
fn supersede_old_read_outputs(messages: &mut [ModelMessage]) {
let read_calls = read_tool_calls_by_id(messages);
let read_calls = tool_calls_by_id(messages);
let mut read_outputs = Vec::new();
for (message_idx, message) in messages.iter().enumerate() {
let ModelMessage::Tool {
@@ -579,16 +582,14 @@ fn tool_content(messages: &[ModelMessage], idx: usize) -> &str {
}
}
fn read_tool_calls_by_id(messages: &[ModelMessage]) -> BTreeMap<String, StoredToolCall> {
fn tool_calls_by_id(messages: &[ModelMessage]) -> BTreeMap<String, StoredToolCall> {
let mut calls = BTreeMap::new();
for message in messages {
let ModelMessage::Assistant { tool_calls, .. } = message else {
continue;
};
for call in tool_calls {
if call.name == "read" {
calls.insert(call.id.clone(), call.clone());
}
calls.insert(call.id.clone(), call.clone());
}
}
calls
@@ -830,43 +831,134 @@ fn context_budget_tokens(config: &Config) -> usize {
.max(MIN_INPUT_BUDGET_TOKENS)
}
#[derive(Debug, Clone, Copy)]
enum ToolOutputTransform {
Compaction,
Truncation,
}
impl ToolOutputTransform {
fn verb(self) -> &'static str {
match self {
ToolOutputTransform::Compaction => "compacted",
ToolOutputTransform::Truncation => "truncated",
}
}
fn reason(self) -> &'static str {
match self {
ToolOutputTransform::Compaction => "to fit the model context",
ToolOutputTransform::Truncation => {
"before sending it to the model because it exceeded the configured model_tool_result_limit"
}
}
}
}
fn apply_model_tool_result_limits(messages: &mut [ModelMessage], limit_chars: usize) {
let tool_calls = tool_calls_by_id(messages);
let target_chars = limit_chars.max(1);
for idx in 1..messages.len() {
transform_tool_output_at(
messages,
idx,
target_chars,
ToolOutputTransform::Truncation,
&tool_calls,
);
}
}
fn compact_tool_outputs(messages: &mut [ModelMessage], budget: usize) {
let newest_tool_idx = messages
.iter()
.rposition(|message| matches!(message, ModelMessage::Tool { .. }));
let tool_calls = tool_calls_by_id(messages);
for idx in 1..messages.len() {
if Some(idx) == newest_tool_idx {
continue;
}
compact_tool_output_at(messages, idx, TOOL_OUTPUT_COMPACT_CHARS);
transform_tool_output_at(
messages,
idx,
TOOL_OUTPUT_COMPACT_CHARS,
ToolOutputTransform::Compaction,
&tool_calls,
);
if estimate_messages_tokens(messages) <= budget {
return;
}
}
for idx in 1..messages.len() {
compact_tool_output_at(messages, idx, TOOL_OUTPUT_TINY_CHARS);
transform_tool_output_at(
messages,
idx,
TOOL_OUTPUT_TINY_CHARS,
ToolOutputTransform::Compaction,
&tool_calls,
);
if estimate_messages_tokens(messages) <= budget {
return;
}
}
}
fn compact_tool_output_at(messages: &mut [ModelMessage], idx: usize, target_chars: usize) {
let Some(ModelMessage::Tool { content, .. }) = messages.get_mut(idx) else {
return;
fn transform_tool_output_at(
messages: &mut [ModelMessage],
idx: usize,
target_chars: usize,
transform: ToolOutputTransform,
tool_calls: &BTreeMap<String, StoredToolCall>,
) {
let replacement = {
let Some(ModelMessage::Tool {
tool_call_id,
name,
content,
}) = messages.get(idx)
else {
return;
};
if content.chars().count() <= target_chars.max(1) {
return;
}
format_tool_output_transform(
name,
tool_calls.get(tool_call_id),
content,
target_chars,
transform,
)
};
if content.chars().count() <= target_chars {
return;
if let Some(ModelMessage::Tool { content, .. }) = messages.get_mut(idx) {
*content = replacement;
}
*content = compact_text(content, target_chars);
}
fn compact_text(content: &str, target_chars: usize) -> String {
fn format_tool_output_transform(
tool_name: &str,
call: Option<&StoredToolCall>,
content: &str,
target_chars: usize,
transform: ToolOutputTransform,
) -> String {
let original_chars = content.chars().count();
let head_chars = (target_chars * 2 / 3).max(1);
let tail_chars = target_chars.saturating_sub(head_chars).max(1);
if original_chars <= target_chars.max(1) {
return content.to_string();
}
let target_chars = target_chars
.max(1)
.min(original_chars.saturating_sub(1).max(1));
let head_chars = if target_chars <= 1 {
1
} else {
(target_chars * 2 / 3).clamp(1, target_chars - 1)
};
let tail_chars = target_chars.saturating_sub(head_chars);
let head: String = content.chars().take(head_chars).collect();
let tail: String = content
.chars()
@@ -876,11 +968,208 @@ fn compact_text(content: &str, target_chars: usize) -> String {
.into_iter()
.rev()
.collect();
let notice = tool_output_transform_notice(
tool_name,
call,
content,
original_chars,
head_chars,
tail_chars,
transform,
);
let mut out = String::new();
out.push_str(&notice);
out.push_str("\n--- retained head excerpt ---\n");
out.push_str(&head);
out.push_str("\n--- omitted middle ---\n");
if tail_chars > 0 {
out.push_str("--- retained tail excerpt ---\n");
out.push_str(&tail);
}
out
}
fn tool_output_transform_notice(
tool_name: &str,
call: Option<&StoredToolCall>,
content: &str,
original_chars: usize,
head_chars: usize,
tail_chars: usize,
transform: ToolOutputTransform,
) -> String {
let retained_shape = if tail_chars > 0 {
format!("{head_chars} chars from the start and {tail_chars} chars from the end")
} else {
format!("{head_chars} chars from the start")
};
let mut sentences = vec![format!(
"Cass {} this tool output from {original_chars} chars {}. Tool: `{tool_name}`. Retained excerpt shape: {retained_shape}.",
transform.verb(),
transform.reason(),
)];
if let Some(provenance) = tool_output_provenance_sentence(tool_name, call, content) {
sentences.push(provenance);
}
sentences.push(tool_output_recovery_guidance(tool_name).to_string());
format!("[{}]", sentences.join(" "))
}
fn tool_output_provenance_sentence(
tool_name: &str,
call: Option<&StoredToolCall>,
content: &str,
) -> Option<String> {
match tool_name {
"read" => read_output_provenance_sentence(call, content),
"grep" => grep_output_provenance_sentence(content),
"shell" => shell_output_provenance_sentence(call),
_ => None,
}
}
fn read_output_provenance_sentence(call: Option<&StoredToolCall>, content: &str) -> Option<String> {
let request_specs = call
.filter(|call| call.name == "read")
.map(|call| read_request_specs(&call.arguments))
.unwrap_or_default();
let sections = parse_read_output_sections(content, &request_specs);
if !sections.is_empty() {
return Some(format!(
"Omitted content came from {}.",
read_sections_summary(&sections)
));
}
read_request_summary(call).map(|summary| format!("The read request targeted {summary}."))
}
fn read_sections_summary(sections: &[ReadOutputSection]) -> String {
let labels: Vec<String> = sections.iter().take(2).map(read_section_label).collect();
match sections.len() {
0 => "no read sections".into(),
1 => labels[0].clone(),
2 => format!("2 read sections: {} and {}", labels[0], labels[1]),
count => format!(
"{count} read sections including {} and {}",
labels[0], labels[1]
),
}
}
fn read_section_label(section: &ReadOutputSection) -> String {
format!(
"[Cass compacted this tool output from {original_chars} chars to fit the model context. Head/tail excerpt follows.]\n{head}\n… omitted …\n{tail}"
"{} lines {}-{}",
section.path, section.start_line, section.end_line
)
}
fn read_request_summary(call: Option<&StoredToolCall>) -> Option<String> {
let call = call.filter(|call| call.name == "read")?;
let mut labels = Vec::new();
if let Some(files) = call
.arguments
.get("files")
.and_then(|files| files.as_array())
{
for file in files.iter().take(2) {
let Some(path) = file.get("path").and_then(|path| path.as_str()) else {
continue;
};
labels.push(read_request_label(
path,
file.get("lines").and_then(|lines| lines.as_str()),
));
}
return match labels.len() {
0 => None,
1 => labels.first().cloned(),
2 if files.len() == 2 => Some(format!("2 files: {} and {}", labels[0], labels[1])),
_ => Some(format!(
"{} files including {} and {}",
files.len(),
labels[0],
labels[1]
)),
};
}
call.arguments
.get("path")
.and_then(|path| path.as_str())
.map(|path| {
read_request_label(
path,
call.arguments.get("lines").and_then(|lines| lines.as_str()),
)
})
}
fn read_request_label(path: &str, lines: Option<&str>) -> String {
match lines.map(str::trim).filter(|lines| !lines.is_empty()) {
Some(lines) => format!("{path} lines {lines}"),
None => path.to_string(),
}
}
fn grep_output_provenance_sentence(content: &str) -> Option<String> {
grep_stopped_after(content)
.map(|count| format!("The grep output reported it stopped after {count} matches."))
}
fn grep_stopped_after(content: &str) -> Option<usize> {
content.lines().find_map(|line| {
let rest = line.trim().strip_prefix("… stopped after ")?;
rest.split_whitespace().next()?.parse().ok()
})
}
fn shell_output_provenance_sentence(call: Option<&StoredToolCall>) -> Option<String> {
let call = call.filter(|call| call.name == "shell")?;
let command = call.arguments.get("command")?.as_str()?.trim();
if command.is_empty() {
return None;
}
let preview = command.split_whitespace().collect::<Vec<_>>().join(" ");
if command_preview_may_contain_sensitive_text(&preview) {
return Some(
"Shell command preview omitted because it may contain sensitive text; inspect the preceding tool-call arguments before rerunning."
.into(),
);
}
let chars = preview.chars().count();
if chars > 160 {
return Some(format!(
"Shell command was {chars} chars; inspect the preceding tool-call arguments before rerunning."
));
}
Some(format!("Shell command: `{preview}`."))
}
fn command_preview_may_contain_sensitive_text(command: &str) -> bool {
let lowered = command.to_ascii_lowercase();
[
"api_key",
"apikey",
"authorization",
"bearer",
"password",
"secret",
"token",
]
.iter()
.any(|needle| lowered.contains(needle))
}
fn tool_output_recovery_guidance(tool_name: &str) -> &'static str {
match tool_name {
"read" => "Recovery: use `read` with a narrower line range, or `grep` for a symbol before reading, before relying on omitted details.",
"grep" => "Recovery: rerun `grep` with a narrower query/path, lower `max_matches`, or `read` around specific matching lines before relying on omitted details.",
"shell" => "Recovery: rerun a narrower command, filter output with grep/head/tail, or inspect specific files named in the excerpt before making edits from omitted lines.",
_ => "Recovery: rerun a narrower tool request or inspect the specific file/range from the excerpt before relying on omitted details.",
}
}
fn trim_to_context_budget(mut messages: Vec<ModelMessage>, budget: usize) -> Vec<ModelMessage> {
let mut omitted = false;
while estimate_messages_tokens(&messages) > budget && messages.len() > 1 {
@@ -894,7 +1183,7 @@ fn trim_to_context_budget(mut messages: Vec<ModelMessage>, budget: usize) -> Vec
if omitted {
let note = ModelMessage::System {
content: "Cass omitted earlier conversation messages to fit the model context budget. Included tool results still follow their matching assistant tool calls; older large tool outputs may be compacted.".to_string(),
content: "Cass omitted earlier conversation messages to fit the model context budget. Included tool results still follow their matching assistant tool calls; older large tool outputs may be compacted with recovery guidance.".to_string(),
};
messages.insert(1, note);
while estimate_messages_tokens(&messages) > budget && messages.len() > 2 {
@@ -1072,7 +1361,7 @@ fn _calls(_calls: Vec<StoredToolCall>) {}
mod tests {
use super::*;
use crate::config::default_model_definition;
use serde_json::json;
use serde_json::{json, Value};
fn small_config(context_length: u64, max_output_tokens: u64) -> Config {
let mut config = Config::default();
@@ -1099,6 +1388,22 @@ mod tests {
}
}
fn named_call(id: &str, name: &str, arguments: Value) -> StoredToolCall {
StoredToolCall {
id: id.to_string(),
name: name.to_string(),
arguments,
}
}
fn numbered_read_output(path: &str, start: usize, end: usize) -> String {
let mut out = format!("--- {path} lines {start}-{end} ---\n");
for line in start..=end {
out.push_str(&format!("{line:>6} | line {line}\n"));
}
out
}
fn assert_valid_tool_structure(messages: &[ModelMessage]) {
let mut idx = 1;
while idx < messages.len() {
@@ -1153,7 +1458,9 @@ mod tests {
ts: now_ts(),
},
];
let messages = build_messages(&records, "system".into(), &small_config(8_000, 512));
let mut config = small_config(8_000, 512);
config.model_tool_result_limit = 100_000;
let messages = build_messages(&records, "system".into(), &config);
assert_valid_tool_structure(&messages);
assert!(messages.iter().any(|message| matches!(
@@ -1162,6 +1469,145 @@ mod tests {
)));
}
#[test]
fn compacted_read_output_includes_range_and_reinspection_guidance() {
let call = call_with_lines("call_1", "src/app.rs", "1-200");
let content = numbered_read_output("src/app.rs", 1, 200);
let compacted = format_tool_output_transform(
"read",
Some(&call),
&content,
320,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("Cass compacted this tool output"));
assert!(compacted.contains("Tool: `read`"));
assert!(compacted.contains("Retained excerpt shape"));
assert!(compacted.contains("src/app.rs lines 1-200"));
assert!(compacted.contains("narrower line range"));
assert!(compacted.contains("grep"));
assert!(compacted.contains("--- retained head excerpt ---"));
assert!(compacted.contains("--- retained tail excerpt ---"));
}
#[test]
fn compacted_multi_file_read_output_summarizes_sections() {
let call = named_call(
"call_1",
"read",
json!({"files":[
{"path":"src/a.rs","lines":"1-80"},
{"path":"src/b.rs","lines":"20-90"}
]}),
);
let content = format!(
"{}{}",
numbered_read_output("src/a.rs", 1, 80),
numbered_read_output("src/b.rs", 20, 90)
);
let compacted = format_tool_output_transform(
"read",
Some(&call),
&content,
360,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("2 read sections"));
assert!(compacted.contains("src/a.rs lines 1-80"));
assert!(compacted.contains("src/b.rs lines 20-90"));
}
#[test]
fn compacted_grep_output_adds_narrowing_guidance() {
let call = named_call(
"call_1",
"grep",
json!({"query":"needle","paths":["src"],"max_matches":300}),
);
let mut content = String::new();
for line in 1..=300 {
content.push_str(&format!("src/lib.rs:{line}: needle {line}\n"));
}
content.push_str("… stopped after 300 matches. Narrow the query or raise max_matches.\n");
let compacted = format_tool_output_transform(
"grep",
Some(&call),
&content,
240,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("Tool: `grep`"));
assert!(compacted.contains("stopped after 300 matches"));
assert!(compacted.contains("narrower query/path"));
assert!(compacted.contains("read` around specific matching lines"));
}
#[test]
fn compacted_shell_output_mentions_command_and_narrowing() {
let call = named_call(
"call_1",
"shell",
json!({"command":"cargo test --locked --all-targets"}),
);
let content = format!("stdout:\n{}\nexit code: 0\n", "test output\n".repeat(200));
let compacted = format_tool_output_transform(
"shell",
Some(&call),
&content,
240,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("Tool: `shell`"));
assert!(compacted.contains("Shell command: `cargo test --locked --all-targets`"));
assert!(compacted.contains("filter output with grep/head/tail"));
assert!(compacted.contains("before making edits from omitted lines"));
}
#[test]
fn model_tool_result_limit_is_model_facing_only() {
let original = numbered_read_output("src/main.rs", 1, 120);
let records = vec![
Record::Assistant {
content: String::new(),
reasoning: String::new(),
reasoning_field: None,
tool_calls: vec![call_with_lines("call_1", "src/main.rs", "1-120")],
ts: now_ts(),
},
Record::Tool {
tool_call_id: "call_1".into(),
name: "read".into(),
ok: true,
content: original.clone(),
ts: now_ts(),
},
];
let mut config = Config::default();
config.model_tool_result_limit = 180;
let messages = build_messages(&records, "system".into(), &config);
assert_valid_tool_structure(&messages);
assert!(messages.iter().any(|message| matches!(
message,
ModelMessage::Tool { content, .. }
if content.contains("Cass truncated this tool output")
&& content.contains("src/main.rs lines 1-120")
)));
assert!(matches!(
&records[1],
Record::Tool { content, .. } if content == &original
));
}
#[test]
fn context_budget_trimming_does_not_leave_orphaned_tool_results() {
let records = vec![
+1 -1
View File
@@ -24,7 +24,7 @@ Make the smallest useful plan, then act. Prefer current project evidence over gu
"## Transcript and tools\n\
Assistant text is streamed to the user. Tool calls, tool results, edit diffs, denials, and approval prompts are visible in the transcript. Request tools directly when they are the right next step; Cassady enforces access policy and shows approval UI separately. Do not ask for chat permission before every tool call, and do not say a tool succeeded before its result arrives. If a tool fails or is denied, adapt instead of repeating the same request.\n\n\
## Tool use\n\
Use tools when the current filesystem or command result matters. Use `ls` for directory orientation, `grep` to locate definitions/usages or inspect large or unknown areas before opening files, `read` for relevant files or ranges, `edit` for focused changes to existing files, `write` for new files or intentional full rewrites, and `shell` for tests, builds, formatting, diagnostics, or project commands when allowed and useful. Prefer targeted inspection and related batched reads over broad exploration. Do not use `shell` for file inspection when `ls`, `grep`, or `read` is safer and sufficient.\n\n\
Use tools when the current filesystem or command result matters. Use `ls` for directory orientation, `grep` to locate definitions/usages or inspect large or unknown areas before opening files, `read` for relevant files or ranges, `edit` for focused changes to existing files, `write` for new files or intentional full rewrites, and `shell` for tests, builds, formatting, diagnostics, or project commands when allowed and useful. Prefer targeted inspection and related batched reads over broad exploration. Treat compacted or truncated tool output as incomplete: re-read a narrower range, search, or rerun a narrower command before editing from omitted details. Do not use `shell` for file inspection when `ls`, `grep`, or `read` is safer and sufficient.\n\n\
## Editing\n\
Inspect before editing. Prefer `edit` for small and medium modifications to existing files. For `edit`, each old text must match exactly and uniquely in the original file; keep replacements minimal, unique, and non-overlapping, and combine related replacements for the same file in one call when practical. Use `write` only for new files or full rewrites where that is safer and intentional. After meaningful code changes, run relevant tests or formatters when allowed, or tell the user what should be run.\n\n\
## Safety and final response\n\
+29 -3
View File
@@ -1,7 +1,7 @@
use super::types::{CompletionResult, ModelMessage};
use crate::agent::AgentEvent;
use crate::codex_auth::load_codex_access_token;
use crate::config::{ReasoningEffort, CHATGPT_CODEX_RESPONSES_URL};
use crate::config::{ReasoningEffort, CHATGPT_CODEX_DEFAULT_MODEL, CHATGPT_CODEX_RESPONSES_URL};
use crate::conversation::StoredToolCall;
use crate::tools::ToolSpec;
use anyhow::{bail, Result};
@@ -194,7 +194,7 @@ fn responses_body(
body["instructions"] = Value::String(instructions.join("\n\n"));
}
if fast_mode {
body["reasoning"] = json!({"effort": "minimal", "summary": "auto"});
body["reasoning"] = json!({"effort": fast_mode_reasoning_effort(model), "summary": "auto"});
} else if let Some(effort) = reasoning_effort.request_value() {
body["reasoning"] = json!({"effort": effort, "summary": "auto"});
} else if reasoning_effort == ReasoningEffort::Off {
@@ -203,6 +203,14 @@ fn responses_body(
body
}
fn fast_mode_reasoning_effort(model: &str) -> &'static str {
if model == CHATGPT_CODEX_DEFAULT_MODEL {
"low"
} else {
"minimal"
}
}
fn tools_to_responses(tools: Vec<ToolSpec>) -> Vec<Value> {
tools
.into_iter()
@@ -489,7 +497,25 @@ mod tests {
}
#[test]
fn responses_body_uses_minimal_reasoning_for_fast_mode() {
fn responses_body_uses_low_reasoning_for_gpt_5_5_fast_mode() {
let body = responses_body(
CHATGPT_CODEX_DEFAULT_MODEL,
vec![ModelMessage::User {
content: "hello".into(),
}],
Vec::new(),
ReasoningEffort::High,
true,
);
assert_eq!(
body["reasoning"],
json!({"effort": "low", "summary": "auto"})
);
}
#[test]
fn responses_body_keeps_minimal_reasoning_for_other_fast_mode_models() {
let body = responses_body(
"gpt-test",
vec![ModelMessage::User {
+1 -4
View File
@@ -342,10 +342,7 @@ mod tests {
ReasoningRequestFormat::ReasoningEffort,
true,
);
assert_eq!(
body["reasoning_effort"],
Value::String("none".to_string())
);
assert_eq!(body["reasoning_effort"], Value::String("none".to_string()));
}
#[test]
+1 -1
View File
@@ -31,7 +31,7 @@ fn default_max() -> usize {
pub fn spec() -> ToolSpec {
ToolSpec {
name: "grep".into(),
description: "Search files or directories for literal text or regex matches. Use before read for large inputs. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
description: "Search files or directories for literal text or regex matches. Use before read for large or unknown inputs; keep queries and paths focused, then read around matching lines. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
parameters: schema::object(json!({
"query": {"type":"string"},
"paths": {"type":"array", "items":{"type":"string"}, "default":["."]},
+5 -1
View File
@@ -227,7 +227,11 @@ fn truncate_model(mut s: String, limit: usize) -> String {
if s.len() <= limit {
return s;
}
s.truncate(limit);
let mut end = limit.min(s.len());
while !s.is_char_boundary(end) {
end = end.saturating_sub(1);
}
s.truncate(end);
s.push_str("\n… truncated by Cass; use grep or narrower line ranges for more.");
s
}
+1 -1
View File
@@ -18,7 +18,7 @@ struct FileArg {
pub fn spec() -> ToolSpec {
ToolSpec {
name: "read".into(),
description: "Read one or more text files, optionally with 1-indexed line ranges like 35-60, 35-, or -60. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
description: "Read one or more text files, optionally with 1-indexed line ranges like 35-60, 35-, or -60. Prefer grep first for unknown locations or large files, and re-read narrower ranges if prior output was compacted or truncated. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
parameters: schema::object(json!({
"files": {
"type":"array",
+1 -1
View File
@@ -16,7 +16,7 @@ struct Args {
pub fn spec() -> ToolSpec {
ToolSpec {
name: "shell".into(),
description: "Run a shell command in the launch cwd. Request this tool directly when shell is useful; do not ask the user for permission in chat. Cass may show a separate approval UI before execution depending on the active access mode. Streams stdout/stderr while running, then returns stdout, stderr, and exit code. Use timeout (seconds) to limit runtime."
description: "Run a shell command in the launch cwd. Request this tool directly when shell is useful; do not ask the user for permission in chat. Cass may show a separate approval UI before execution depending on the active access mode. Streams stdout/stderr while running, then returns stdout, stderr, and exit code. Use timeout (seconds) to limit runtime. Prefer narrow commands or filtering (grep/head/tail) for broad output, and rerun narrower commands if prior output was compacted or truncated."
.into(),
parameters: schema::object(
json!({
+36 -5
View File
@@ -674,11 +674,26 @@ fn collapsed_tool_summary(content: &str) -> String {
return "no output".into();
}
let lines = content.lines().count();
format!(
"{} · {} · tool output hidden",
pluralize(lines, "line"),
human_bytes(content.len())
)
let mut parts = vec![pluralize(lines, "line"), human_bytes(content.len())];
if let Some(marker) = tool_incompleteness_marker(content) {
parts.push(marker.into());
}
parts.push("tool output hidden".into());
parts.join(" · ")
}
fn tool_incompleteness_marker(content: &str) -> Option<&'static str> {
if content.contains("Cass compacted this tool output") {
Some("compacted")
} else if content.contains("truncated by Cass")
|| content.contains("Cass truncated this tool output")
{
Some("truncated")
} else if content.contains("… stopped after ") {
Some("stopped early")
} else {
None
}
}
fn pluralize(count: usize, unit: &str) -> String {
@@ -899,6 +914,22 @@ mod tests {
assert!(!text.contains("one\ntwo\nthree"));
}
#[test]
fn collapsed_tool_output_marks_incomplete_results() {
let transcript = vec![TranscriptBlock {
kind: TranscriptKind::Tool,
title: "grep ✓ (call_1)".into(),
content: "match\n… stopped after 1 matches. Narrow the query or raise max_matches."
.into(),
}];
let rendered = transcript_lines_from(&transcript, false, false);
let text = rendered_text(&rendered);
assert!(text.contains("stopped early"));
assert!(text.contains("tool output hidden"));
}
#[test]
fn successful_ls_shows_summary_when_tools_are_collapsed() {
let transcript = vec![TranscriptBlock {
+80
View File
@@ -368,6 +368,86 @@ async fn empty_final_response_is_reprompted_and_persisted() {
));
}
#[tokio::test]
async fn tool_results_are_stored_full_but_sent_to_model_with_limit_guidance() {
let server = MockServer::start().await;
Mock::given(method("POST"))
.and(path("/chat/completions"))
.and(body_string_contains("Cass truncated this tool output"))
.and(body_string_contains("large.txt lines 1-200"))
.respond_with(sse(
"data: {\"choices\":[{\"index\":0,\"delta\":{\"content\":\"Done.\"}}]}\r\n\r\ndata: [DONE]\r\n\r\n",
))
.with_priority(1)
.expect(1)
.mount(&server)
.await;
Mock::given(method("POST"))
.and(path("/chat/completions"))
.respond_with(tool_call_sse(
"call_read",
"read",
r#"{"files":[{"path":"large.txt"}]}"#,
))
.with_priority(10)
.expect(1)
.mount(&server)
.await;
let root = tempdir().unwrap();
let cwd = tempdir().unwrap();
let docs = tempdir().unwrap();
let large = (1..=200)
.map(|line| format!("line {line}"))
.collect::<Vec<_>>()
.join("\n");
std::fs::write(cwd.path().join("large.txt"), large).unwrap();
let config = Config {
root: root.path().to_path_buf(),
docs_dir: docs.path().to_path_buf(),
model: "test-model".into(),
model_tool_result_limit: 180,
active_provider: cassady::config::ResolvedProviderConfig {
base_url: server.uri(),
api_key: "test-key".into(),
..Config::default().active_provider
},
..Config::default()
};
let conversation = Conversation::create(
&config.conversations_dir(),
&config.model,
cwd.path(),
"base prompt".into(),
)
.unwrap();
let (tx, _rx) = mpsc::unbounded_channel::<AgentEvent>();
let updated = run_turn(
conversation,
"read the large file".into(),
AgentSettings {
config,
cwd: cwd.path().to_path_buf(),
mode: AccessMode::ReadOnly,
reasoning_effort: ReasoningEffort::Off,
},
tx,
)
.await
.unwrap();
assert!(updated.records.iter().any(|record| matches!(
record,
Record::Tool { name, content, .. }
if name == "read"
&& content.contains("line 200")
&& !content.contains("Cass truncated this tool output")
)));
}
#[tokio::test]
async fn workspace_edit_shell_does_not_execute_until_approved() {
let server = MockServer::start().await;
+16
View File
@@ -1,5 +1,6 @@
use cassady::access::AccessMode;
use cassady::prompt::{build_base_system_prompt, build_effective_system_prompt};
use cassady::tools;
use std::path::Path;
fn approximate_token_count(s: &str) -> usize {
@@ -44,6 +45,8 @@ fn base_prompt_has_required_sections_without_runtime_context() {
"Tool calls, tool results, edit diffs, denials, and approval prompts are visible"
));
assert!(prompt.contains("each old text must match exactly and uniquely"));
assert!(prompt.contains("compacted or truncated tool output as incomplete"));
assert!(prompt.contains("re-read a narrower range"));
assert!(prompt.contains("End every turn with a concise user-facing response"));
assert!(!prompt.contains("## Runtime context"));
assert!(!prompt.contains("Model:"));
@@ -131,6 +134,19 @@ fn runtime_constraints_stay_after_global_instructions() {
assert!(access_index < authority_index);
}
#[test]
fn tool_specs_bias_toward_narrow_reinspection() {
let specs = tools::specs(AccessMode::FullAccess);
let read = specs.iter().find(|spec| spec.name == "read").unwrap();
let grep = specs.iter().find(|spec| spec.name == "grep").unwrap();
let shell = specs.iter().find(|spec| spec.name == "shell").unwrap();
assert!(read.description.contains("Prefer grep first"));
assert!(read.description.contains("compacted or truncated"));
assert!(grep.description.contains("large or unknown inputs"));
assert!(shell.description.contains("grep/head/tail"));
}
#[test]
fn effective_prompt_size_remains_intentional() {
for mode in [