5 Commits
Author SHA1 Message Date
owen b4ff4f9f14 Prepare Cassady v0.3.5
CI / Build (push) Waiting to run
CI / Test (push) Waiting to run
2026-06-26 11:53:01 -05:00
owen ac86ee3933 Updating README 2026-06-25 22:03:46 -05:00
IrrelevantandGitHub da577aed1c Update ROADMAP with completion status for versions 2026-06-25 21:58:39 -05:00
owen ab0c45aff4 Correct roadmap: real v0.3.4 release notes, rename planned Tool Output Context Reliability to v0.3.5
The v0.3.4 tag is the released 'Reasoning Off Handling' release (commit
b35c4fb), distinct from the planned Tool Output Context Reliability work.
Add proper v0.3.4 release notes based on the shipped commit, move the Tool
Output Context Reliability entry to the top as v0.3.5, and rename its plan
file to match.
2026-06-25 21:56:44 -05:00
owen b35c4fb7e1 Send reasoning effort 'none' when reasoning is off
CI / Build (push) Waiting to run
CI / Test (push) Waiting to run
When a reasoning-capable model has reasoning effort set to off, send
'none' rather than omitting the field or sending 'off'. This applies
to both the ChatGPT Codex Responses API and OpenAI-compatible chat
completions (both reasoning_effort and reasoning object formats).
Non-reasoning models continue to send no reasoning field at all.

Bumps version to 0.3.4.
2026-06-25 20:08:15 -05:00
21 changed files with 825 additions and 80 deletions
Generated
+1 -1
View File
@@ -226,7 +226,7 @@ checksum = "8ae3f5d315924270530207e2a68396c3cc547f6dca3fbdca317cfb1a51edb593"
[[package]]
name = "cassady"
version = "0.3.3"
version = "0.3.5"
dependencies = [
"anyhow",
"async-trait",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "cassady"
version = "0.3.3"
version = "0.3.5"
edition = "2021"
description = "Cassady/Cass minimal terminal coding agent"
license = "MIT"
+7 -1
View File
@@ -8,7 +8,7 @@ The project installs two equivalent commands, `cass` and `cassady`; examples use
- Provider support includes OpenAI-compatible chat/completions APIs plus the `ChatGPT Codex` preset for users already signed in to Codex.
- The primary interface is an interactive terminal UI.
- v0.2.6 adds an experimental Rust embedding API for headless sessions; it is useful for early integrations but not yet a stable long-term library contract.
- An experimental Rust embedding API for headless sessions is available (added in v0.2.6); it is useful for early integrations but not yet a stable long-term library contract.
- Config and conversation state live under `~/.cass`.
- Windows binaries are built for releases, but deeper Windows terminal, path, shell, and filesystem polish is planned for a later release.
- `cass update` can update release-archive installs from official GitHub releases; external package managers should still be updated through their own tools.
@@ -108,6 +108,12 @@ Cassady exposes tools according to the active access mode:
Use `--readonly`, `--workspace-edit`, or `--full-access` to choose a mode at launch, or press `Shift-Tab` while idle.
## Tool output and context recovery
Cassady keeps tool calls reviewable while fitting provider context windows. Very large tool results may be sent to the model as compacted or truncated head/tail excerpts with a note that names the tool, retained excerpt shape, file ranges or command provenance when available, and suggested follow-up reads/searches. Treat those notes as incomplete evidence: ask Cass to re-read a narrower line range, run a focused `grep`, or rerun a narrower shell command before editing from omitted details.
Conversation files keep the recorded tool result content; model-facing compaction happens when preparing provider messages. The compact/full tool-output UI toggle only changes display.
## Branch and restore
Press `Esc` twice while idle, or type `/branch`, to browse the current conversation's branch family. Selecting an earlier user message, assistant message, tool call, or tool result creates a new branch conversation instead of truncating the original chat. The menu also lets you switch back to related branches later.
+50 -33
View File
@@ -1,6 +1,53 @@
# Cassady (Cass) Roadmap
## v0.3.3 — Codex Fast-Mode Compatibility
## v0.3.5 — Tool Output Context Reliability
This release focuses on making large tool outputs easier for the assistant to recover from when model-context compaction or truncation hides important details. Cassady should guide the assistant toward smaller, targeted reads and searches, preserve enough provenance for follow-up inspection, and add regression coverage for broad-output workflows that previously stalled safe edits. See `plans/V0_3_5_TOOL_OUTPUT_CONTEXT_RELIABILITY_PLAN.md`.
### Model Context Recovery
- [x] **Improve compacted tool-output guidance.** Replace generic head/tail compaction notices with actionable guidance that tells the assistant what was omitted and how to inspect it again safely.
- Include tool name, output size, retained excerpt shape, and suggested narrower follow-up reads or searches when available.
- Keep model-facing guidance concise enough that it does not worsen context pressure.
- [x] **Preserve targeted reinspection metadata.** Track enough structured context for large reads and command output so the assistant can recover omitted details without repeating broad requests.
- For file reads, preserve path and line-range coverage even after compaction.
- For shell and search output, prefer guidance toward narrower commands or `grep`/`read` follow-ups rather than blindly rerunning the same broad command.
### Tool Behavior and Prompting
- [x] **Bias tool use toward smaller inspections.** Update tool descriptions, prompt guidance, and result messages so broad reads become a fallback rather than the default.
- Encourage search-first workflows for large files and unknown locations.
- Mention result limits before or at truncation points so the assistant knows when context may be incomplete.
- [x] **Make truncation and compaction visible across layers.** Align model-facing messages, stored conversation records, and UI summaries so users and the assistant can tell when output was incomplete.
- Do not let UI-only collapsed output change what is stored or sent to the model.
- Keep existing conversation files readable and resumable.
### Validation
- [x] **Add regression coverage for broad-output recovery.** Test workflows where an early broad read or command output is compacted before the assistant needs exact context for an edit.
- Cover superseded reads, compacted non-newest tool outputs, provider-message validity, and suggested follow-up guidance.
- Verify `cargo fmt` and `cargo test --locked --all-targets` pass before handoff.
## v0.3.4 — Reasoning Off Handling ✅ Completed
This release focuses on sending the correct reasoning effort value when a reasoning-capable model has reasoning turned off. Cassady now sends `none` rather than omitting the field or sending `off` to both the ChatGPT Codex Responses API and OpenAI-compatible chat completions, while non-reasoning models continue to send no reasoning field at all.
### Reasoning Effort Off
- [x] **Send `none` when reasoning is off for supported models.** For reasoning-capable models with reasoning effort set to off, send `none` as the effort value instead of omitting the field or sending `off`.
- Apply to the ChatGPT Codex Responses API (`reasoning.effort`) and OpenAI-compatible chat completions in both `reasoning_effort` and `reasoning` object request formats.
- Keep non-reasoning models sending no reasoning field.
- [x] **Gate reasoning fields on model capability.** Track whether the active model supports reasoning from its metadata so reasoning request fields are only added for reasoning-capable models.
### Validation
- [x] **Add regression coverage for reasoning-off requests.** Test that supported models send `none` when reasoning is off across both provider kinds and both request formats, and that unsupported models send no reasoning field.
- Verify `cargo fmt` and `cargo test --locked --all-targets` pass before handoff.
## v0.3.3 — Codex Fast-Mode Compatibility ✅ Completed
This release focuses on keeping fast mode available for ChatGPT Codex users when local model metadata predates the fast-mode capability flag. Cassady should treat active `chatgpt-codex` provider models, including `gpt-5.5`, as fast-capable while leaving OpenAI-compatible and custom providers capability-gated by model metadata.
@@ -16,37 +63,7 @@ This release focuses on keeping fast mode available for ChatGPT Codex users when
- [x] **Add regression coverage for legacy metadata.** Test that a ChatGPT Codex model with older `fast_mode.supported: false` metadata still reports fast mode as supported and active when preferred.
- Verify `cargo fmt` and `cargo test --locked --all-targets` pass before handoff.
## v0.3.4 — Tool Output Context Reliability
This release focuses on making large tool outputs easier for the assistant to recover from when model-context compaction or truncation hides important details. Cassady should guide the assistant toward smaller, targeted reads and searches, preserve enough provenance for follow-up inspection, and add regression coverage for broad-output workflows that previously stalled safe edits. See `plans/V0_3_4_TOOL_OUTPUT_CONTEXT_RELIABILITY_PLAN.md`.
### Model Context Recovery
- [ ] **Improve compacted tool-output guidance.** Replace generic head/tail compaction notices with actionable guidance that tells the assistant what was omitted and how to inspect it again safely.
- Include tool name, output size, retained excerpt shape, and suggested narrower follow-up reads or searches when available.
- Keep model-facing guidance concise enough that it does not worsen context pressure.
- [ ] **Preserve targeted reinspection metadata.** Track enough structured context for large reads and command output so the assistant can recover omitted details without repeating broad requests.
- For file reads, preserve path and line-range coverage even after compaction.
- For shell and search output, prefer guidance toward narrower commands or `grep`/`read` follow-ups rather than blindly rerunning the same broad command.
### Tool Behavior and Prompting
- [ ] **Bias tool use toward smaller inspections.** Update tool descriptions, prompt guidance, and result messages so broad reads become a fallback rather than the default.
- Encourage search-first workflows for large files and unknown locations.
- Mention result limits before or at truncation points so the assistant knows when context may be incomplete.
- [ ] **Make truncation and compaction visible across layers.** Align model-facing messages, stored conversation records, and UI summaries so users and the assistant can tell when output was incomplete.
- Do not let UI-only collapsed output change what is stored or sent to the model.
- Keep existing conversation files readable and resumable.
### Validation
- [ ] **Add regression coverage for broad-output recovery.** Test workflows where an early broad read or command output is compacted before the assistant needs exact context for an edit.
- Cover superseded reads, compacted non-newest tool outputs, provider-message validity, and suggested follow-up guidance.
- Verify `cargo fmt` and `cargo test --locked --all-targets` pass before handoff.
## v0.3.2 — Provider Fast Mode
## v0.3.2 — Provider Fast Mode ✅ Completed
This release focuses on adding a `/fast` command that lets users prefer faster inference when the active provider/model supports it. The first supported provider is `ChatGPT Codex`; other providers can add their own fast-mode request behavior later without changing the user-facing command. See `plans/V0_3_2_FAST_MODE_PLAN.md`.
@@ -78,7 +95,7 @@ This release focuses on adding a `/fast` command that lets users prefer faster i
- [ ] **Test fast-mode preference, switching, and provider requests.** Cover command parsing, persistence, status rendering, model/provider switching, and Codex request body behavior.
- Verify `cargo fmt` and `cargo test --locked --all-targets` pass before handoff.
## v0.3.1 — Transcript Scroll Stability
## v0.3.1 — Transcript Scroll Stability ✅ Completed
This release focuses on keeping the live transcript anchored correctly above the input and footer during long sessions with blank reasoning or tool-output lines.
+1 -1
View File
@@ -60,7 +60,7 @@ Fields:
- `default_fast_mode`: optional boolean, defaults to `false`. When `true`, Cassady requests faster inference only for provider/model combinations that advertise fast-mode support.
- `default_access_mode`: `"read-only"`, `"workspace-edit"`, or `"full-access"`.
- `context_message_limit`: optional legacy upper bound for recent non-system messages. Cassady primarily budgets context from model metadata and trims along valid tool-call boundaries.
- `model_tool_result_limit`: optional max bytes of tool output sent back to the model.
- `model_tool_result_limit`: optional approximate max characters of each tool output sent back to the model. Larger results are model-facing head/tail excerpts with recovery guidance; conversation records keep the tool result content.
- `ui_tool_result_limit`: optional max bytes of tool output shown in the UI unless full output is toggled.
- `show_reasoning`: optional boolean, defaults to `false`. Shows provider-streamed reasoning in the transcript.
- `confirm_destructive_operations`: optional compatibility preference currently stored in config.
+4
View File
@@ -12,6 +12,8 @@
**Config root**: The `~/.cass` directory containing config, conversations, global instructions, and installed docs.
**Compacted tool output**: A model-facing replacement for a large tool result that keeps a head/tail excerpt plus provenance and recovery guidance so the assistant can re-read or re-search narrowly before relying on omitted details.
**Exact edit**: An `edit` tool replacement where each `old_text` must match exactly once in the original file before anything is written.
**Fast mode**: A saved preference enabled with `/fast`. It is active only when the current provider/model advertises fast-mode support; otherwise Cassady keeps the preference but reports it as unavailable.
@@ -28,4 +30,6 @@
**Tool call**: A model-requested operation such as `ls`, `read`, `grep`, `write`, `edit`, or `shell`.
**Truncated tool output**: A model-facing shortened tool result produced when output exceeds `model_tool_result_limit`. Cassady tells the model that output was incomplete and suggests narrower follow-up inspection.
**Workspace**: The launch cwd, either the current directory or the path passed with `--cwd`. In workspace-edit mode, writes must stay inside this root.
+8
View File
@@ -136,6 +136,14 @@ Likely cause: the command itself failed, the working directory is wrong, depende
Fix: inspect stdout/stderr, verify cwd in `/status`, and ask Cassady to rerun the smallest relevant command.
## Tool output was compacted or truncated
Symptom: a tool result note says Cassady compacted or truncated output, shows retained head/tail excerpts, or says a search stopped after a match limit.
Likely cause: the raw tool output exceeded the model-facing result limit or the active model context budget.
Fix: treat omitted output as incomplete. Ask Cassady to re-read the exact file line range named in the note, run a narrower `grep` query/path, lower `max_matches`, or rerun a shell command with a more focused flag/filter before making edits based on omitted lines.
## Exact-text edit failed
Symptom: edit reports `old_text not found`, `old_text is not unique`, or overlapping edits.
+19 -1
View File
@@ -29,7 +29,25 @@ Start in read-only mode or press `Shift-Tab` until the status shows `read-only`.
Find where configuration is loaded and summarize the precedence rules.
```
Cassady can use `ls`, `read`, and `grep` to inspect the workspace and bundled docs.
Cassady can use `ls`, `read`, and `grep` to inspect the workspace and bundled docs. For large files or unknown locations, prefer a search-first flow: `grep` for a symbol or phrase, then `read` a small line range around the relevant match.
## Recover from compacted or truncated output
When a tool result is too large for the model context, Cassady sends the model a head/tail excerpt with a recovery note. The note includes the tool name, retained excerpt shape, and file range or shell-command provenance when available.
If Cassady reports compacted or truncated output, do not rely on omitted lines for edits. Ask Cass to narrow the inspection instead:
```text
Re-read src/app.rs lines 220-280 before editing that function.
```
```text
Search only src/ for "load_config" and then read around the matching lines.
```
```text
Rerun the test command with a focused package/filter, or pipe the noisy output through grep/head/tail.
```
## Apply a focused edit
@@ -1,8 +1,8 @@
# v0.3.4 Tool Output Context Reliability Implementation Plan
# v0.3.5 Tool Output Context Reliability Implementation Plan
## Goal
v0.3.4 makes Cassady more reliable after broad tool output has been truncated, compacted, or superseded in the model context. The assistant should be able to tell when details are missing, understand which file range or command produced them, and quickly recover by using narrower reads or searches instead of stalling or making unsafe edits from incomplete context.
v0.3.5 makes Cassady more reliable after broad tool output has been truncated, compacted, or superseded in the model context. The assistant should be able to tell when details are missing, understand which file range or command produced them, and quickly recover by using narrower reads or searches instead of stalling or making unsafe edits from incomplete context.
Success statement:
+467 -21
View File
@@ -113,7 +113,9 @@ pub async fn run_turn_with_commands(
cwd: settings.cwd.clone(),
read_roots: vec![settings.cwd.clone(), docs_dir.clone()],
blocked_write_roots: vec![docs_dir.clone()],
model_result_limit: settings.config.model_tool_result_limit,
// Keep stored/UI tool results intact; build_messages applies the
// model-facing result limit when preparing provider messages.
model_result_limit: usize::MAX,
runtime_tx: None,
};
@@ -476,6 +478,7 @@ fn build_messages(records: &[Record], system: String, config: &Config) -> Vec<Mo
messages.extend(records.iter().filter_map(record_to_model_message));
messages = sanitize_tool_message_structure(messages);
supersede_old_read_outputs(&mut messages);
apply_model_tool_result_limits(&mut messages, config.model_tool_result_limit);
let budget = context_budget_tokens(config);
if estimate_messages_tokens(&messages) > budget {
@@ -518,7 +521,7 @@ fn record_to_model_message(record: &Record) -> Option<ModelMessage> {
}
fn supersede_old_read_outputs(messages: &mut [ModelMessage]) {
let read_calls = read_tool_calls_by_id(messages);
let read_calls = tool_calls_by_id(messages);
let mut read_outputs = Vec::new();
for (message_idx, message) in messages.iter().enumerate() {
let ModelMessage::Tool {
@@ -579,16 +582,14 @@ fn tool_content(messages: &[ModelMessage], idx: usize) -> &str {
}
}
fn read_tool_calls_by_id(messages: &[ModelMessage]) -> BTreeMap<String, StoredToolCall> {
fn tool_calls_by_id(messages: &[ModelMessage]) -> BTreeMap<String, StoredToolCall> {
let mut calls = BTreeMap::new();
for message in messages {
let ModelMessage::Assistant { tool_calls, .. } = message else {
continue;
};
for call in tool_calls {
if call.name == "read" {
calls.insert(call.id.clone(), call.clone());
}
calls.insert(call.id.clone(), call.clone());
}
}
calls
@@ -830,43 +831,134 @@ fn context_budget_tokens(config: &Config) -> usize {
.max(MIN_INPUT_BUDGET_TOKENS)
}
#[derive(Debug, Clone, Copy)]
enum ToolOutputTransform {
Compaction,
Truncation,
}
impl ToolOutputTransform {
fn verb(self) -> &'static str {
match self {
ToolOutputTransform::Compaction => "compacted",
ToolOutputTransform::Truncation => "truncated",
}
}
fn reason(self) -> &'static str {
match self {
ToolOutputTransform::Compaction => "to fit the model context",
ToolOutputTransform::Truncation => {
"before sending it to the model because it exceeded the configured model_tool_result_limit"
}
}
}
}
fn apply_model_tool_result_limits(messages: &mut [ModelMessage], limit_chars: usize) {
let tool_calls = tool_calls_by_id(messages);
let target_chars = limit_chars.max(1);
for idx in 1..messages.len() {
transform_tool_output_at(
messages,
idx,
target_chars,
ToolOutputTransform::Truncation,
&tool_calls,
);
}
}
fn compact_tool_outputs(messages: &mut [ModelMessage], budget: usize) {
let newest_tool_idx = messages
.iter()
.rposition(|message| matches!(message, ModelMessage::Tool { .. }));
let tool_calls = tool_calls_by_id(messages);
for idx in 1..messages.len() {
if Some(idx) == newest_tool_idx {
continue;
}
compact_tool_output_at(messages, idx, TOOL_OUTPUT_COMPACT_CHARS);
transform_tool_output_at(
messages,
idx,
TOOL_OUTPUT_COMPACT_CHARS,
ToolOutputTransform::Compaction,
&tool_calls,
);
if estimate_messages_tokens(messages) <= budget {
return;
}
}
for idx in 1..messages.len() {
compact_tool_output_at(messages, idx, TOOL_OUTPUT_TINY_CHARS);
transform_tool_output_at(
messages,
idx,
TOOL_OUTPUT_TINY_CHARS,
ToolOutputTransform::Compaction,
&tool_calls,
);
if estimate_messages_tokens(messages) <= budget {
return;
}
}
}
fn compact_tool_output_at(messages: &mut [ModelMessage], idx: usize, target_chars: usize) {
let Some(ModelMessage::Tool { content, .. }) = messages.get_mut(idx) else {
return;
fn transform_tool_output_at(
messages: &mut [ModelMessage],
idx: usize,
target_chars: usize,
transform: ToolOutputTransform,
tool_calls: &BTreeMap<String, StoredToolCall>,
) {
let replacement = {
let Some(ModelMessage::Tool {
tool_call_id,
name,
content,
}) = messages.get(idx)
else {
return;
};
if content.chars().count() <= target_chars.max(1) {
return;
}
format_tool_output_transform(
name,
tool_calls.get(tool_call_id),
content,
target_chars,
transform,
)
};
if content.chars().count() <= target_chars {
return;
if let Some(ModelMessage::Tool { content, .. }) = messages.get_mut(idx) {
*content = replacement;
}
*content = compact_text(content, target_chars);
}
fn compact_text(content: &str, target_chars: usize) -> String {
fn format_tool_output_transform(
tool_name: &str,
call: Option<&StoredToolCall>,
content: &str,
target_chars: usize,
transform: ToolOutputTransform,
) -> String {
let original_chars = content.chars().count();
let head_chars = (target_chars * 2 / 3).max(1);
let tail_chars = target_chars.saturating_sub(head_chars).max(1);
if original_chars <= target_chars.max(1) {
return content.to_string();
}
let target_chars = target_chars
.max(1)
.min(original_chars.saturating_sub(1).max(1));
let head_chars = if target_chars <= 1 {
1
} else {
(target_chars * 2 / 3).clamp(1, target_chars - 1)
};
let tail_chars = target_chars.saturating_sub(head_chars);
let head: String = content.chars().take(head_chars).collect();
let tail: String = content
.chars()
@@ -876,11 +968,208 @@ fn compact_text(content: &str, target_chars: usize) -> String {
.into_iter()
.rev()
.collect();
let notice = tool_output_transform_notice(
tool_name,
call,
content,
original_chars,
head_chars,
tail_chars,
transform,
);
let mut out = String::new();
out.push_str(&notice);
out.push_str("\n--- retained head excerpt ---\n");
out.push_str(&head);
out.push_str("\n--- omitted middle ---\n");
if tail_chars > 0 {
out.push_str("--- retained tail excerpt ---\n");
out.push_str(&tail);
}
out
}
fn tool_output_transform_notice(
tool_name: &str,
call: Option<&StoredToolCall>,
content: &str,
original_chars: usize,
head_chars: usize,
tail_chars: usize,
transform: ToolOutputTransform,
) -> String {
let retained_shape = if tail_chars > 0 {
format!("{head_chars} chars from the start and {tail_chars} chars from the end")
} else {
format!("{head_chars} chars from the start")
};
let mut sentences = vec![format!(
"Cass {} this tool output from {original_chars} chars {}. Tool: `{tool_name}`. Retained excerpt shape: {retained_shape}.",
transform.verb(),
transform.reason(),
)];
if let Some(provenance) = tool_output_provenance_sentence(tool_name, call, content) {
sentences.push(provenance);
}
sentences.push(tool_output_recovery_guidance(tool_name).to_string());
format!("[{}]", sentences.join(" "))
}
fn tool_output_provenance_sentence(
tool_name: &str,
call: Option<&StoredToolCall>,
content: &str,
) -> Option<String> {
match tool_name {
"read" => read_output_provenance_sentence(call, content),
"grep" => grep_output_provenance_sentence(content),
"shell" => shell_output_provenance_sentence(call),
_ => None,
}
}
fn read_output_provenance_sentence(call: Option<&StoredToolCall>, content: &str) -> Option<String> {
let request_specs = call
.filter(|call| call.name == "read")
.map(|call| read_request_specs(&call.arguments))
.unwrap_or_default();
let sections = parse_read_output_sections(content, &request_specs);
if !sections.is_empty() {
return Some(format!(
"Omitted content came from {}.",
read_sections_summary(&sections)
));
}
read_request_summary(call).map(|summary| format!("The read request targeted {summary}."))
}
fn read_sections_summary(sections: &[ReadOutputSection]) -> String {
let labels: Vec<String> = sections.iter().take(2).map(read_section_label).collect();
match sections.len() {
0 => "no read sections".into(),
1 => labels[0].clone(),
2 => format!("2 read sections: {} and {}", labels[0], labels[1]),
count => format!(
"{count} read sections including {} and {}",
labels[0], labels[1]
),
}
}
fn read_section_label(section: &ReadOutputSection) -> String {
format!(
"[Cass compacted this tool output from {original_chars} chars to fit the model context. Head/tail excerpt follows.]\n{head}\n… omitted …\n{tail}"
"{} lines {}-{}",
section.path, section.start_line, section.end_line
)
}
fn read_request_summary(call: Option<&StoredToolCall>) -> Option<String> {
let call = call.filter(|call| call.name == "read")?;
let mut labels = Vec::new();
if let Some(files) = call
.arguments
.get("files")
.and_then(|files| files.as_array())
{
for file in files.iter().take(2) {
let Some(path) = file.get("path").and_then(|path| path.as_str()) else {
continue;
};
labels.push(read_request_label(
path,
file.get("lines").and_then(|lines| lines.as_str()),
));
}
return match labels.len() {
0 => None,
1 => labels.first().cloned(),
2 if files.len() == 2 => Some(format!("2 files: {} and {}", labels[0], labels[1])),
_ => Some(format!(
"{} files including {} and {}",
files.len(),
labels[0],
labels[1]
)),
};
}
call.arguments
.get("path")
.and_then(|path| path.as_str())
.map(|path| {
read_request_label(
path,
call.arguments.get("lines").and_then(|lines| lines.as_str()),
)
})
}
fn read_request_label(path: &str, lines: Option<&str>) -> String {
match lines.map(str::trim).filter(|lines| !lines.is_empty()) {
Some(lines) => format!("{path} lines {lines}"),
None => path.to_string(),
}
}
fn grep_output_provenance_sentence(content: &str) -> Option<String> {
grep_stopped_after(content)
.map(|count| format!("The grep output reported it stopped after {count} matches."))
}
fn grep_stopped_after(content: &str) -> Option<usize> {
content.lines().find_map(|line| {
let rest = line.trim().strip_prefix("… stopped after ")?;
rest.split_whitespace().next()?.parse().ok()
})
}
fn shell_output_provenance_sentence(call: Option<&StoredToolCall>) -> Option<String> {
let call = call.filter(|call| call.name == "shell")?;
let command = call.arguments.get("command")?.as_str()?.trim();
if command.is_empty() {
return None;
}
let preview = command.split_whitespace().collect::<Vec<_>>().join(" ");
if command_preview_may_contain_sensitive_text(&preview) {
return Some(
"Shell command preview omitted because it may contain sensitive text; inspect the preceding tool-call arguments before rerunning."
.into(),
);
}
let chars = preview.chars().count();
if chars > 160 {
return Some(format!(
"Shell command was {chars} chars; inspect the preceding tool-call arguments before rerunning."
));
}
Some(format!("Shell command: `{preview}`."))
}
fn command_preview_may_contain_sensitive_text(command: &str) -> bool {
let lowered = command.to_ascii_lowercase();
[
"api_key",
"apikey",
"authorization",
"bearer",
"password",
"secret",
"token",
]
.iter()
.any(|needle| lowered.contains(needle))
}
fn tool_output_recovery_guidance(tool_name: &str) -> &'static str {
match tool_name {
"read" => "Recovery: use `read` with a narrower line range, or `grep` for a symbol before reading, before relying on omitted details.",
"grep" => "Recovery: rerun `grep` with a narrower query/path, lower `max_matches`, or `read` around specific matching lines before relying on omitted details.",
"shell" => "Recovery: rerun a narrower command, filter output with grep/head/tail, or inspect specific files named in the excerpt before making edits from omitted lines.",
_ => "Recovery: rerun a narrower tool request or inspect the specific file/range from the excerpt before relying on omitted details.",
}
}
fn trim_to_context_budget(mut messages: Vec<ModelMessage>, budget: usize) -> Vec<ModelMessage> {
let mut omitted = false;
while estimate_messages_tokens(&messages) > budget && messages.len() > 1 {
@@ -894,7 +1183,7 @@ fn trim_to_context_budget(mut messages: Vec<ModelMessage>, budget: usize) -> Vec
if omitted {
let note = ModelMessage::System {
content: "Cass omitted earlier conversation messages to fit the model context budget. Included tool results still follow their matching assistant tool calls; older large tool outputs may be compacted.".to_string(),
content: "Cass omitted earlier conversation messages to fit the model context budget. Included tool results still follow their matching assistant tool calls; older large tool outputs may be compacted with recovery guidance.".to_string(),
};
messages.insert(1, note);
while estimate_messages_tokens(&messages) > budget && messages.len() > 2 {
@@ -1072,7 +1361,7 @@ fn _calls(_calls: Vec<StoredToolCall>) {}
mod tests {
use super::*;
use crate::config::default_model_definition;
use serde_json::json;
use serde_json::{json, Value};
fn small_config(context_length: u64, max_output_tokens: u64) -> Config {
let mut config = Config::default();
@@ -1099,6 +1388,22 @@ mod tests {
}
}
fn named_call(id: &str, name: &str, arguments: Value) -> StoredToolCall {
StoredToolCall {
id: id.to_string(),
name: name.to_string(),
arguments,
}
}
fn numbered_read_output(path: &str, start: usize, end: usize) -> String {
let mut out = format!("--- {path} lines {start}-{end} ---\n");
for line in start..=end {
out.push_str(&format!("{line:>6} | line {line}\n"));
}
out
}
fn assert_valid_tool_structure(messages: &[ModelMessage]) {
let mut idx = 1;
while idx < messages.len() {
@@ -1153,7 +1458,9 @@ mod tests {
ts: now_ts(),
},
];
let messages = build_messages(&records, "system".into(), &small_config(8_000, 512));
let mut config = small_config(8_000, 512);
config.model_tool_result_limit = 100_000;
let messages = build_messages(&records, "system".into(), &config);
assert_valid_tool_structure(&messages);
assert!(messages.iter().any(|message| matches!(
@@ -1162,6 +1469,145 @@ mod tests {
)));
}
#[test]
fn compacted_read_output_includes_range_and_reinspection_guidance() {
let call = call_with_lines("call_1", "src/app.rs", "1-200");
let content = numbered_read_output("src/app.rs", 1, 200);
let compacted = format_tool_output_transform(
"read",
Some(&call),
&content,
320,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("Cass compacted this tool output"));
assert!(compacted.contains("Tool: `read`"));
assert!(compacted.contains("Retained excerpt shape"));
assert!(compacted.contains("src/app.rs lines 1-200"));
assert!(compacted.contains("narrower line range"));
assert!(compacted.contains("grep"));
assert!(compacted.contains("--- retained head excerpt ---"));
assert!(compacted.contains("--- retained tail excerpt ---"));
}
#[test]
fn compacted_multi_file_read_output_summarizes_sections() {
let call = named_call(
"call_1",
"read",
json!({"files":[
{"path":"src/a.rs","lines":"1-80"},
{"path":"src/b.rs","lines":"20-90"}
]}),
);
let content = format!(
"{}{}",
numbered_read_output("src/a.rs", 1, 80),
numbered_read_output("src/b.rs", 20, 90)
);
let compacted = format_tool_output_transform(
"read",
Some(&call),
&content,
360,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("2 read sections"));
assert!(compacted.contains("src/a.rs lines 1-80"));
assert!(compacted.contains("src/b.rs lines 20-90"));
}
#[test]
fn compacted_grep_output_adds_narrowing_guidance() {
let call = named_call(
"call_1",
"grep",
json!({"query":"needle","paths":["src"],"max_matches":300}),
);
let mut content = String::new();
for line in 1..=300 {
content.push_str(&format!("src/lib.rs:{line}: needle {line}\n"));
}
content.push_str("… stopped after 300 matches. Narrow the query or raise max_matches.\n");
let compacted = format_tool_output_transform(
"grep",
Some(&call),
&content,
240,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("Tool: `grep`"));
assert!(compacted.contains("stopped after 300 matches"));
assert!(compacted.contains("narrower query/path"));
assert!(compacted.contains("read` around specific matching lines"));
}
#[test]
fn compacted_shell_output_mentions_command_and_narrowing() {
let call = named_call(
"call_1",
"shell",
json!({"command":"cargo test --locked --all-targets"}),
);
let content = format!("stdout:\n{}\nexit code: 0\n", "test output\n".repeat(200));
let compacted = format_tool_output_transform(
"shell",
Some(&call),
&content,
240,
ToolOutputTransform::Compaction,
);
assert!(compacted.contains("Tool: `shell`"));
assert!(compacted.contains("Shell command: `cargo test --locked --all-targets`"));
assert!(compacted.contains("filter output with grep/head/tail"));
assert!(compacted.contains("before making edits from omitted lines"));
}
#[test]
fn model_tool_result_limit_is_model_facing_only() {
let original = numbered_read_output("src/main.rs", 1, 120);
let records = vec![
Record::Assistant {
content: String::new(),
reasoning: String::new(),
reasoning_field: None,
tool_calls: vec![call_with_lines("call_1", "src/main.rs", "1-120")],
ts: now_ts(),
},
Record::Tool {
tool_call_id: "call_1".into(),
name: "read".into(),
ok: true,
content: original.clone(),
ts: now_ts(),
},
];
let mut config = Config::default();
config.model_tool_result_limit = 180;
let messages = build_messages(&records, "system".into(), &config);
assert_valid_tool_structure(&messages);
assert!(messages.iter().any(|message| matches!(
message,
ModelMessage::Tool { content, .. }
if content.contains("Cass truncated this tool output")
&& content.contains("src/main.rs lines 1-120")
)));
assert!(matches!(
&records[1],
Record::Tool { content, .. } if content == &original
));
}
#[test]
fn context_budget_trimming_does_not_leave_orphaned_tool_results() {
let records = vec![
+1 -1
View File
@@ -24,7 +24,7 @@ Make the smallest useful plan, then act. Prefer current project evidence over gu
"## Transcript and tools\n\
Assistant text is streamed to the user. Tool calls, tool results, edit diffs, denials, and approval prompts are visible in the transcript. Request tools directly when they are the right next step; Cassady enforces access policy and shows approval UI separately. Do not ask for chat permission before every tool call, and do not say a tool succeeded before its result arrives. If a tool fails or is denied, adapt instead of repeating the same request.\n\n\
## Tool use\n\
Use tools when the current filesystem or command result matters. Use `ls` for directory orientation, `grep` to locate definitions/usages or inspect large or unknown areas before opening files, `read` for relevant files or ranges, `edit` for focused changes to existing files, `write` for new files or intentional full rewrites, and `shell` for tests, builds, formatting, diagnostics, or project commands when allowed and useful. Prefer targeted inspection and related batched reads over broad exploration. Do not use `shell` for file inspection when `ls`, `grep`, or `read` is safer and sufficient.\n\n\
Use tools when the current filesystem or command result matters. Use `ls` for directory orientation, `grep` to locate definitions/usages or inspect large or unknown areas before opening files, `read` for relevant files or ranges, `edit` for focused changes to existing files, `write` for new files or intentional full rewrites, and `shell` for tests, builds, formatting, diagnostics, or project commands when allowed and useful. Prefer targeted inspection and related batched reads over broad exploration. Treat compacted or truncated tool output as incomplete: re-read a narrower range, search, or rerun a narrower command before editing from omitted details. Do not use `shell` for file inspection when `ls`, `grep`, or `read` is safer and sufficient.\n\n\
## Editing\n\
Inspect before editing. Prefer `edit` for small and medium modifications to existing files. For `edit`, each old text must match exactly and uniquely in the original file; keep replacements minimal, unique, and non-overlapping, and combine related replacements for the same file in one call when practical. Use `write` only for new files or full rewrites where that is safer and intentional. After meaningful code changes, run relevant tests or formatters when allowed, or tell the user what should be run.\n\n\
## Safety and final response\n\
+49 -3
View File
@@ -1,7 +1,7 @@
use super::types::{CompletionResult, ModelMessage};
use crate::agent::AgentEvent;
use crate::codex_auth::load_codex_access_token;
use crate::config::{ReasoningEffort, CHATGPT_CODEX_RESPONSES_URL};
use crate::config::{ReasoningEffort, CHATGPT_CODEX_DEFAULT_MODEL, CHATGPT_CODEX_RESPONSES_URL};
use crate::conversation::StoredToolCall;
use crate::tools::ToolSpec;
use anyhow::{bail, Result};
@@ -194,13 +194,23 @@ fn responses_body(
body["instructions"] = Value::String(instructions.join("\n\n"));
}
if fast_mode {
body["reasoning"] = json!({"effort": "minimal", "summary": "auto"});
body["reasoning"] = json!({"effort": fast_mode_reasoning_effort(model), "summary": "auto"});
} else if let Some(effort) = reasoning_effort.request_value() {
body["reasoning"] = json!({"effort": effort, "summary": "auto"});
} else if reasoning_effort == ReasoningEffort::Off {
body["reasoning"] = json!({"effort": "none", "summary": "auto"});
}
body
}
fn fast_mode_reasoning_effort(model: &str) -> &'static str {
if model == CHATGPT_CODEX_DEFAULT_MODEL {
"low"
} else {
"minimal"
}
}
fn tools_to_responses(tools: Vec<ToolSpec>) -> Vec<Value> {
tools
.into_iter()
@@ -487,7 +497,25 @@ mod tests {
}
#[test]
fn responses_body_uses_minimal_reasoning_for_fast_mode() {
fn responses_body_uses_low_reasoning_for_gpt_5_5_fast_mode() {
let body = responses_body(
CHATGPT_CODEX_DEFAULT_MODEL,
vec![ModelMessage::User {
content: "hello".into(),
}],
Vec::new(),
ReasoningEffort::High,
true,
);
assert_eq!(
body["reasoning"],
json!({"effort": "low", "summary": "auto"})
);
}
#[test]
fn responses_body_keeps_minimal_reasoning_for_other_fast_mode_models() {
let body = responses_body(
"gpt-test",
vec![ModelMessage::User {
@@ -504,6 +532,24 @@ mod tests {
);
}
#[test]
fn responses_body_sends_none_effort_when_reasoning_is_off() {
let body = responses_body(
"gpt-test",
vec![ModelMessage::User {
content: "hello".into(),
}],
Vec::new(),
ReasoningEffort::Off,
false,
);
assert_eq!(
body["reasoning"],
json!({"effort": "none", "summary": "auto"})
);
}
#[test]
fn stream_parser_collects_text_and_function_call() {
let (tx, _rx) = mpsc::unbounded_channel();
+6 -3
View File
@@ -28,11 +28,13 @@ impl ProviderClient {
match config.active_provider.kind.as_str() {
DEFAULT_PROVIDER_KIND => {
let api_key = config.resolved_api_key()?;
let reasoning_request_format = config
.model_metadata
.as_ref()
let model_metadata = config.model_metadata.as_ref();
let reasoning_request_format = model_metadata
.map(|model| model.reasoning.request_format)
.unwrap_or_default();
let reasoning_supported = model_metadata
.map(|model| model.reasoning.supported)
.unwrap_or(false);
Ok(Self::OpenAiCompatible(OpenAiCompatibleProvider::new(
OpenAiCompatibleSettings {
model: config.model.clone(),
@@ -40,6 +42,7 @@ impl ProviderClient {
api_key,
reasoning_effort: options.reasoning_effort,
reasoning_request_format,
reasoning_supported,
},
)))
}
+69 -3
View File
@@ -18,6 +18,7 @@ pub struct OpenAiCompatibleProvider {
api_key: String,
reasoning_effort: ReasoningEffort,
reasoning_request_format: ReasoningRequestFormat,
reasoning_supported: bool,
}
#[derive(Debug, Clone)]
@@ -27,6 +28,7 @@ pub struct OpenAiCompatibleSettings {
pub api_key: String,
pub reasoning_effort: ReasoningEffort,
pub reasoning_request_format: ReasoningRequestFormat,
pub reasoning_supported: bool,
}
#[derive(Debug, Default)]
@@ -45,6 +47,7 @@ impl OpenAiCompatibleProvider {
api_key: settings.api_key,
reasoning_effort: settings.reasoning_effort,
reasoning_request_format: settings.reasoning_request_format,
reasoning_supported: settings.reasoning_supported,
}
}
@@ -65,6 +68,7 @@ impl OpenAiCompatibleProvider {
&mut body,
self.reasoning_effort,
self.reasoning_request_format,
self.reasoning_supported,
);
let resp = self
.client
@@ -219,9 +223,17 @@ fn apply_reasoning_request(
body: &mut Value,
effort: ReasoningEffort,
format: ReasoningRequestFormat,
supported: bool,
) {
let Some(effort) = effort.request_value() else {
if !supported {
return;
}
let effort_str = match effort {
ReasoningEffort::Off => "none",
_ => match effort.request_value() {
Some(value) => value,
None => return,
},
};
let Value::Object(obj) = body else {
return;
@@ -230,11 +242,11 @@ fn apply_reasoning_request(
ReasoningRequestFormat::ReasoningEffort => {
obj.insert(
"reasoning_effort".to_string(),
Value::String(effort.to_string()),
Value::String(effort_str.to_string()),
);
}
ReasoningRequestFormat::ReasoningObject => {
obj.insert("reasoning".to_string(), json!({ "effort": effort }));
obj.insert("reasoning".to_string(), json!({ "effort": effort_str }));
}
}
}
@@ -316,3 +328,57 @@ fn chat_url(base: &str) -> String {
format!("{}/chat/completions", base.trim_end_matches('/'))
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn reasoning_effort_format_sends_none_when_off_and_supported() {
let mut body = json!({"model": "test"});
apply_reasoning_request(
&mut body,
ReasoningEffort::Off,
ReasoningRequestFormat::ReasoningEffort,
true,
);
assert_eq!(body["reasoning_effort"], Value::String("none".to_string()));
}
#[test]
fn reasoning_object_format_sends_none_when_off_and_supported() {
let mut body = json!({"model": "test"});
apply_reasoning_request(
&mut body,
ReasoningEffort::Off,
ReasoningRequestFormat::ReasoningObject,
true,
);
assert_eq!(body["reasoning"], json!({ "effort": "none" }));
}
#[test]
fn reasoning_sends_nothing_when_unsupported_even_if_off() {
let mut body = json!({"model": "test"});
apply_reasoning_request(
&mut body,
ReasoningEffort::Off,
ReasoningRequestFormat::ReasoningEffort,
false,
);
assert!(body.get("reasoning_effort").is_none());
assert!(body.get("reasoning").is_none());
}
#[test]
fn reasoning_sends_nothing_when_unsupported_even_if_high() {
let mut body = json!({"model": "test"});
apply_reasoning_request(
&mut body,
ReasoningEffort::High,
ReasoningRequestFormat::ReasoningObject,
false,
);
assert!(body.get("reasoning").is_none());
}
}
+1 -1
View File
@@ -31,7 +31,7 @@ fn default_max() -> usize {
pub fn spec() -> ToolSpec {
ToolSpec {
name: "grep".into(),
description: "Search files or directories for literal text or regex matches. Use before read for large inputs. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
description: "Search files or directories for literal text or regex matches. Use before read for large or unknown inputs; keep queries and paths focused, then read around matching lines. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
parameters: schema::object(json!({
"query": {"type":"string"},
"paths": {"type":"array", "items":{"type":"string"}, "default":["."]},
+5 -1
View File
@@ -227,7 +227,11 @@ fn truncate_model(mut s: String, limit: usize) -> String {
if s.len() <= limit {
return s;
}
s.truncate(limit);
let mut end = limit.min(s.len());
while !s.is_char_boundary(end) {
end = end.saturating_sub(1);
}
s.truncate(end);
s.push_str("\n… truncated by Cass; use grep or narrower line ranges for more.");
s
}
+1 -1
View File
@@ -18,7 +18,7 @@ struct FileArg {
pub fn spec() -> ToolSpec {
ToolSpec {
name: "read".into(),
description: "Read one or more text files, optionally with 1-indexed line ranges like 35-60, 35-, or -60. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
description: "Read one or more text files, optionally with 1-indexed line ranges like 35-60, 35-, or -60. Prefer grep first for unknown locations or large files, and re-read narrower ranges if prior output was compacted or truncated. In read-only and workspace-edit modes, paths must stay inside the launch cwd or bundled docs directory.".into(),
parameters: schema::object(json!({
"files": {
"type":"array",
+1 -1
View File
@@ -16,7 +16,7 @@ struct Args {
pub fn spec() -> ToolSpec {
ToolSpec {
name: "shell".into(),
description: "Run a shell command in the launch cwd. Request this tool directly when shell is useful; do not ask the user for permission in chat. Cass may show a separate approval UI before execution depending on the active access mode. Streams stdout/stderr while running, then returns stdout, stderr, and exit code. Use timeout (seconds) to limit runtime."
description: "Run a shell command in the launch cwd. Request this tool directly when shell is useful; do not ask the user for permission in chat. Cass may show a separate approval UI before execution depending on the active access mode. Streams stdout/stderr while running, then returns stdout, stderr, and exit code. Use timeout (seconds) to limit runtime. Prefer narrow commands or filtering (grep/head/tail) for broad output, and rerun narrower commands if prior output was compacted or truncated."
.into(),
parameters: schema::object(
json!({
+36 -5
View File
@@ -674,11 +674,26 @@ fn collapsed_tool_summary(content: &str) -> String {
return "no output".into();
}
let lines = content.lines().count();
format!(
"{} · {} · tool output hidden",
pluralize(lines, "line"),
human_bytes(content.len())
)
let mut parts = vec![pluralize(lines, "line"), human_bytes(content.len())];
if let Some(marker) = tool_incompleteness_marker(content) {
parts.push(marker.into());
}
parts.push("tool output hidden".into());
parts.join(" · ")
}
fn tool_incompleteness_marker(content: &str) -> Option<&'static str> {
if content.contains("Cass compacted this tool output") {
Some("compacted")
} else if content.contains("truncated by Cass")
|| content.contains("Cass truncated this tool output")
{
Some("truncated")
} else if content.contains("… stopped after ") {
Some("stopped early")
} else {
None
}
}
fn pluralize(count: usize, unit: &str) -> String {
@@ -899,6 +914,22 @@ mod tests {
assert!(!text.contains("one\ntwo\nthree"));
}
#[test]
fn collapsed_tool_output_marks_incomplete_results() {
let transcript = vec![TranscriptBlock {
kind: TranscriptKind::Tool,
title: "grep ✓ (call_1)".into(),
content: "match\n… stopped after 1 matches. Narrow the query or raise max_matches."
.into(),
}];
let rendered = transcript_lines_from(&transcript, false, false);
let text = rendered_text(&rendered);
assert!(text.contains("stopped early"));
assert!(text.contains("tool output hidden"));
}
#[test]
fn successful_ls_shows_summary_when_tools_are_collapsed() {
let transcript = vec![TranscriptBlock {
+80
View File
@@ -368,6 +368,86 @@ async fn empty_final_response_is_reprompted_and_persisted() {
));
}
#[tokio::test]
async fn tool_results_are_stored_full_but_sent_to_model_with_limit_guidance() {
let server = MockServer::start().await;
Mock::given(method("POST"))
.and(path("/chat/completions"))
.and(body_string_contains("Cass truncated this tool output"))
.and(body_string_contains("large.txt lines 1-200"))
.respond_with(sse(
"data: {\"choices\":[{\"index\":0,\"delta\":{\"content\":\"Done.\"}}]}\r\n\r\ndata: [DONE]\r\n\r\n",
))
.with_priority(1)
.expect(1)
.mount(&server)
.await;
Mock::given(method("POST"))
.and(path("/chat/completions"))
.respond_with(tool_call_sse(
"call_read",
"read",
r#"{"files":[{"path":"large.txt"}]}"#,
))
.with_priority(10)
.expect(1)
.mount(&server)
.await;
let root = tempdir().unwrap();
let cwd = tempdir().unwrap();
let docs = tempdir().unwrap();
let large = (1..=200)
.map(|line| format!("line {line}"))
.collect::<Vec<_>>()
.join("\n");
std::fs::write(cwd.path().join("large.txt"), large).unwrap();
let config = Config {
root: root.path().to_path_buf(),
docs_dir: docs.path().to_path_buf(),
model: "test-model".into(),
model_tool_result_limit: 180,
active_provider: cassady::config::ResolvedProviderConfig {
base_url: server.uri(),
api_key: "test-key".into(),
..Config::default().active_provider
},
..Config::default()
};
let conversation = Conversation::create(
&config.conversations_dir(),
&config.model,
cwd.path(),
"base prompt".into(),
)
.unwrap();
let (tx, _rx) = mpsc::unbounded_channel::<AgentEvent>();
let updated = run_turn(
conversation,
"read the large file".into(),
AgentSettings {
config,
cwd: cwd.path().to_path_buf(),
mode: AccessMode::ReadOnly,
reasoning_effort: ReasoningEffort::Off,
},
tx,
)
.await
.unwrap();
assert!(updated.records.iter().any(|record| matches!(
record,
Record::Tool { name, content, .. }
if name == "read"
&& content.contains("line 200")
&& !content.contains("Cass truncated this tool output")
)));
}
#[tokio::test]
async fn workspace_edit_shell_does_not_execute_until_approved() {
let server = MockServer::start().await;
+16
View File
@@ -1,5 +1,6 @@
use cassady::access::AccessMode;
use cassady::prompt::{build_base_system_prompt, build_effective_system_prompt};
use cassady::tools;
use std::path::Path;
fn approximate_token_count(s: &str) -> usize {
@@ -44,6 +45,8 @@ fn base_prompt_has_required_sections_without_runtime_context() {
"Tool calls, tool results, edit diffs, denials, and approval prompts are visible"
));
assert!(prompt.contains("each old text must match exactly and uniquely"));
assert!(prompt.contains("compacted or truncated tool output as incomplete"));
assert!(prompt.contains("re-read a narrower range"));
assert!(prompt.contains("End every turn with a concise user-facing response"));
assert!(!prompt.contains("## Runtime context"));
assert!(!prompt.contains("Model:"));
@@ -131,6 +134,19 @@ fn runtime_constraints_stay_after_global_instructions() {
assert!(access_index < authority_index);
}
#[test]
fn tool_specs_bias_toward_narrow_reinspection() {
let specs = tools::specs(AccessMode::FullAccess);
let read = specs.iter().find(|spec| spec.name == "read").unwrap();
let grep = specs.iter().find(|spec| spec.name == "grep").unwrap();
let shell = specs.iter().find(|spec| spec.name == "shell").unwrap();
assert!(read.description.contains("Prefer grep first"));
assert!(read.description.contains("compacted or truncated"));
assert!(grep.description.contains("large or unknown inputs"));
assert!(shell.description.contains("grep/head/tail"));
}
#[test]
fn effective_prompt_size_remains_intentional() {
for mode in [