Pick 3a of the Discord file-passing plan: extend the URL-flavored File
variant with optional mime and size metadata so adapters can pass
attachment context through to bridges. FileData (bytes-flavored) is
unchanged; size is implicit in data.len() and mime_type already exists.
Match-arm sites in bridge.rs, telegram.rs, whatsapp.rs use `..` to stay
forward-compatible. Construction sites in telegram.rs and kernel.rs
pass `mime: None, size: None` for now; Discord inbound (PR-A) will
populate them.
Refs: projects/openfang-fork/discord-file-passing-plan.md
A typo in any binding's match_rule no longer drops the entire bindings
table. Each entry is parsed independently; malformed entries log an
ERROR with index, agent name, and the underlying serde error, then are
skipped. A single WARN summarizes total dropped vs. surviving bindings.
Per-entry deny_unknown_fields is preserved so silent typos still fail
loudly — just no longer catastrophically.
Before this change, a single misspelled field anywhere in [[bindings]]
caused the whole table to fail parsing, silently unbinding every
agent — the worst possible failure mode for a routing config.
- New `lenient_extract_bindings` runs after include-merge / [api]
migration, before `try_into::<KernelConfig>()`.
- 7 new config tests cover the reproducer, happy path, all-malformed,
no-bindings, missing-agent, survivor-order preservation, and
top-level field typos:
* test_lenient_bindings_drops_typo_keeps_rest
* test_lenient_bindings_all_valid_unchanged
* test_lenient_bindings_all_malformed_yields_empty_but_keeps_rest_of_config
* test_lenient_bindings_no_bindings_section_is_noop
* test_lenient_bindings_missing_agent_field_dropped
* test_lenient_bindings_preserves_survivor_order — locks in that
first-match-wins routing semantics cannot silently regress when
a middle entry is dropped
* test_lenient_bindings_top_level_field_typo_dropped — locks in
that deny_unknown_fields catches operator typos at the binding
top level (e.g. \`agnt = ...\`), not just inside match_rule
- TODO marker added on the remaining \`warn!\` fallback in \`load_config\`
for the non-binding silent-default path (follow-up work).
Tested live: typo'd \`hannel\` field on binding #2 logged as expected;
remaining 5 bindings loaded and routed correctly.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The TOML-vs-DB change detection at boot only checked a subset of fields,
causing edits to workspace, schedule, resources, autonomous, and exec_policy
to be silently ignored after a restart.
Add the missing fields to the changed-detection predicate so the kernel
properly reflects all agent.toml edits in the SQL database. The workspace
comparison is intentionally guarded — if the TOML omits workspace (None),
the kernel-assigned default path already stored in the DB is kept rather
than being overwritten with None.
Derive PartialEq on ScheduleMode, AutonomousConfig, ResourceQuota, and
ExecPolicy to enable the comparisons without manual field-by-field expansion.
Co-authored-by: octo-patch <octo-patch@github.com>
Follow-up to 79aa34c. The previous commit added the public surface
(DriverConfig field + OPENFANG_SUBPROCESS_TIMEOUT_SECS env var) but
left every DriverConfig construction site hardcoded to None — so the
struct field was wired but had no on-disk source feeding it. The env
var was the only operator-facing knob.
This commit plumbs the missing layer: the timeout is now deserializable
from config.toml on both the primary and global-fallback providers.
Public surface
- DefaultModelConfig.subprocess_timeout_secs: Option<u64>
- FallbackProviderConfig.subprocess_timeout_secs: Option<u64>
- Both fields are #[serde(default)] — existing config.toml files
without the field deserialize cleanly to None (no breaking change).
Placement rationale
- Per-provider on each config struct, not a top-level field or a new
[driver] section. This matches the existing per-provider config shape
and lets operators set different timeouts for primary vs. fallback
(e.g. tighter timeout on a fast fallback to fail over sooner). If a
second driver-level setting ever lands, refactoring two struct fields
into a [driver] section is cheap; we don't pre-pay for it now.
Wiring (kernel.rs)
- L663 primary driver ........ pulls config.default_model.subprocess_timeout_secs
- L687 auto-detect path ...... inherits default_model intent (the swap
is replacing the *provider*, not the
timeout policy)
- L736 global fallback loop .. pulls fb.subprocess_timeout_secs
- L5031 agent primary ......... inherits effective_default's value when
agent_provider == default_provider;
None for cross-provider overrides
- L5108 agent manifest fallback inherits dm's value when the manifest
fallback resolves to "default" (matching
the existing fb.provider sentinel logic);
None for explicit cross-provider entries
- L5139 global fallback (per-agent loop) — pulls fb.subprocess_timeout_secs
Sites kept as None (intentional)
- agent_loop.rs:1146, 1330: ModelNotFound recovery iterates over the
agent manifest's fallback_models (FallbackModel, not the config-toml
type) — no per-provider config in scope.
- routes.rs:7701: provider connectivity test endpoint; no config source.
- routes.rs:7529: dashboard hot-update path constructs a fresh DM with
defaults (None) — operator sets timeout via config.toml, not via the
set-key flow.
Tests
- test_subprocess_timeout_secs_in_toml: round-trips a TOML doc with
default_model.subprocess_timeout_secs = 600 and one fallback at 180,
one fallback omitted; asserts each value (or None) reaches the parsed
config struct.
- test_subprocess_timeout_secs_omitted_defaults_to_none: asserts a
legacy-shaped config.toml (no timeout fields) parses cleanly with
both fields = None — backward-compat guard.
- 4 existing claude_code driver timeout tests still pass.
Mechanical pass-throughs
- 8 test fixtures across openfang-kernel/tests and openfang-api/tests
gain subprocess_timeout_secs: None on their DefaultModelConfig
literals.
- 1 production literal in routes.rs gains the same field.
- The existing FallbackProviderConfig serde-roundtrip test gains
subprocess_timeout_secs: None plus an assertion.
Precedence comment in drivers/mod.rs::create_driver updated to reflect
that the config-field path is now real, with explicit pointers to the
kernel.rs wiring sites for future contributors.
Validated: cargo check --workspace --tests is clean; openfang-types
(362), openfang-runtime (933), and openfang-kernel (260) lib tests
all pass.
The claude-code driver hardcodes its per-message turn timeout inside
ClaudeCodeDriver and exposed no operator-facing knob, so long-running
CC subprocess turns (large prompt-caches, deep tool chains) hit the
internal default with no escape hatch. Adds a public config surface,
honored today only by the claude-code driver, designed so future
subprocess drivers can opt in without re-shaping the API.
Public surface
- DriverConfig.subprocess_timeout_secs: Option<u64> (llm_driver.rs)
- OPENFANG_SUBPROCESS_TIMEOUT_SECS env var (drivers/mod.rs)
- Precedence in create_driver(): env var > config field > driver default
Naming rationale
- Field/env are scope-flavored, not semantic, on purpose: the name
telegraphs that HTTP providers (default/Anthropic, openai, bedrock,
qwen-code) accept-but-silently-ignore the field today. A semantic
name (message_timeout_secs) would have invited the same silent-no-op
footgun on those providers.
- Driver-internal field in claude_code.rs intentionally kept as
message_timeout_secs — it's not on the public boundary and the
semantic name accurately describes what it stores.
Tests (drivers/mod.rs)
- default_when_unset: no env, no config -> driver default
- config_set: config field flows through
- env_overrides_config: env var wins over config (construction-only
assertion; trait-object opacity prevents reading the value back)
- malformed_env_falls_through: unparseable env silently falls through
to config, matching the .parse::<u64>().ok() chain in production
- All four tests scrub OPENFANG_SUBPROCESS_TIMEOUT_SECS pre/post to
avoid cross-test pollution
Mechanical pass-throughs
- 12 x DriverConfig { .. } test fixtures in drivers/mod.rs gain
subprocess_timeout_secs: None
- routes.rs (1), kernel.rs (6), agent_loop.rs (2): same pass-through
fills in DriverConfig literals; no logic touched
Forward-compat note
- A NOTE block in drivers/mod.rs flags the scope-vs-implementation
gap so the next contributor adding a subprocess driver knows
exactly where to wire the config in.
Validated end-to-end against a live daemon: dry-run + full deploy
(deploy-local.sh, all 7 phases) + post-swap agent_send round-trip
through the claude-code dispatch path.
The WASM sandbox host_net_fetch() had its own SSRF implementation
(is_ssrf_target) that was incomplete compared to the canonical
check_ssrf() in web_fetch.rs:
- Missing 6 blocked hostnames (ip6-localhost, Alibaba/Azure IMDS,
0.0.0.0, ::1, [::1])
- Missing cloud metadata IP detection (is_metadata_ip)
- Missing IPv6 bracket notation support
- Ignoring ssrf_allowed_hosts from config.toml entirely
- Duplicate is_private_ip() and extract_host_from_url() functions
This meant a WASM agent could bypass SSRF protections that the
builtin web_fetch tool correctly enforced.
Changes:
- Remove duplicated is_ssrf_target(), is_private_ip(), and
extract_host_from_url() from host_functions.rs
- Delegate to web_fetch::check_ssrf() which has the complete
implementation with allowlist, CIDR matching, and metadata
IP detection
- Add ssrf_allowed_hosts to SandboxConfig and GuestState so the
config propagates to WASM host calls
- Make extract_host() pub(crate) for reuse
- Update tests to exercise the unified code path, including new
coverage for IPv6 and cloud metadata endpoints
All 908 runtime tests pass. Zero clippy warnings.
The schedule_create tool, its sibling schedule_list and schedule_delete,
and the matching /api/schedules HTTP routes were all writing to a
shared-memory key that no executor ever read. Jobs registered that way
silently never fired.
Route all three tools and all /api/schedules endpoints through the real
cron scheduler in openfang-kernel. Add a one-shot idempotent migration
at kernel startup that imports legacy __openfang_schedules entries into
the cron scheduler and clears the old key.
Tests:
- Unit tests for sanitize_schedule_name and sanitize_cron_job_name
- Tool wrapper tests using a fake KernelHandle that verify
schedule_create/list/delete route into cron_create/list/cancel
- Migration tests cover the happy path, idempotency via the marker key,
and skipping entries whose target agent is not in the registry
Quality gates: cargo check + test + clippy -D warnings + fmt clean on
openfang-kernel, openfang-runtime, openfang-api.
Made-with: Cursor
When an agent had a context.md file updated externally (e.g. a cron job
refreshing live market data), the updated content never reached the LLM
during an active session. The file was effectively cached for the
lifetime of the conversation.
The runtime now reads context.md from the agent workspace once per turn,
right before the system prompt is built, and injects it as a dedicated
'Live Context' section. Agents that still want the old cache-at-start
behaviour can opt back in with 'cache_context = true' on the manifest.
- new openfang-runtime::agent_context module with a small per-path cache
- if a re-read fails after a previous success, fall back to the cached
content with a warning instead of dropping context mid-conversation
- new PromptContext.context_md field wired up in both kernel streaming
and non-streaming paths
- one small disk read per agent turn (not per streaming token); file
size capped at 32 KB like the other identity files
Made-with: Cursor
Bug: in activate_hand(), kill_agent() is called on the existing agent
BEFORE the new agent is spawned. kill_agent() invokes
cron_scheduler.remove_agent_jobs() which deletes all cron jobs from memory
AND persists [] to cron_jobs.json. The reassign_agent_jobs() call further
down was meant to migrate jobs from old to new (per #461), but it always
runs as a no-op because the jobs are already gone — the order of
operations defeats the fix.
Symptom: every daemon restart silently destroys cron jobs for hand-style
agents. cron_jobs.json is rewritten as []. /api/cron/jobs returns empty.
No error message.
Fix: snapshot the cron jobs into a local Vec BEFORE kill_agent (same
pattern as saved_triggers above), then re-add them under the new agent_id
AFTER spawn_agent_with_parent. Runtime state (next_run, last_run) is
reset so jobs get a fresh start. The existing reassign_agent_jobs()
block is kept as a defensive safety net but is now redundant in the
common path.
Verified with cargo check -p openfang-kernel --lib (clean compile, no
warnings).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The Copilot LLM driver was broken - it expected users to provide a
GITHUB_TOKEN env var, but no standard token type (PAT, gh CLI token)
works with the Copilot token exchange endpoint.
Changes:
- Full rewrite of copilot.rs with OAuth device flow using Copilot's
client ID (Iv1.b507a08c87ecfe98)
- Three-layer token chain: ghu_ (8h) -> Copilot API token (30min),
with automatic caching and refresh
- Dynamic model fetching from Copilot API on daemon startup and on
model_not_supported error
- Init wizard: TUI auth screen with device code display, live model
picker after authentication
- set-key command: interactive device flow for github-copilot provider
- Doctor: detects Copilot auth via persisted token file
- Removed static Copilot model entries (now fetched dynamically)
- Simplified driver instantiation (no env vars needed)
Tested end-to-end with Copilot Enterprise: auth, token exchange,
43 models fetched, completions working with claude-opus-4.6-1m.
Closes#1014
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The agent config reload logic was missing skills and mcp_servers from
the change detection, so edits to these fields in agent.toml weren't
being picked up when loading agents from SQLite.
Added both fields to the comparison to ensure proper hot-reload.
- #875: Install script uses robust sed parsing instead of fragile cut for version detection
- #872: Session endpoint returns full tool results (removed 2000-char truncation)
- #867: agent_send/agent_spawn get 600s timeout (was 120s), regular tools keep 120s
- #824: Doctor workspace skills count uses direct return value from load_workspace_skills
- #833: Model switching respects provider via new find_model_for_provider() lookup
- #766: Closed as resolved by combined heartbeat fixes (v0.5.3 + merged PRs)
All tests passing. Live tested with daemon.
- Add missing budget_config field to AppState in all 3 test files
- Fix redundant closures and unwrap_or_else in openfang-memory semantic.rs
- Fix needless_borrow in openfang-api routes.rs (toml::from_str)
- Update parse_researcher_hand test to match new max_iterations = 25
- Update tar to 0.4.45 to fix RUSTSEC-2026-0067 and RUSTSEC-2026-0068
- Apply cargo fmt fixes in ws.rs and feishu.rs