mirror of
https://github.com/tinyhumansai/openhuman.git
synced 2026-07-28 13:32:23 +00:00
403f239ca58ee0f9347ca096deeacbe03b20cdc9
453
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
403f239ca5 |
refactor: remove QuickJS skills runtime (#508)
* refactor: remove quickjs skills runtime * style: apply repo formatting * refactor: clean up error reporting and connection handling - Removed the 'skill' source option from the error report structure to streamline error reporting. - Refactored the ConnectionsPanel component to simplify connection status badge rendering and improve clarity. - Updated the CronJobsPanel to enhance logging for cron job loading processes. - Adjusted SkillCard component to use a more consistent type for icons. - Deleted outdated end-to-end tests for Gmail and Notion skills, improving test suite maintainability. * fix: remove unnecessary ESLint disable comment in Conversations component - Cleaned up the Conversations component by removing the ESLint disable comment for exhaustive dependencies in the useEffect hook, improving code clarity and maintainability. * fix: remove unnecessary whitespace in Conversations component - Eliminated an extra line of whitespace in the Conversations component, enhancing code readability and maintainability. * refactor: streamline SkillCard imports for improved clarity - Combined import statements in the SkillCard component to enhance code readability and maintainability. |
||
|
|
759691e380 |
feat(composio): improve toolkit sync and connection handling (#507)
* Enhance release workflow with build target input and improved job structure - Added a new input parameter `build_target` to specify the environment (production or staging) for the release process. - Made `release_type` input optional with a default value of `patch`. - Refactored job names and dependencies to reflect the new build target logic, including conditional steps for production and staging environments. - Introduced a `resolve` step to determine build outputs based on the selected environment, enhancing the workflow's flexibility and clarity. - Updated the `create-release` job to depend on the new `prepare-build` job, ensuring proper execution flow based on the build target. * feat(composio): enhance Composio integration with toolkit management and testing - Added `KNOWN_COMPOSIO_TOOLKITS` constant to facilitate access to available toolkits. - Implemented unit tests for `useComposioIntegrations` to ensure correct behavior during toolkit and connection fetching, including error handling scenarios. - Updated `Skills` page to utilize the new `KNOWN_COMPOSIO_TOOLKITS` for improved toolkit display logic. - Refactored hooks to handle connection errors gracefully and maintain toolkit visibility. - Enhanced backend integration by updating Composio client configuration to streamline toolkit management. * refactor(dispatch): remove channel delivery instructions for Telegram - Deleted the `channel_delivery_instructions` function, which provided response guidelines for Telegram messages. This change simplifies the message processing logic in the `process_channel_message` function by eliminating unnecessary instructions, enhancing clarity and maintainability. * refactor(composio): simplify Composio client configuration and remove toggles - Updated the `build_composio_client` function to remove unnecessary configuration checks, as Composio is always enabled when the user is signed in. - Revised the `resolve_client` function to clarify error handling related to user authentication. - Streamlined the `IntegrationsConfig` structure by removing toggles for Composio and related backend settings, ensuring a consistent configuration approach across integrations. - Adjusted tests to reflect the removal of integration toggles and focus on core API key usage. * refactor(composio): remove composio disabled state and improve error handling - Eliminated the `disabled` state from the `useComposioIntegrations` hook, as Composio is always enabled when the user is authenticated. - Updated error handling to surface backend connection issues directly, replacing previous checks for a disabled state. - Revised tests to reflect the new error handling logic, ensuring clarity in toolkit fetch error reporting. * refactor(integrations): update authentication handling for client configuration - Revised the `build_client` function to prioritize app-session JWT for user authentication, enhancing clarity in the fallback mechanism to `config.api_key`. - Improved error messages in `resolve_client` and `build_client` to provide clearer guidance on authentication issues related to session tokens. - Streamlined comments and documentation to reflect the updated authentication flow, ensuring consistency across integration components. * refactor(composio): unwrap CLI envelope for API responses - Introduced a new `unwrapCliEnvelope` function to handle the response format from the Rust side, allowing for easier access to the flat shapes defined in `./types`. - Updated `listToolkits`, `listConnections`, `listTools`, `authorize`, `deleteConnection`, and `execute` functions to utilize the new unwrapping logic, improving response handling consistency across the Composio API. - Enhanced error handling by ensuring that responses without logs pass through unchanged, maintaining backward compatibility. * refactor(skills_agent): update agent description and enhance Composio tool integration - Revised the `when_to_use` description in `agent.toml` to clarify the role of the Skills Agent as a service integration specialist, emphasizing its capability to execute both Composio and QuickJS skill tools. - Expanded the `prompt.md` documentation to detail available tool surfaces and typical Composio flow, improving clarity on how to interact with external services. - Implemented category overrides for Composio tools in `tools.rs` to ensure they are recognized as part of the Skill category, allowing proper access through the skills sub-agent. - Added tests to verify that Composio tools are correctly filtered and accessible by the skills sub-agent, ensuring robust integration and functionality. * style: apply formatter output from pre-push checks * refactor(tests): update Gmail and Notion integration tests for composio - Refactored tests for the Skills page to integrate Gmail and Notion as composio tools, enhancing the testing framework. - Removed mock data and streamlined the test setup to reflect the current state of available skills. - Updated assertions to verify the rendering of connected and disconnected states for Gmail and Notion integrations, respectively. - Improved clarity and maintainability of test cases by consolidating mock implementations and removing redundant code. * refactor(skills): update toolkit categorization and enhance test assertions - Modified the Skills component to assign categories dynamically based on toolkit metadata, improving organization of displayed tools. - Updated test cases to reflect the new categorization, ensuring accurate rendering of tools under their respective categories instead of a generic 'Other' group. - Enhanced assertions in tests to verify the presence of specific tools and their categories, improving test coverage and reliability. * refactor(skills): streamline item creation in Skills component - Simplified the item creation logic in the Skills component by removing unnecessary line breaks, enhancing code readability without altering functionality. - This change contributes to cleaner code structure and maintainability in the Skills page. * refactor(tests): enhance Gmail and Notion integration tests for improved clarity - Updated the integration tests for Gmail and Notion on the Skills page to utilize the `within` function for more precise querying of elements within their respective sections. - Improved test assertions to ensure that the connected and disconnected states are accurately verified, enhancing the reliability of the tests. - Streamlined the test setup to better reflect the current structure of the Skills component, contributing to overall test maintainability. * feat(skills): enhance Composio integration error handling and logging - Added error handling for Composio integrations in the Skills component, displaying a user-friendly message when integration status is stale or an error occurs. - Implemented logging in development mode to provide insights into the state of Composio toolkits and connections, aiding in debugging. - Updated the item rendering logic to reflect the error state, ensuring users can retry fetching integrations when an error is detected. - Enhanced tests to verify the display of error messages and the functionality of the retry mechanism, improving overall test coverage and reliability. |
||
|
|
362d0a014f |
fix(context): preserve cache-boundary metadata through prompt builders (#506)
* feat(prompt): enhance system prompt handling with cache boundary support - Updated the `SystemPromptBuilder` to include a new method for building prompts with cache metadata, returning a `RenderedPrompt` struct that contains the prompt text and an optional cache boundary index. - Introduced a constant for the cache boundary marker to improve readability and maintainability. - Modified the `extract_cache_boundary` function to cleanly remove the cache boundary marker from the rendered prompt while returning the relevant index. - Updated tests to validate the new cache boundary functionality and ensure that the system prompt does not leak internal markers. - Adjusted the agent's session management to track the cache boundary, enhancing the overall prompt handling and memory context management. * feat(context): add method to build system prompt with cache metadata - Introduced `build_system_prompt_with_cache_metadata` method in `ContextManager` to assemble the opening system prompt while preserving cache-boundary metadata for improved request prefix caching. - Updated imports to include `RenderedPrompt` for the new functionality, enhancing the context management capabilities. * feat(prompt): include cache boundary in subagent system prompt rendering - Updated `render_subagent_system_prompt` to insert a cache boundary marker before the workspace section, allowing typed sub-agents to maintain static instructions above this boundary for efficient prompt reuse. - Added a new test to verify that the cache boundary is correctly included in the rendered prompt, ensuring that the system prompt structure supports improved caching behavior. - Adjusted comments for clarity and updated the runtime banner numbering to reflect the new structure. |
||
|
|
afb95e72a5 |
refactor(event-bus): native typed request/response surface + channels encapsulation (#505)
* feat(event-bus): introduce typed request/response API for enhanced inter-module communication - Added a typed request/response surface to the existing event bus, allowing modules to execute requests through a shared controller registry. - Implemented `request_global` and `request_controller_global` functions for executing typed requests, enhancing the API's usability. - Updated documentation to reflect the new capabilities and usage patterns for the event bus, including when to use the request API versus traditional event publishing. - Added tests to validate the functionality of the new request/response features, ensuring robust integration with existing event bus operations. * refactor(event-bus): restructure event bus module and update references - Moved the event bus implementation from `src/openhuman/event_bus/` to `src/core/event_bus/`, establishing a clearer module hierarchy. - Updated all references throughout the codebase to reflect the new location of the event bus, ensuring consistency and reducing confusion. - Enhanced documentation to clarify the usage of the event bus and its core types, improving developer experience. - Introduced new files for event handling, requests, and subscribers, streamlining the event bus functionality and making it more modular. - Added tests to validate the new structure and ensure that the event bus operates correctly after the refactor. * refactor(event-bus): enhance event bus with native request/response surface - Updated the event bus to include a native, in-process typed request/response surface, allowing for zero serialization of Rust types and direct communication between modules. - Replaced the previous request API with a more streamlined approach using `register_native_global` and `request_native_global` functions. - Improved documentation to clarify the usage of the event bus, detailing when to use broadcast events versus native requests. - Removed the old request/response implementation to simplify the event bus structure and enhance maintainability. - Added examples and guidelines for registering and using native request handlers, improving developer experience and usability. * feat(agent): introduce native request handlers for agentic turns - Added a new `bus` module to encapsulate native event-bus handlers for the agent domain, including the `agent.run_turn` handler for executing agentic turns. - Updated the event bus registration process to include the new agent handlers, allowing for direct in-process request/response communication without serialization. - Refactored the channel message processing to dispatch agentic turns through the native bus, enhancing modularity and testability. - Improved documentation to clarify the usage of the new agent handlers and their integration with the event bus. - Added tests to validate the functionality of the new request handlers and ensure proper routing through the event bus. * refactor(tests): implement global bus handler lock for channel dispatch tests - Introduced a `use_real_agent_handler` function to manage the global bus handler lock during channel dispatch tests, ensuring exclusive access to the `agent.run_turn` handler. - Updated multiple test files to utilize the new handler function, improving test reliability by preventing race conditions during concurrent test execution. - Enhanced documentation to clarify the usage of the bus handler lock in tests that interact with the global native request registry. * feat(tests): add integration tests for Discord channel dispatch - Introduced a new test file `discord_integration.rs` to validate the end-to-end functionality of the Discord dispatch path within the channels module. - Implemented tests to ensure proper handling of inbound messages, reaction capabilities, and conversation history management specific to Discord. - Updated `mod.rs` to include the new Discord integration tests, enhancing overall test coverage for the channels module. * feat(tests): add Telegram integration tests for channel dispatch - Introduced new tests in `telegram_integration.rs` to validate the end-to-end functionality of the Telegram dispatch path within the channels module. - Implemented tests to ensure proper handling of threaded inbound messages, automatic acknowledgment reactions, and response routing through the agent bus handler. - Enhanced test coverage for Telegram, ensuring that the `supports_reactions()` capability is honored and that the dispatch pipeline operates correctly for both Telegram and Discord channels. * feat(tests): add testing utilities for event bus stubbing - Introduced a new `testing` module in the event bus to provide shared utilities for stubbing the global native bus registry. - Implemented `mock_bus_stub` and `MockBusGuard` to facilitate safe installation and restoration of stub handlers in tests, preventing race conditions. - Updated existing tests to utilize the new mocking utilities, enhancing test reliability and clarity in handling agent bus interactions. - Improved documentation to guide users on using the new testing features effectively. * style(event-bus): clean up formatting and remove unnecessary line breaks - Removed trailing whitespace and unnecessary line breaks in the event bus module files to improve code readability and maintainability. - Consolidated import statements and function definitions for a cleaner code structure across the event bus and agent modules. * fix(memory): point conversations/bus.rs at crate::core::event_bus Incoming `memory/conversations/bus.rs` from upstream/main still imports from the old `openhuman::event_bus` path. This branch relocated the bus to `core::event_bus`, so the merge left the import unresolved and the crate failed to compile. Rewire both references (`use` + fully-qualified `subscribe_global` call) to the canonical `crate::core::event_bus` path. |
||
|
|
986ff7ae4b |
feat(context): global context management module (#504)
* refactor(context): implement layered context pipeline with memory management - Introduced a new context pipeline orchestrator that manages context reduction stages before each provider call, including tool-result budgeting, microcompaction, and session memory management. - Added `ContextGuard` to monitor context utilization and trigger auto-compaction when thresholds are exceeded. - Implemented `microcompact` to replace older tool result payloads with placeholders, preserving API invariants while managing memory effectively. - Created `SessionMemory` to maintain persistent notes across sessions, updated by a background process to avoid impacting performance during user interactions. - Enhanced overall context management to improve memory efficiency and ensure smoother operation of the agent's functionality. * chore(context): add context module and refactor imports - Introduced a new `context` module to centralize context management for agent sessions, enhancing organization and maintainability. - Updated various files to replace references to the deprecated `context_pipeline` with the new `context` module, ensuring consistency across the codebase. - Bumped the version of the `openhuman` package to 0.52.4 in `Cargo.lock` to reflect these changes. * refactor(imports): update prompt module paths to new context structure - Refactored import statements across multiple files to replace references to the deprecated `agent::prompt` module with the new `context::prompt` module. - This change enhances code organization and aligns with the recent restructuring of the context management system. * feat(context): introduce comprehensive context management configuration - Added a new `ContextConfig` struct to manage global context settings, including budget thresholds, summarization triggers, and session-memory extraction parameters. - Implemented environment variable overrides for context management settings, allowing dynamic configuration. - Refactored the configuration loading process to accommodate the new context management structure, ensuring backward compatibility with existing configurations. - Introduced a `Summarizer` trait and default implementation for handling conversation summarization, enhancing the agent's ability to manage context effectively. - Updated related modules to integrate the new context management features, improving overall organization and maintainability of the codebase. * feat(context): implement ContextManager for enhanced session context handling - Introduced a new `ContextManager` struct to manage session context, including prompt assembly, context reduction, and summarization dispatch. - Added methods to `ContextGuard` for tracking input and output token counts, as well as the context window size. - Updated the `mod.rs` files to include the new `manager` module and its exports, improving organization and maintainability of the context management system. - Enhanced overall context management capabilities, allowing for more efficient handling of conversation context during agent interactions. * refactor(context): centralize system prompt assembly for channel runtimes - Moved the system prompt construction logic for channel runtimes into a new `channels_prompt` module within the `context` directory, ensuring all prompt-building code is organized in one location. - Updated various files to replace deprecated references to the old prompt module, enhancing code clarity and maintainability. - Introduced a new `build_system_prompt` function tailored for channel-specific requirements, including tool descriptions and channel-specific preambles, while maintaining byte-stability for cache efficiency. - Refactored related imports and tests to align with the new structure, improving overall organization of the context management system. * style(context): apply rustfmt to context module and integration sites Clippy caught a never-looping `while` in `snap_split_forward` that was really an `if` in disguise — collapse it. Everything else is pure `cargo fmt` cleanup on files touched during the context/ refactor. Full `cargo test --lib` sweep: 2253 tests passing. * delete(config): remove obsolete user configuration file - Deleted the `config.toml` file for user `69ccc8e95692bb0ddd56c10f`, which was no longer needed, streamlining the user configuration management. * refactor(SOUL.md): remove emoji usage guidelines to streamline prompt instructions - Deleted the section on emoji usage to simplify the prompt guidelines, focusing on clarity and direct communication with users. This change aims to enhance the overall effectiveness of the prompts by reducing unnecessary complexity. * refactor(config): enhance environment variable handling for context compaction settings - Improved the parsing and validation of environment variables for `OPENHUMAN_CONTEXT_COMPACTION_TRIGGER_PCT` and `OPENHUMAN_CONTEXT_HARD_LIMIT_PCT`, ensuring that the compaction trigger percentage is strictly less than the hard limit. - Added warnings for invalid percentage values and retained existing configuration values on failure. - Updated the context management logic to handle migration from deprecated fields, ensuring backward compatibility and clearer user guidance. * refactor(context): update tool result budget handling in Agent - Changed the source of the tool result budget to be derived from the ContextManager, ensuring it reflects the resolved `context.tool_result_budget_bytes` with any environment or configuration overrides. - This update removes reliance on the deprecated `agent.tool_result_budget_bytes` field, centralizing budget management for improved consistency across the context management system. * refactor(prompt): enhance subagent rendering options and system prompt construction - Introduced `SubagentRenderOptions` to manage per-definition rendering flags, allowing for more flexible control over the inclusion of identity, safety preamble, and skills catalog in the system prompt. - Updated `render_subagent_system_prompt` to utilize these options, ensuring that the rendering behavior aligns with the specified flags. - Modified `build_system_prompt` to accept an optional `channel_name`, improving clarity in channel capabilities messaging and removing hardcoded references to "Discord". - Adjusted tests to reflect changes in prompt construction, ensuring comprehensive coverage of the new rendering logic. * refactor(context): improve session memory handling in ContextManager - Updated session memory methods in `ContextManager` to use a locked reference for thread safety during extraction state changes. - Introduced a new method `session_memory_handle` to facilitate background task management of session memory extraction. - Adjusted the context statistics to utilize a snapshot of session memory, enhancing data integrity and consistency. * refactor(turn): improve extraction state management in Agent - Updated the extraction state handling to transition to "in-progress" instead of "complete" immediately, allowing for better control during background extraction tasks. - Introduced a shared handle for session memory to manage extraction completion and failure states, ensuring deltas are preserved for retries on failure. - Enhanced logging to provide clearer insights into the extraction process outcomes. |
||
|
|
462b4d2895 |
Move conversation persistence into workspace memory store (#503)
* Move conversation persistence into workspace memory store * Apply formatter updates for conversation persistence * Enhance conversation persistence by registering a new subscriber for channel events. Update domain subscriber registration to include workspace directory, ensuring proper message handling and persistence across channels. Refactor event structure to include message ID and reply target for improved tracking. Additionally, adjust module visibility for context management. |
||
|
|
73f8d1287a |
refactor: remove hardware-related components and streamline service management (#502)
* feat(config): introduce pre-login user directory structure - Added support for a pre-login user directory to encapsulate configuration, memory, and state before any user logs in. This ensures that all initial data is scoped under a dedicated user directory (`users/local`), preventing direct writes to the root `.openhuman` path. - Implemented the `pre_login_user_dir` function to return the appropriate path for the pre-login user. - Updated configuration loading logic to defer disk state creation until the first successful login, enhancing user data management and isolation. - Added tests to verify the correct behavior of the pre-login directory structure. * refactor: remove hardware-related components and streamline service management - Deleted hardware configuration and related tools from the codebase, including `HardwareConfig`, `HardwareTransport`, and associated memory management tools. - Introduced a new `service.ts` module for managing service and daemon commands, consolidating service-related functionalities. - Updated import paths across the application to reflect the removal of hardware references and the addition of the new service management module. - Refactored the `build_system_prompt` function to remove hardware access instructions, focusing on action instructions instead. - Cleaned up the Cargo.toml and Cargo.lock files by removing unused dependencies related to hardware management. * chore: apply formatting and tauri lockfile sync * refactor(tests): extract config file writing logic into a reusable function - Introduced a `write_config_file` function to encapsulate the logic for creating directories and writing configuration files, improving code reuse and readability. - Updated test cases to utilize the new function for writing configuration files, ensuring consistency and reducing duplication. - Added handling for pre-login user directory structure to ensure configuration is correctly written to the appropriate paths. |
||
|
|
1c7318603e |
feat(composio): backend-proxied Composio integration end-to-end (#501)
* feat(composio): add Composio integration module with OAuth support - Introduced a new Composio module for backend-proxied access to various OAuth integrations. - Implemented ComposioClient for handling API requests related to toolkits, connections, and actions. - Added RPC operations for listing toolkits, managing connections, and executing actions. - Registered Composio controllers and schemas for integration with the existing system. - Created a debug script for testing Composio OAuth flow against the live backend. - Updated Cargo.lock to version 0.52.3 to reflect the new changes. * feat(composio): implement Composio connection modal and API integration - Added ComposioConnectModal for managing Composio toolkit connections, mirroring the user experience of existing modals. - Introduced composioApi for backend communication, including functions for listing toolkits, managing connections, and handling OAuth authorization. - Created toolkitMeta for displaying metadata of Composio toolkits in the Skills grid. - Developed hooks for fetching and managing Composio integrations, ensuring real-time updates on connection status. - Updated Skills page to integrate Composio toolkits, providing a seamless user experience for connecting and managing integrations. * chore: update OpenHuman version to 0.52.3 and add debug-composio-trigger script - Bumped OpenHuman package version in Cargo.lock to 0.52.3. - Introduced a new script, debug-composio-trigger.mjs, for Socket.IO live listening of Composio trigger events, facilitating testing and debugging of webhook interactions. * style(composio): apply Prettier + rustfmt from pre-push hook * feat(composio): enhance ComposioConnectModal and improve polling logic - Updated ComposioConnectModal to handle various connection phases more effectively, including 'waiting' for pending connections. - Introduced new refs for managing polling state and in-flight requests to prevent overlapping executions. - Improved error handling during polling, providing clearer feedback on OAuth timeouts. - Added logic to resume polling if the modal opens while an OAuth handoff is in progress. - Refactored connection handling in the modal for better clarity and maintainability. * refactor(composio): improve JSON payload handling and connection verification - Updated debug-composio-login.sh to build JSON payloads using jq for safer handling of toolkit data. - Enhanced connection verification logic in debug-composio-trigger.mjs to prioritize newly created connections and provide warnings for missing toolkits. - Refactored composio operations in ops.rs to return a more explicit error type, improving clarity in error handling across RPC operations. |
||
|
|
57d307ad1e |
feat(agent): simplify harness + trim bundled prompts (#500)
* Add built-in agent definitions and prompts for orchestrator, planner, code executor, skills agent, researcher, critic, archivist, and tool maker - Introduced a new module for built-in agent definitions, allowing for easy addition of agents through a structured format. - Created TOML configuration files for each agent, detailing their properties such as id, display name, usage context, and tool specifications. - Developed corresponding prompt files that outline the responsibilities and rules for each agent, enhancing their functionality and usability. - Implemented a loading function to parse these definitions into usable agent structures, ensuring all agents are correctly initialized and validated. * Remove deprecated archetypes and related components from the multi-agent harness - Deleted the `archetypes.rs`, `dag.rs`, `executor.rs`, and `types.rs` files, which contained the definitions and implementations for agent archetypes, task DAGs, and execution logic. - Updated the `builtin_definitions.rs` and `definition.rs` files to remove references to the now-removed archetypes, streamlining the agent definition process. - Refactored the `mod.rs` file to eliminate unused modules and clarify the structure of the multi-agent harness. - Adjusted comments and documentation to reflect the removal of the DAG orchestration flow, emphasizing the current sub-agent delegation model. * Remove unused `ArchetypeConfig` from schema exports in `mod.rs` to streamline configuration management. * Implement context guard and memory management features in the harness module - Introduced `ContextGuard` to monitor context utilization and trigger auto-compaction when usage exceeds defined thresholds. - Added `scrub_credentials` function to sanitize sensitive information from tool outputs. - Implemented history management functions to trim conversation history and build compaction transcripts. - Created `build_tool_instructions` to generate tool usage guidelines for the system prompt. - Developed `memory_context` functions to manage user memory and relevant context for conversations. - Enhanced the `tool_loop` to integrate the new context management and tool invocation logic. This update improves the agent's ability to manage context effectively, ensuring efficient memory usage and secure handling of sensitive data. * Refactor agent module imports to use the harness instead of loop_ - Updated import paths in `dispatcher.rs`, `memory_loader.rs`, `mod.rs`, `dispatch.rs`, and `startup.rs` to replace references from the `loop_` module to the `harness` module. - This change enhances module organization and aligns with the recent architectural updates in the agent structure. * Remove cost tracking and context assembly modules from the agent structure - Deleted `cost.rs` and `context_assembly.rs` files, which handled token cost tracking and context assembly for the multi-agent harness, respectively. - Updated `mod.rs` files to remove references to the deleted modules, streamlining the agent's organization and focusing on essential components. - This cleanup enhances maintainability and aligns with the current architectural direction of the project. * Remove identity module and related configurations from the agent structure - Deleted the `identity.rs` file, which contained the AIEOS identity handling logic and related structures. - Updated `mod.rs` and other files to remove references to the deleted identity module, streamlining the agent's organization. - This change simplifies the codebase and aligns with the decision to rely solely on OpenClaw markdown files for identity management. * Enhance agent prompts with improved formatting and clarity - Added spacing for better readability in the prompts of the Archivist, Code Executor, Critic, Orchestrator, Researcher, Skills Agent, and Tool Maker agents. - Updated the table formatting in the Orchestrator prompt for consistency and clarity. - These changes improve the overall presentation and usability of agent documentation, making it easier for users to understand agent responsibilities and rules. * Remove REPL functionality and related components from the agent structure - Deleted the `repl.rs` file, which contained the implementation for the interactive REPL. - Removed references to the REPL in `cli.rs`, `mod.rs`, and other related files, streamlining the agent's organization. - Eliminated unused REPL session handling logic from the agent schemas and local AI operations, enhancing maintainability and focusing on essential components. - This cleanup aligns with the current architectural direction of the project, simplifying the codebase. * Refactor imports and clean up configuration exports - Updated import paths in `startup.rs` to include `host_runtime` from the correct module. - Streamlined the export statements in `mod.rs` and `schema/mod.rs` for better organization and clarity. - Removed unnecessary line breaks in the export lists to enhance readability and maintainability of the configuration schema. * Remove unused history management functions and clean up agent module exports - Deleted the `history.rs` and `session.rs` files, which contained functions for managing conversation history and session handling. - Updated `mod.rs` to remove references to the deleted modules and streamline the agent's organization. - Cleaned up import statements in `mod.rs` and other files to enhance clarity and maintainability of the codebase. * Add agent harness module with builder, runtime, and turn management - Introduced a new `harness` module for the `Agent`, encapsulating the `AgentBuilder`, runtime accessors, and turn lifecycle management. - Implemented the `AgentBuilder` fluent API for constructing `Agent` instances with customizable configurations. - Developed the `turn` method to handle user interactions, tool execution, and context management within the agent. - Created a `runtime` module for public accessors and utility functions related to agent operations. - Added comprehensive unit and integration tests to ensure the functionality of the new agent structure and its components. * Remove classifier and traits modules from the agent structure - Deleted the `classifier.rs` and `traits.rs` files, which contained the classification logic and core agent traits, respectively. - Updated `mod.rs` to remove references to the deleted modules, streamlining the agent's organization. - This cleanup enhances maintainability and focuses on essential components of the agent architecture. * Refactor agent prompt structure and remove unused files - Updated the `prompt.rs` file to streamline the orchestrator's logic by removing unnecessary tool documentation references and simplifying the file list. - Deleted several prompt files (`AGENTS.md`, `BOOTSTRAP.md`, `CONSCIOUS_LOOP.md`, `MEMORY.md`, `README.md`, `TOOLS.md`) to declutter the project and focus on essential components. - Adjusted the `harness/mod.rs` file to change module visibility from public to crate-level for better encapsulation. - These changes enhance maintainability and align with the current architectural direction of the project. * Refactor bootstrap file handling and update tests - Simplified the list of bundled prompt files in `prompt.rs` by removing references to `AGENTS.md` and `TOOLS.md`, focusing on essential identity files. - Adjusted the logic for injecting `MEMORY.md` to be optional, ensuring it only appears if it exists. - Updated test cases in `common.rs`, `identity.rs`, and `prompt.rs` to reflect the changes in the bootstrap file structure and ensure accurate validation of the new logic. - These modifications enhance clarity and maintainability of the codebase while aligning with the current project architecture. * Implement agent delegation tools and browser automation features - Introduced new tools for agent delegation, including `ArchetypeDelegationTool`, `AskClarificationTool`, `DelegateTool`, and `SkillDelegationTool`, enhancing the agent's ability to manage tasks through specialized sub-agents. - Added `SpawnSubagentTool` for delegating tasks to sub-agents, allowing for more complex workflows and improved task management. - Implemented browser automation capabilities with `BrowserOpenTool`, enabling secure opening of approved URLs in the Brave Browser. - Enhanced the `BrowserTool` with pluggable backends for improved automation and user interaction. - Added `ImageInfoTool` for extracting metadata from images, supporting future multimodal capabilities. - Organized tools into a new `impl` module structure for better maintainability and clarity in the codebase. * Refactor imports and clean up tool implementations - Updated import statements across various tool implementations to ensure consistency and clarity, particularly in the `traits.rs` file. - Removed unnecessary line breaks and adjusted module paths for better organization. - Cleaned up test files by removing trailing whitespace, enhancing code readability and maintainability. - These changes streamline the codebase and improve the overall structure of tool-related components. * Update built-in definitions and prompt handling for clarity and consistency - Revised the test for built-in agent definitions to dynamically reference the length of `BUILTINS`, enhancing maintainability. - Clarified documentation for the `system_prompt` field in `AgentDefinition`, emphasizing the default behavior and usage of inline prompts. - Simplified the `build_system_prompt` function by removing unnecessary conditional logic, ensuring a more straightforward return of the prompt string. * Enhance agent definition loading and documentation clarity - Updated the `load_file` function to reject definitions with missing or empty `system_prompt`, ensuring custom definitions are properly validated. - Added a new test to verify that definitions lacking a `system_prompt` are correctly rejected, improving robustness. - Clarified documentation for built-in definitions and TOML parsing, emphasizing compile-time guarantees and runtime checks. - Improved the logic for checking the existence of `MEMORY.md` to prevent errors from stray directories, enhancing file handling reliability. |
||
|
|
31297ad19d |
refactor(memory): remove GLiNER/GLiREL ingestion phase (#499)
* refactor(memory): replace GLiNER model with heuristic extraction - Removed GLiNER-related code and dependencies from the memory ingestion pipeline, transitioning to a heuristic-only extraction approach. - Updated documentation and comments to reflect changes in extraction methods. - Adjusted tests to ensure compatibility with the new heuristic extraction configuration. - Bumped version of the tokenizers dependency and updated Cargo.lock accordingly. * chore: apply formatting and tauri lockfile sync |
||
|
|
9118bfb5d6 |
Fix/skill start issue (#498)
* chore: update .gitignore and bump openhuman version to 0.52.2 - Added `overlay/src-tauri/target/` to .gitignore to prevent tracking of build artifacts. - Updated the openhuman package version from 0.52.0 to 0.52.2 in Cargo.lock files for both the main and app/src-tauri directories. - Enhanced entitlements for macOS Hardened Runtime to allow outbound HTTPS calls and server connections. - Refactored registry operations to use rustls explicitly, improving network reliability on macOS. * refactor(logging): improve debug message formatting in fetch_url_bytes function - Updated the logging statement in the fetch_url_bytes function to enhance readability by formatting the debug message across multiple lines. This change improves clarity in log outputs, making it easier to track the number of bytes fetched from URLs. * feat(skill-setup): enhance OAuth handling and skill status synchronization - Introduced a managed OAuth auto-advance mechanism to ensure it runs only once per login attempt, improving user experience during authentication. - Updated the SkillSetupWizard to handle skill runtime checks more effectively, ensuring that the skill starts correctly and transitions to the setup phase seamlessly. - Enhanced the useSkillSnapshot hook to provide a synthesized offline snapshot when the skill is not yet running, preventing UI stalls during loading. - Implemented background synchronization after OAuth completion to ensure users see fresh data immediately without blocking the UI. - Added tests to validate the new behavior for skills setup completion and status retrieval without requiring the skill to be started first. * refactor(skills): streamline setup_complete retrieval in handle_skills_status function - Simplified the retrieval of the `setup_complete` variable by removing unnecessary line breaks, enhancing code readability and maintainability. - This change improves the clarity of the function's logic without altering its functionality. * refactor(skills): simplify success message and remove initial sync from OAuth flow - Updated the success message in the SkillSetupWizard to remove references to background syncing, streamlining user communication. - Removed the initial sync trigger from the SkillManager after OAuth completion, shifting the responsibility for data synchronization to the user interface or cron jobs. - Adjusted comments in the desktopDeepLinkListener to reflect the new sync behavior, clarifying that initial data sync is no longer automatic. * fix(pr-498): address CodeRabbit review and CI failures - SkillSetupWizard: only show complete after startSetup succeeds; error on failures - hooks: merge prior snapshot into offline fallback; use const arrow for helper - E2E: reset skills_set_setup_complete in finally for isolation - json_rpc_e2e: assert oauth/complete returns start() result; add minimal start() - Skills page tests: mock screen-intelligence/autocomplete/voice hooks (CoreStateProvider) Made-with: Cursor * fix: address follow-up CodeRabbit (readiness poll, shared test mocks) - SkillSetupWizard: waitForSkillRunning after startSkill before startSetup/auth RPC - json_rpc_e2e: poll skills_status until running instead of fixed 400ms sleep - Consolidate Skills page vi.mocks in test/mockDefaultSkillStatusHooks.ts Made-with: Cursor * fix: CodeRabbit — legacy OAuth awaits setSetupComplete, const waitForSkillRunning, mock base - Legacy OAuth: await persistence before complete; error on failure; guard ref + reset on skillId - waitForSkillRunning: const arrow per TS style - mockDefaultSkillStatusHooks: offlineStatusBase spread for shared literals Made-with: Cursor |
||
|
|
7b457aa27d |
feat(agent): pure orchestrator pattern with per-skill delegation tools (#496)
* feat(agent): pure orchestrator pattern with per-skill delegation tools (#478) Refactors the main agent from a direct tool-calling model to a pure orchestrator that delegates all work through dynamically generated tools. Architecture changes: - Orchestrator only sees generated tools (notion, gmail, research, run_code, review_code, plan, spawn_subagent) — skill tools are architecturally unreachable from the main agent - Each installed skill auto-generates a delegation tool at build time (SkillDelegationTool) that routes to skills_agent with the correct skill_filter - Static archetype tools (research, run_code, etc.) delegate to their respective sub-agents - visible_tool_specs filters the function-calling schema sent to the provider, enforcing the orchestrator boundary at the API level Prompt changes: - Rewrote AGENTS.md as a lean orchestrator prompt — no more routing tables or agent_id instructions - Orchestrator skips TOOLS.md, MEMORY.md, HEARTBEAT.md (~6k tokens saved per turn) — subagents get tool specs from the registry - Workspace .md files auto-sync via builtin-hash mechanism so prompt updates ship automatically to existing installs Bug fixes: - ModelSpec::Hint now resolves to {hint}-v1 (e.g. agentic-v1) instead of hint:agentic which the backend rejected - validate_skill_filter now uses skill_id from the engine tuple instead of splitting on __ in the raw tool name (which always failed) - Memory context forwarded to subagents via ParentExecutionContext Observability: - Added [agent] tagged logs for tool responses, agent state transitions, and delegation decisions throughout turn.rs See docs/agent-prompt-architecture.excalidraw for the visual diagram. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: rustfmt orchestrator_tools.rs Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: cargo fmt Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address CodeRabbit review — dispatch guard, fork specs, decouple sync - Enforce visible-tool allowlist at dispatch time (not just schema) - Fork mode uses visible_tool_specs (not full registry) - De-duplicate spawn_subagent when extending orchestrator tools - Raw tool output moved to debug level, info level logs metadata only - Decouple workspace file sync from prompt rendering so skipped files still get synced to disk Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
8e8da17ad9 |
fix(voice): cross-platform microphone permission handling (#489) (#491)
* fix(voice): add cross-platform microphone permission handling (#489) Voice dictation in release DMG silently fails because the macOS hardened runtime enforces entitlements and the sidecar plist lacked the audio-input entitlement. This adds the entitlement, NSMicrophoneUsageDescription for the system permission prompt, and cross-platform microphone permission detection (CPAL device probe) with clear error messages on macOS, Windows, and Linux. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(permissions): use plist file for infoPlist and fix cross-platform warnings (#489) - infoPlist expects a file path, not inline JSON — create Info.plist with NSMicrophoneUsageDescription and reference it as a string - Move Microphone permission request out of macOS-only cfg block since request_microphone_access() is cross-platform (fixes unused import warning on Linux CI) - Treat persistent Unknown mic permission as Denied per CodeRabbit review Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Cyrus Gray <144336577+graycyrus@users.noreply.github.com> |
||
|
|
49cbbbebaf |
Fix/skill start issue (#493)
* chore: update .gitignore and bump openhuman version to 0.52.2 - Added `overlay/src-tauri/target/` to .gitignore to prevent tracking of build artifacts. - Updated the openhuman package version from 0.52.0 to 0.52.2 in Cargo.lock files for both the main and app/src-tauri directories. - Enhanced entitlements for macOS Hardened Runtime to allow outbound HTTPS calls and server connections. - Refactored registry operations to use rustls explicitly, improving network reliability on macOS. * refactor(logging): improve debug message formatting in fetch_url_bytes function - Updated the logging statement in the fetch_url_bytes function to enhance readability by formatting the debug message across multiple lines. This change improves clarity in log outputs, making it easier to track the number of bytes fetched from URLs. |
||
|
|
a2fb1119ea |
refactor(skills): unify auth/oauth handshake on start({validate}) and drop enabled flag (#484)
* refactor(preferences): streamline skill preference management by removing the enabled toggle - Removed the `enabled` field from `SkillPreference`, simplifying the preference model to focus solely on `setup_complete`. - Updated related methods and RPC handlers to reflect this change, ensuring that skills are automatically started based on the completion of their setup process. - Adjusted tests to validate the new behavior, ensuring consistent functionality without the `enabled` toggle. - Enhanced documentation to clarify the new preference management approach. * refactor(auth): simplify authentication flow by removing onAuthComplete hook - Updated the `handle_auth_complete` function to eliminate the separate `onAuthComplete` JavaScript hook, streamlining the authentication process. - Revised the flow to directly inject new credentials and validate them using the `start` function, which now handles both validation and activation. - Enhanced rollback logic to ensure temporary credentials are cleared if validation fails, preventing persistence on disk. - Improved documentation to clarify the new authentication steps and their implications for credential management. * refactor(oauth): streamline OAuth credential handling in `handle_oauth_complete` - Removed the `build_start_credentials_arg` function and integrated its logic directly into `handle_oauth_complete`, simplifying the flow. - Updated the OAuth handling to validate credentials before persisting them, ensuring that only successful validations are saved. - Enhanced rollback logic to clear temporary credentials if validation fails, preventing incorrect state persistence. - Improved documentation to clarify the new steps in the OAuth process and their implications for credential management. |
||
|
|
0cd0f7a670 |
feat(voice): sync overlay orb with chat voice button state (#487) (#490)
* feat(voice): sync overlay orb with chat voice button state (#487) The overlay orb already reacts to hotkey-based dictation via Socket.IO events, but the chat "Start Talking" button used local React state only. Add a new RPC method `openhuman.overlay_stt_notify` that the chat button calls at each voice state transition, which publishes to the existing DICTATION_BUS / TRANSCRIPTION_BUS broadcast channels — so the overlay reflects recording/transcribing/idle from both input paths with zero changes to the Socket.IO bridge or overlay event handlers. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(voice): address CI formatting and CodeRabbit review feedback - Run cargo fmt and prettier to fix formatting violations - Use typed enum OverlaySttState instead of raw String for state param (serde rejects invalid states at deserialization, eliminating the unknown state branch) - Require `text` field for transcription_done state (return error if missing instead of silently ignoring) - Replace raw transcript logging with metadata-only (has_text, text_len) to avoid logging sensitive user speech content - Use "Voice input active" aria-label (covers recording + linger phases) - Convert notifyOverlaySttState to arrow function with async/await per repo TS conventions Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
aee9c52e88 |
fix(channels): Telegram threading, live listeners, core restart, and webhook cleanup (#485)
* feat(config): enhance world-readable config warning mechanism - Introduced a new static variable to track previously warned world-readable config files, preventing duplicate warnings. - Updated the warning logic to only log a warning for each unique world-readable config file, improving log clarity and reducing noise. - Added new `ChannelReactionReceived` and `ChannelReactionSent` events to the DomainEvent enum, expanding event handling capabilities in the event bus. - Included tests for the new reaction events to ensure proper functionality and integration. * feat(logging): add log file constraints and event filtering - Introduced functions to parse log file constraints from environment variables and filter log events based on these constraints. - Enhanced the `init_for_cli_run` function to apply the new filtering logic, improving log management and clarity. - Updated the `conversation_history_key` function to include thread context for Telegram, ensuring accurate message targeting. - Added a new trait method `supports_reactions` to the `Channel` trait, indicating support for emoji reactions. - Implemented integration tests for Telegram channel features, including reaction handling and thread message forwarding. * feat(telegram): enhance message handling with reactions and typing indicators - Added support for emoji reactions in Telegram responses, allowing for contextual acknowledgment of user messages. - Implemented a decision heuristic for when to use reactions, improving user interaction quality. - Introduced a typing indicator that activates immediately upon receiving a message, providing instant feedback to users. - Updated the channel delivery instructions to include new reaction syntax and guidelines for usage. - Enhanced tests to cover new reaction handling and message acknowledgment features, ensuring robust functionality. * fix(tests): update route key for Telegram message handling tests - Changed the route key in tests from `telegram_alice` to `telegram_alice_chat-1` to match the updated `conversation_history_key` format for Telegram. - This adjustment ensures accurate routing and consistency in message handling tests. * refactor(tests): streamline message handling in runtime tool calls - Refactored the message handling tests to utilize a `ChannelMessage` struct for improved clarity and maintainability. - Updated the route key generation to use the `conversation_history_key` function, ensuring consistency in message routing. - Simplified the invocation of `process_channel_message` by directly passing the constructed message, enhancing readability. * fix(telegram): enhance finalize_draft method to support thread context - Updated the `finalize_draft` method in the `Channel` trait and its implementation for `TelegramChannel` to accept an optional `thread_ts` parameter, allowing for message threading. - Adjusted related message handling functions to utilize the new parameter, ensuring proper message context during sending. - Modified tests to reflect changes in the `finalize_draft` method signature, enhancing the robustness of message handling in threaded conversations. * refactor(tests): ran format * feat(discord, telegram): implement core process restart on channel connection - Added functionality to restart the core process when a channel connection requires a restart, enhancing the user experience by automating the process. - Implemented error handling to log any issues during the restart, ensuring users are informed to restart the app if necessary. - Updated both Discord and Telegram configuration components to include this new behavior, improving consistency across channel integrations. * feat(core-update): enhance core update logging and error handling - Added warnings for outdated sidecar versions and potential mismatches in UI features, improving user awareness of version compatibility. - Implemented detailed error logging for failed attempts to fetch the latest core release, providing users with clear instructions for manual updates if necessary. - Enhanced logging for reusing existing core RPC endpoints, alerting users to potential issues with stale connections. * feat(channels): implement real-time channel listeners and enhance logging - Added support for real-time channel listeners for Telegram and Discord, ensuring inbound bot messages are polled during `openhuman run`. - Introduced a method to check for configured listening integrations, preventing unnecessary listener spawning when not needed. - Enhanced logging for channel connection events and message handling, providing better visibility into channel operations and user interactions. - Updated the Telegram channel connection to log the count of allowed users and mention-only settings for improved debugging. * chore(dependencies): update openhuman version to 0.51.18 and refactor imports in channel config components - Bumped the openhuman package version from 0.49.17 to 0.51.18 in Cargo.toml and Cargo.lock. - Refactored import statements in DiscordConfig.tsx and TelegramConfig.tsx to maintain consistency and ensure proper functionality. * Implement webhook deletion for long polling in TelegramChannel - Added `delete_webhook_for_long_polling` method to clear the Bot API webhook, enabling `getUpdates` long polling. - Updated error handling in `fetch_bot_username` to call the new method when a 409 conflict indicates an active webhook, allowing for retries after webhook deletion. - Enhanced logging for better traceability of webhook deletion and polling conflicts. * Refactor Discord and Telegram connection handling to ensure channel connection updates are dispatched regardless of restart requirement. Improved error handling during core process restart and enhanced logging for connection status. |
||
|
|
6410db1fad |
feat(thu-fullrun): overlay attention, skills sync, credits & settings refresh (#479)
* Update Conversations component to enhance user messaging for budget limits. Changed the warning text for exhausted weekly inference budget to improve clarity and user experience. * feat(schemas): add new configuration option for vision model usage - Introduced a new optional boolean field `use_vision_model` in the schemas for enabling vision LLM for screenshot analysis. - Updated the screen intelligence schemas to include a required `consent` field for starting sessions, replacing the previous `sample_interval_ms` field. - Enhanced the `ttl_secs` field description for clarity and modified the `capture_policy` to `screen_monitoring` for better understanding of its purpose. * feat(CoreStateProvider): enhance state management with optimistic updates and error handling - Implemented optimistic local commits for `setAnalyticsEnabled` and `setOnboardingCompletedFlag` to provide instant UI feedback while ensuring state consistency through authoritative snapshot refreshes. - Added error handling for the `refresh` function calls in `setAnalyticsEnabled`, `setOnboardingCompletedFlag`, and `clearSession` to log failures, improving robustness in state management during user interactions. - Updated dependencies in the `useCallback` hooks to include `refresh`, ensuring proper state updates and synchronization with the core. * feat(paths): centralize runtime path resolution for user-scoped skills data - Introduced a new module `paths.rs` to handle the resolution of runtime paths for skills, ensuring that `skills_data` and `workspace` directories are scoped per user. - Updated `bootstrap_skill_runtime` and `bootstrap_skills_runtime` functions to utilize the new path resolution logic, improving consistency and clarity in directory management. - Enhanced error logging for directory creation failures to include the specific path that failed, aiding in debugging. - Added a new optional field `overlay_ttl_ms` in the autocomplete schemas to support overlay time-to-live configuration. * feat(ScreenIntelligencePanel): optimize config synchronization to prevent user edit clobbering - Introduced a reference to track the last synced configuration signature, ensuring that user edits are preserved during periodic updates from the CoreStateProvider. - Updated the effect to compare serialized configuration values, allowing for re-sync only when actual changes occur, enhancing user experience and preventing unintended data loss. * feat(SkillManager): implement initial sync after OAuth completion - Added functionality to trigger an initial data sync immediately after OAuth completion, ensuring users see fresh data without waiting for the next scheduled sync. - Updated comments to clarify the change in sync behavior due to recent modifications in the Rust core, which no longer auto-triggers sync on OAuth completion. * fix(UsageLimitModal, Conversations): enhance user messaging for budget limits - Updated warning messages in both UsageLimitModal and Conversations components to provide clearer information regarding weekly limits and reset times. - Improved clarity in user notifications to enhance overall experience when budget limits are reached. * refactor(SkillSetupModal): improve session mode handling for skill configuration - Updated the SkillSetupModal component to lock the mode at mount time, ensuring users remain in the setup wizard during their session even if the skill is marked as complete. - Simplified mode management by replacing the forceSetup state with a sessionMode state, allowing explicit mode switching while maintaining a consistent user experience. * feat(SkillManager): enhance setup flow for OAuth-based skills - Updated the `startSetup` method to handle OAuth-based skills more effectively by implementing a fallback to core RPC for skills without a frontend runtime. - Improved error handling to treat missing `onSetupStart` implementations as successful completion for pure OAuth skills, allowing the setup wizard to display the "Connected!" screen. - Added detailed logging for both local runtime and core RPC fallback scenarios to improve traceability during the setup process. * feat(Home): enhance local AI status handling and asset management - Introduced a new state for local AI assets, allowing for better tracking of model file readiness. - Updated the loading logic to fetch both local AI status and assets concurrently, improving performance and error handling. - Implemented a mechanism to hide the Local Model Runtime card once all models are fully downloaded, enhancing user experience. - Added comprehensive comments to clarify the logic behind model readiness checks based on asset states. * refactor(Credits): update credit balance structure and terminology - Renamed credit categories in the RewardsCouponSection and PayAsYouGoCard components for clarity, changing "General credits" to "Promo credits" and "Top-up credits" to "Team top-up." - Updated the credit balance API to reflect the new structure, replacing `balanceUsd` and `topUpBalanceUsd` with `promotionBalanceUsd` and `teamTopupUsd`. - Adjusted normalization logic in the credits API to accommodate the new credit balance fields. - Modified tests to ensure correct handling of the updated credit balance structure. * feat(Settings): reorganize billing settings and update descriptions - Added a new top-level billing section to the settings, promoting it out of the Account & Security category for better visibility. - Updated the description for the Account & Security section to remove billing references, focusing on recovery phrase, team management, and linked account access. - Adjusted the settings navigation to accommodate the new billing section, ensuring proper routing and user experience. * refactor(Config): change logging level from info to debug for environment overrides - Updated logging statements in the Config implementation to use debug level instead of info, reducing verbosity during runtime while maintaining necessary traceability for configuration loading. * feat(Overlay): implement overlay attention event handling and refactor overlay app structure - Introduced a new overlay module to manage attention events, allowing the core to publish messages to the overlay window. - Enhanced the OverlayApp component to handle dictation and attention events, improving user interaction with the overlay. - Refactored the overlay state management to support different modes (idle, stt, attention) and added auto-dismiss functionality for attention messages. - Removed the Browser Access Toggle from the Skills page, streamlining the UI and focusing on core functionalities. - Updated tests to reflect changes in the Skills component and removed unnecessary mocks related to browser access. * fix(OverlayBubbleChip): reset typewriter animation on new bubble identity - Updated the OverlayBubbleChip component to reset the typewriter animation correctly when a new bubble is displayed by using the `key` prop. - Refactored the cleanup logic in the useEffect hook to ensure proper interval management and state reset, enhancing the user experience with bubble transitions. * refactor(rest): streamline key_bytes_from_string function and improve readability - Simplified the condition for checking the ASCII key length and character restrictions in the key_bytes_from_string function. - Consolidated the import statements for base64 engines into a single line for better clarity. - Adjusted test data formatting for improved readability in the key_bytes_from_string_tests module. * enhance(logging): improve color detection logic for terminal output - Updated the color detection logic in the logging module to prioritize environment variables (`NO_COLOR`, `FORCE_COLOR`, `CLICOLOR_FORCE`) for better control over color output. - Added detailed comments explaining the color resolution order, enhancing code clarity and maintainability. * test(Home): add mock for openhumanLocalAiAssetsStatus in tests - Enhanced the Home and HomeBootstrapButtons test files by adding a mock implementation for openhumanLocalAiAssetsStatus, which resolves to an object with null result and empty logs. This improves the test setup for local AI asset status handling. * refactor(SkillSetupModal): improve session mode handling and loading state - Updated the SkillSetupModal component to ensure session mode is determined after the first snapshot resolution, preventing premature defaults to the setup wizard. - Introduced a loading state to display a message while waiting for the skill setup status, enhancing user experience during the modal's initial render. - Refactored the SkillManager to throw errors for real failures during setup, ensuring proper error handling and user feedback. * refactor(Config): simplify logging for invalid proxy scope values - Consolidated the logging statement for invalid OPENHUMAN_PROXY_SCOPE values into a single line, improving code readability while maintaining the warning functionality. |
||
|
|
6465f3d314 |
feat(agent): sub-agents, reasoning→agentic routing, layered context pipeline (#474)
* Enhance agent architecture with sub-agent support and memory optimizations - Refactored the `Agent` struct to use `Arc` for shared ownership of the provider, tools, and tool specifications, enabling efficient memory management and concurrent access. - Introduced a `NullMemoryLoader` to optimize memory usage for sub-agents, allowing them to operate without incurring the cost of memory recall. - Added new methods in the `Agent` implementation to facilitate sharing of the provider, tools, and tool specifications with sub-agents, enhancing their operational efficiency. - Implemented a new `SystemPromptBuilder` method for constructing prompts specifically for sub-agents, ensuring they receive tailored context while minimizing unnecessary information. - Established a framework for loading custom agent definitions from TOML files, allowing for dynamic agent configuration and specialization. - Introduced a `ForkContext` to support efficient sub-agent execution in fork mode, leveraging shared resources for improved performance and reduced token usage. * Enhance agent definition management and sub-agent functionality - Introduced a global `AgentDefinitionRegistry` to manage built-in and custom agent definitions from TOML files, ensuring idempotent initialization. - Added new RPC handlers for listing, fetching, and reloading agent definitions, improving the flexibility of agent management. - Refactored the `Agent` struct to streamline sub-agent execution, including enhancements to the task execution flow and context handling. - Updated the orchestrator configuration to support fork mode for sub-agents, optimizing resource usage and performance. - Improved error handling and logging for agent definition loading and initialization processes, enhancing system reliability. * Add end-to-end test for sub-agent spawning and response integration - Implemented a new asynchronous test to validate the full path of a parent agent issuing a `spawn_subagent` tool call. - The test ensures that the sub-agent's output is correctly folded into the parent's response, verifying the interaction between the parent agent and the sub-agent. - Enhanced the `AgentDefinitionRegistry` to support global initialization of built-in agents, ensuring consistent behavior across tests. - Updated the handling of tool calls and memory configuration to facilitate the new test scenario, improving overall test coverage for agent interactions. * Enhance agent tool filtering with category support - Introduced a new `category_filter` in `AgentDefinition` to restrict tool visibility based on their category (System or Skill). - Updated the `from_archetype` function to apply the category filter for the `SkillsAgent` archetype. - Modified `SubagentRunOptions` to include a `category_filter_override` for dynamic filtering during sub-agent execution. - Enhanced the `filter_tool_indices` function to incorporate category filtering logic, ensuring tools are correctly filtered based on their defined categories. - Updated relevant tests to validate the new category filtering functionality, improving overall test coverage for agent interactions. * Implement layered context reduction pipeline for agent - Introduced a new `context_pipeline` module to manage a layered context reduction strategy, enhancing memory efficiency during agent interactions. - Added stages for tool-result budgeting, history trimming, microcompaction, autocompaction, and session memory extraction, each with specific triggers and cache implications. - Updated the `Agent` struct to include a `context_pipeline` field, ensuring state persistence across turns. - Enhanced the `AgentBuilder` to initialize the context pipeline by default. - Implemented tests to validate the functionality and stability of the context pipeline, ensuring consistent behavior across agent sessions. - Refactored relevant components to integrate the new context management features, improving overall agent performance and memory handling. * Enhance agent context pipeline with tool result budgeting and microcompaction - Renamed variable for clarity in tool execution result handling. - Implemented a new stage in the context pipeline to apply a byte budget to tool results, ensuring efficient memory usage. - Added logging for budget application, including details on original and final byte sizes. - Integrated microcompaction stages before tool calls to manage history and reduce memory footprint, with appropriate logging for outcomes. - Updated the agent's session memory management to track turn counts, facilitating better resource handling across iterations. * Refactor agent context pipeline for session memory extraction and tool call management - Simplified method calls in the `Agent` struct for clarity and efficiency. - Enhanced session memory extraction logic to spawn a background archivist sub-agent when thresholds are met. - Improved context pipeline handling for tool call recording and usage tracking. - Updated documentation and comments for better understanding of session memory extraction process. - Refactored microcompact function for cleaner code structure and readability. * refactor(agent): split agent.rs into focused submodules - Convert agent.rs (1988 lines) into agent/ folder with six files: * types.rs — Agent + AgentBuilder struct defs * builder.rs — AgentBuilder fluent API + Agent::from_config factory * turn.rs — turn lifecycle, tool dispatch, context pipeline wiring * runtime.rs — public accessors, run_single/run_interactive, helpers * tests.rs — integration tests with shared fakes * mod.rs — glue + top-level `run` convenience function - Drop misc external inspiration references from doc comments in the context_pipeline module and fork_context — the files now stand on their own design language. * fix(agent): address review comments on sub-agent + context pipeline PR Inline comment fixes: - definition.rs: YAML/TOML inconsistency — module doc, PromptSource, source bookkeeping, and load() all now uniformly document the TOML format. - subagent_runner.rs: render_subagent_system_prompt previously only appended PromptSource::Inline bodies, silently dropping PromptSource::File content. Thread the preloaded archetype_body through as an explicit &str parameter so both source variants render. Drops the unused SystemPromptBuilder + tools_for_prompt wiring while we're here. - prompt.rs: remove DateTimeSection from SystemPromptBuilder::for_subagent. Local::now() would make the sub-agent system prompt change per call and break KV-prefix cache stability. Document the invariant. Nitpick fixes: - agent/turn.rs: drop the no-op mark_extraction_started() call — the immediately-following mark_extraction_complete() clears the in-flight flag anyway. - context_pipeline/pipeline.rs: call guard.record_compaction_success() on microcompact success so a prior streak of autocompaction failures doesn't leave the circuit breaker tripped after a successful reduction. - context_pipeline/tool_result_budget.rs: remove the unnecessary out.clone() in the truncation return — capture final_bytes first, then move out. - definition_loader.rs: replace brittle reg.len() == 10 with assert!(reg.len() > 1) plus the existing targeted .get() checks. - executor.rs: tracing::warn! on unknown sandbox override values so typos surface during development; explicitly accept "none" and empty string as valid defaults. - fork_context.rs: add parent_context_visible_inside_scope test mirroring fork_context_visible_inside_scope, with minimal stub Provider/Memory impls so the test stays self-contained. - schemas.rs: drop redundant serde_json::to_value(serde_json::json!(...)) wrapping in handle_list_definitions + handle_get_definition. - event_bus/events.rs: add SubagentSpawned/SubagentCompleted/SubagentFailed cases to all_variants_have_correct_domain. - tools/ops.rs: add all_tools_includes_spawn_subagent regression test. - tools/spawn_subagent.rs: sort/dedup the known skill list in place instead of cloning into a second vec. Tests: 2316 passed / 0 failed (up from 2314; two new tests added). fmt + clippy clean on all touched files. * udpate prompts * Enhance AgentBuilder and runtime with event context and interactive CLI improvements - Added `event_context` method to `AgentBuilder` for setting `session_id` and `channel` for `DomainEvent`s, improving event tagging and correlation. - Updated `run_interactive` method in `Agent` to dispatch messages through `run_single`, ensuring consistent lifecycle event handling and error sanitization for interactive turns. * Optimize configuration handling in AgentBuilder by lazily creating Arc for full config in reflection hook. This change reduces unnecessary cloning when learning is enabled, improving performance. * Add fork-mode test for sub-agent spawning in agent - Introduced a new test, `turn_dispatches_spawn_subagent_in_fork_mode`, to validate the behavior of the agent when spawning a sub-agent in fork mode. - The test ensures that the parent agent correctly processes the sub-agent's output and maintains the expected response sequence. - Enhanced the test setup with a mock provider and memory configuration to simulate the agent's environment effectively. * Refactor sync RPC handling in skills to treat missing onSync as no-op - Updated the `handle_sync` function to log a debug message and return a no-op response when a skill does not implement the `onSync` handler, preventing unnecessary RPC errors for skills that do not require periodic syncs. - This change improves logging clarity and reduces error noise in logs and dashboards for skills that are not designed to handle sync operations. |
||
|
|
a1f8bc55c4 |
Improve upsell flow and retire legacy overlay app (#473)
* fix(autocomplete): disable autocomplete feature by default - Updated the default configuration for the Autocomplete feature to be disabled instead of enabled in both the frontend and backend configurations. This change aims to improve user experience by preventing unintended activations of the autocomplete functionality upon application startup. * feat(overlay): enhance overlay window functionality and responsiveness - Updated the OverlayApp component to dynamically adjust its size and position based on the overlay status (idle/active), improving user interaction. - Introduced new constants for overlay dimensions and margins, enhancing visual consistency. - Implemented hover effects to adjust opacity, providing better visual feedback. - Refactored window resizing and positioning logic to ensure the overlay remains user-friendly and visually appealing across different scenarios. - Updated macOS window level to NSScreenSaverWindowLevel for improved behavior in fullscreen and multi-space environments. * feat(dependencies): add objc2-core-graphics and related packages - Introduced `objc2-core-graphics` as a new dependency in the Cargo.toml for macOS support. - Updated Cargo.lock to include `objc2-metal`, `block2`, and `libc` as dependencies for enhanced functionality. - Refactored overlay window configuration to utilize `CGShieldingWindowLevel`, improving window behavior in macOS environments. * refactor(billing): remove storage limits and update plan budgets - Removed `storageLimitBytes` from the `PlanMeta` interface and all plan definitions, simplifying the billing structure. - Updated the `Free` plan to have zero budgets for monthly and weekly usage, aligning with the new billing strategy. - Adjusted the `BillingPanel` and related components to conditionally display budget information based on the updated plan values. - Enhanced the `InferenceBudget` and `PayAsYouGoCard` components to reflect changes in budget handling and improve user messaging. - Updated tests to ensure consistency with the new billing logic and removed references to storage limits. * feat(upsell): enhance GlobalUpsellBanner and PayAsYouGoCard components - Added the GlobalUpsellBanner component to the App, improving user visibility of upgrade options. - Refactored PayAsYouGoCard to better handle credit balance calculations, separating promo and top-up credits for clarity. - Updated the UpsellBanner styles for a more consistent visual presentation. - Introduced normalization functions in creditsApi to ensure robust handling of credit balance data. - Added tests for creditsApi to validate the normalization logic and prevent UI crashes with missing data. * feat(upsell): reintroduce GlobalUpsellBanner in App and enhance UpsellBanner styling - Added the GlobalUpsellBanner back into the App component to improve user visibility of upgrade options. - Updated the UpsellBanner component to include a new `rounded` prop for customizable styling. - Removed dismissible functionality from the GlobalUpsellBanner, streamlining the user experience. - Enhanced visual presentation by adjusting CSS styles for better consistency. * refactor(overlay): remove obsolete overlay files and configurations - Deleted unused files including .gitignore, index.html, package.json, postcss.config.js, README.md, tailwind.config.js, tsconfig.json, vite.config.ts, and yarn.lock from the overlay directory. - Removed all source files related to the overlay functionality, including App.tsx, main.tsx, parentCoreRpc.ts, styles.css, and various components. - Cleaned up the src-tauri directory by removing configuration files, icons, and capabilities related to the overlay, streamlining the project structure. - This commit enhances maintainability by eliminating legacy code and unused resources. * refactor(rewards): move DISCORD_INVITE_URL to a separate utility file - Refactored the Rewards component to import the DISCORD_INVITE_URL from a new links utility file, improving code organization and maintainability. - Created a new links.ts file to centralize URL constants, enhancing clarity and reusability across the application. * style(app): apply formatter hook fixes * chore(dependencies): remove @heroicons/react from yarn.lock - Deleted the entry for @heroicons/react@^2.2.0 from yarn.lock, streamlining dependency management and reducing potential conflicts. |
||
|
|
3d17bfb23d |
Refactor core cross-domain flows onto event bus (#468)
* Refactor core cross-domain flows onto event bus * feat(agent): enhance error handling and message processing - Introduced a new constant for maximum error message length in the Agent. - Added methods for comparing conversation messages and determining new entries for turns. - Implemented a function to sanitize error messages, improving clarity and consistency in error reporting. - Updated the `run_single` method to utilize the new error handling and message processing logic, ensuring better tracking of conversation history and error states. * feat(supervision): initialize global event bus and register health subscriber - Added initialization of the global event bus with default capacity to ensure channel health events have a live bus and subscriber target. - Registered a health subscriber to enhance monitoring capabilities within the supervised listener context. |
||
|
|
63b17ca55c |
refactor: remove dead frontend modules and obsolete API layers (#469)
* refactor: remove unused socket, agent tool registry, daemon health, and API service files - Deleted the `useSocket` hook, `AgentToolRegistry`, `DaemonHealthService`, and various API service files including `actionableItemsApi`, `apiKeysApi`, `feedbackApi`, `inferenceApi`, `managedDmApi`, and `settingsApi`. - This cleanup reduces code complexity and improves maintainability by removing obsolete components that are no longer in use. * feat: add knip configuration and commands to package.json - Introduced new scripts for running knip in development and production modes in both package.json files. - Added a knip.json configuration file to specify entry points and project files for dependency analysis. - Updated yarn.lock to include new dependencies related to knip, enhancing the project's dependency management capabilities. This commit improves the project's tooling for managing dependencies and ensures better code quality through automated checks. * refactor: remove unused components and clean up dependencies - Deleted several unused components related to intelligence features, including ActionPanel, InputGroup, SectionCard, and ValidatedField, to streamline the codebase. - Removed mock data and country data files that are no longer in use, enhancing maintainability. - Cleaned up the package.json by removing the @heroicons/react dependency, which is no longer required. - This commit improves the overall project structure and reduces complexity by eliminating obsolete code. * refactor: remove IntelligenceApiService and redefine ConnectedTool interface - Deleted the `IntelligenceApiService` class and its associated backend API methods to streamline the codebase. - Introduced a local definition of the `ConnectedTool` interface in `useIntelligenceApiFallback.ts` for better encapsulation and clarity. - This refactor enhances maintainability by eliminating unused code and consolidating relevant types within the appropriate context. * style: format knip config * chore(dependencies): update OpenHuman to version 0.52.0 in Cargo.lock * feat: implement Daemon Health Service for polling and state management - Introduced the `DaemonHealthService` class to poll the Rust core health snapshot and synchronize the frontend daemon store. - Added methods for setting up a health listener, parsing health snapshots, and updating the daemon store based on health data. - Implemented a timeout mechanism to handle disconnection scenarios, enhancing the reliability of the daemon's health monitoring. - This addition improves the application's ability to maintain an accurate representation of the daemon's health status in real-time. * chore(knip): update entry points in knip configuration - Modified the `entry` field in `knip.json` to include `src/main.tsx` alongside existing test specifications. - This change ensures that the main application file is included in dependency analysis, improving project structure and tooling. * refactor: streamline code and enhance readability across multiple modules - Consolidated multiple `replace` calls into single calls using arrays for improved efficiency in text processing. - Simplified default implementations for several structs, removing redundant code. - Updated query mapping in database interactions to enhance clarity and maintainability. - Improved logging and error handling by refining how state and error messages are processed. - Enhanced the readability of various functions by restructuring conditional checks and simplifying logic. These changes collectively improve code maintainability and performance across the application. * Merge remote-tracking branch 'origin/fix/cleanup' into fix/cleanup * fix: update error handling in bootstrap_after_login function - Changed the error parameter in the inspect_err closure to an underscore to indicate it is unused. - This minor adjustment improves code clarity and adheres to Rust conventions for unused variables. |
||
|
|
30eec2ad88 |
docs: comprehensive documentation for core Rust modules (#470)
* feat(core): enhance RPC controller and dispatch logic - Added comprehensive documentation for the core RPC controller and dispatch modules, detailing their purpose and functionality. - Introduced new functions for managing registered controllers and schemas, including validation and invocation methods. - Improved the structure of the `RpcOutcome` type to standardize response formats across domain-specific handlers. - Enhanced the local AI operations module with additional functionalities for agent interactions, model management, and audio processing. - Updated the dispatcher to better route RPC calls to their respective handlers, ensuring a more robust and maintainable architecture. These changes improve the clarity and usability of the RPC system, facilitating easier integration and interaction with various components of the OpenHuman platform. * docs: comprehensive documentation for core Rust modules * chore(dependencies): update OpenHuman to version 0.52.0 in Cargo.lock and add knip dependency in package.json |
||
|
|
acc6246e59 |
Refactor core-polled app state and screen intelligence status (#464)
* refactor(accessibility): remove device control and predictive input features from accessibility settings - Updated accessibility-related components and tests to eliminate device control and predictive input features. - Adjusted AccessibilityPanel and ScreenIntelligencePanel to reflect the removal of these features. - Modified related tests to ensure consistency with the updated accessibility status structure. - Cleaned up accessibility session parameters and state management to focus solely on screen monitoring. * refactor(accessibility): streamline featureOverrides state initialization - Simplified the initialization of featureOverrides state in AccessibilityPanel and ScreenIntelligencePanel components for better readability. - Consolidated parameter definitions in startAccessibilitySession to enhance clarity and maintainability. - Removed unnecessary re-exports in the screen_intelligence engine module to clean up the codebase. * chore(dependencies): update OpenHuman to version 0.51.19 in Cargo.lock * chore(dependencies): update OpenHuman version to 0.51.19 in Cargo.lock * feat(restart): implement core process restart functionality - Added a new `SystemRestartRequested` event to the `DomainEvent` enum to handle restart requests. - Introduced a `RestartSubscriber` that listens for restart events and manages the process respawn. - Created a `service_restart` function to publish restart requests via the event bus. - Updated service schemas to include a new `restart` controller with parameters for source and reason. - Enhanced documentation to reflect changes in behavior and added necessary code comments. * feat(accessibility): add last restart summary to Screen Intelligence Panel - Introduced `lastRestartSummary` to the accessibility state and updated relevant components to display the last successful core restart information. - Modified `PermissionsSection` and `ScreenIntelligencePanel` to include the new summary. - Updated tests to validate the display of the last restart summary and ensure proper state management during core restarts. - Refactored accessibility slice to handle the new restart summary in state updates. * feat(core): enhance startup process with restart delay and subscriber registration - Added a call to apply startup restart delay from environment variables in `run_core_from_args`. - Updated the `bootstrap_skill_runtime` function to register a `RestartSubscriber` for handling restart requests, ensuring consistent respawn logic across triggers. - Introduced a new `core_process` field in the `AccessibilityEngine` to track the core process status, including its PID and start time. - Implemented a helper function to capture the core process start time using `OnceLock` for efficient initialization. * feat(screen-intelligence): refactor accessibility state management and UI components - Replaced direct Redux state access with a new `useScreenIntelligenceState` hook across multiple components, including `AccessibilityPanel`, `ScreenIntelligencePanel`, and their respective subcomponents. - Streamlined permission and session handling by consolidating related functions and removing unnecessary dispatch calls. - Updated tests to mock the new state management approach, ensuring consistent behavior and validation of UI elements. - Removed the `SessionAndVisionSection` component to simplify the structure and improve maintainability. - Introduced a new API file for screen intelligence to encapsulate related functionality and improve code organization. * refactor(tests): clean up and optimize test files for accessibility and screen intelligence panels - Removed redundant imports and streamlined the structure of test files for `AccessibilityPanel` and `ScreenIntelligencePanel`. - Consolidated core process state initialization in test mocks for better readability. - Updated dependency imports and ensured consistent mocking of state management hooks across tests. - Enhanced the `ScreenPermissionsStep` component by improving the dependency array in the useEffect hook for better performance. * refactor(store): remove unused authentication and user management code - Deleted the `UserProvider`, `authSlice`, `authSelectors`, `userSlice`, `teamSlice`, and related test files to streamline the codebase. - This cleanup enhances maintainability by removing legacy code that is no longer in use. - Updated the store configuration to reflect the removal of these slices and ensure proper state management. * refactor(webhooks): reorganize types and remove legacy state management - Moved `TunnelRegistration` and `WebhookActivityEntry` types to a new `types.ts` file for better organization. - Updated imports in `TunnelList` and `WebhookActivity` components to reference the new types location. - Refactored `useWebhooks` hook to eliminate Redux state management in favor of local state, enhancing performance and reducing complexity. - Removed unused `aiSlice`, `inviteSlice`, and `webhooksSlice` along with their associated tests to streamline the codebase. * refactor(daemon): migrate state management from Redux to a custom store - Introduced a new `store.ts` file to manage daemon state, replacing the previous Redux slice. - Updated components and hooks to utilize the new state management approach, enhancing performance and reducing complexity. - Removed the legacy `daemonSlice` and associated Redux logic, streamlining the codebase. - Adjusted imports in various components and hooks to reference the new store structure. * refactor(screen-intelligence): integrate core state management and enhance status handling - Replaced direct state management in `useScreenIntelligenceState` with a new core state approach, utilizing `useCoreState` for improved performance and consistency. - Updated status fetching and permission handling to leverage the core state snapshot, streamlining the logic and reducing redundant API calls. - Introduced a new `CoreRuntimeSnapshot` interface to encapsulate runtime statuses, including screen intelligence, local AI, autocomplete, and service states. - Adjusted related components and hooks to align with the new state management structure, enhancing maintainability and readability. - Updated tests to validate the new runtime state structure and ensure proper functionality across the application. * refactor(components): reorganize imports and streamline function formatting - Moved the import of `Tunnel` and `tunnelsApi` in `TunnelList.tsx` for better organization. - Reformatted function definitions in `store.ts`, `useDaemonHealth.ts`, `useDaemonLifecycle.ts`, `useWebhooks.ts` for improved readability. - Cleaned up the structure of test files in `coreRpcClient.test.ts` by consolidating object properties for clarity. - These changes enhance code maintainability and readability across the application. * test(screen-intelligence): fix duplicate hook imports * fix(tests): update ScreenIntelligenceDebugPanel test to use baseState for refresh status and vision calls * refactor(invites): simplify error message rendering in Invites component - Consolidated the conditional rendering of the load error message in the Invites component for improved readability. - This change enhances the clarity of the code without altering functionality. * refactor(daemon): streamline state management and function definitions - Removed the `healthTimeoutId` from the `DaemonUserState` interface and related functions to simplify state management. - Converted several functions in `store.ts` to arrow function syntax for consistency and improved readability. - Updated the `Invites` component to handle asynchronous loading and error states more effectively, ensuring that in-flight requests are properly managed. - Refactored the `CoreStateProvider` to enhance the refresh logic and prevent multiple simultaneous refreshes. - Introduced a new `register_domain_subscribers` function in `jsonrpc.rs` to centralize event bus subscriber registration, improving code organization and maintainability. * fix: add debug logging, atomic restart guard, and idempotent subscriber registration - CoreStateProvider: add namespaced debug logger for polling failure diagnostics - service/bus.rs: add AtomicBool gate to prevent duplicate restart spawns - service/bus.rs: use OnceLock for idempotent RestartSubscriber registration - Invites.tsx: add debug log in loadInviteCodes catch block * style: apply prettier formatting to CoreStateProvider * fix: sanitize error logging, serialize refresh, and demote restart logs - CoreStateProvider: sanitize error objects in poll failure logs to avoid leaking tokens/headers - CoreStateProvider: move in-flight guard into refresh() via shared promise so all callers (poll, updateLocalState, storeSessionToken) are serialized - CoreStateProvider: log refreshTeams errors instead of swallowing them - service/bus.rs: demote duplicate-restart log to debug, omit reason from log output to avoid free-form text emission * style: apply cargo fmt to service/bus.rs |
||
|
|
d66ee0d4de |
feat: consolidate overlay into desktop app and add compact orb demo (#450)
* feat(overlay): implement overlay window functionality and related RPC integration - Added a new OverlayApp component to handle overlay-specific UI and functionality. - Introduced an overlay window configuration in tauri.conf.json, allowing for a transparent, always-on-top overlay. - Implemented parent RPC communication for the overlay to interact with the main application. - Updated main.tsx to conditionally render the OverlayApp based on the current window context. - Enhanced CSS styles to support the overlay's visual requirements. This commit establishes the foundation for overlay functionality, improving user experience with a dedicated interface for specific tasks. * refactor(overlay): remove overlay functionality and related configurations - Deleted the overlay module and its associated files, including process management and configuration settings. - Removed environment variable checks and overlay-related logic from the core process and configuration schema. - Updated documentation to reflect the removal of overlay features, simplifying the codebase and improving maintainability. This commit streamlines the application by eliminating unused overlay components, enhancing overall performance. * feat(overlay): enhance overlay bubble functionality and styling - Added a new CSS animation for overlay bubble appearance, improving visual feedback. - Introduced an OverlayBubble interface to manage bubble properties such as tone and text. - Updated OverlayApp component to include a new OverlayBubbleChip for displaying messages with dynamic styling based on tone. - Adjusted overlay window dimensions in tauri.conf.json for a more compact design. This commit improves the user experience by providing visually distinct overlay messages and a refined interface. * feat(rotating-tetrahedron): add inverted color support and refactor canvas component - Introduced an optional `inverted` prop to the `RotatingTetrahedronCanvas` component, allowing for dynamic color changes based on the prop value. - Updated fill and edge materials to reflect the inverted state, enhancing visual customization. - Refactored the component to improve readability and maintainability by utilizing the new prop in the rendering logic. - Adjusted the effect dependencies to include the `inverted` prop for proper reactivity. This commit enhances the user experience by providing a more flexible and visually appealing tetrahedron display. * fix(rotating-tetrahedron): adjust opacity and emissive intensity for inverted colors - Updated the opacity of the fill material in the RotatingTetrahedronCanvas component to enhance visual clarity when the inverted prop is true. - Reduced the emissive intensity for the inverted state to improve the overall appearance of the tetrahedron. This commit refines the visual representation of the rotating tetrahedron, ensuring better contrast and aesthetics based on user preferences. * feat(rotating-tetrahedron): enhance dynamic color handling and performance - Implemented useRef hooks for fill and edge materials in the RotatingTetrahedronCanvas component to optimize rendering performance. - Updated the useEffect hook to adjust material properties based on the inverted state, improving visual consistency. - Refactored animation speed handling to utilize a reference for smoother updates. - Cleaned up resource management by ensuring materials are disposed of correctly when the component unmounts. This commit enhances the visual fidelity and performance of the rotating tetrahedron, providing a more responsive and visually appealing experience. * fix(overlay): adjust bubble alignment and overlay positioning - Changed the text alignment of the OverlayBubbleChip component from left to right for improved readability. - Updated the vertical positioning logic in the Tauri overlay to account for a right margin, ensuring consistent placement of the overlay window. These adjustments enhance the visual presentation and positioning of overlay elements, contributing to a better user experience. * fix(overlay): update overlay dimensions and bubble styling - Adjusted the overlay dimensions for a more refined appearance. - Modified the bubble tone classes for improved color consistency and readability. - Enhanced the text size and line height in the OverlayBubbleChip component for better visual clarity. - Updated the orb button size and styling to enhance user interaction. These changes contribute to a more polished and user-friendly overlay experience. * feat(overlay): enhance overlay scenario handling and text display - Introduced a new scenario management system in the OverlayApp component, allowing for dynamic cycling between different overlay states. - Added a new text display feature for scenario two, providing real-time feedback as the text is typed out. - Refactored the bubble rendering logic to accommodate the new scenario structure, improving the overall user interaction experience. These changes enhance the functionality and interactivity of the overlay, making it more engaging for users. * fix(overlay): reorder demo scenarios * fix(overlay): update overlay dimensions and improve text display logic - Adjusted overlay dimensions for better visual consistency. - Enhanced text display in OverlayBubbleChip to show text progressively based on bubble content. - Refactored the handling of scenario text in OverlayApp to streamline the display logic. These changes contribute to a more polished and engaging user experience in the overlay. * refactor(overlay): simplify type parameters in RPC functions - Removed unnecessary generic type parameter from `unwrapCliCompatibleJson` and `callParentCoreRpc` functions for improved clarity and conciseness. - These changes enhance code readability and maintainability without altering functionality. |
||
|
|
0b23ae6b96 |
fix(voice): resolve dictation pipeline in embedded Tauri app (#466)
* fix(voice): guard dictation_listener when voice_server is active macOS only supports one rdev::listen() global event tap per process. When voice_server.auto_start is true, skip starting the separate dictation_listener — the voice server owns the single listener and forwards hotkey events itself. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(voice): add TRANSCRIPTION_BUS broadcast channel Make publish_dictation_event public so the voice server can forward hotkey events. Add a new TRANSCRIPTION_BUS broadcast channel with subscribe_transcription_results() and publish_transcription() for delivering completed transcriptions to frontend clients via Socket.IO. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(voice): forward hotkey events and route transcription delivery In the voice server's hotkey handler, forward events to the dictation bus so Socket.IO clients receive dictation:toggle even without the separate dictation_listener running. In process_recording_bg, detect when the OpenHuman app is focused and deliver transcription via Socket.IO instead of OS-level Cmd+V paste, preventing text from disappearing into the unfocused WebView. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(voice): bridge transcription results to Socket.IO Subscribe to TRANSCRIPTION_BUS and emit dictation:transcription events to all connected Socket.IO clients. This completes the Rust-side pipeline for delivering transcribed text to the frontend. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(socket): queue listeners registered before socket connects Add a pendingListeners queue to socketService so that on()/once() calls made before the socket is established are replayed once the connection opens, preventing silently dropped event listeners. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(voice): use dedicated unauthenticated socket for dictation events Replace the auth-gated socketService dependency with a direct Socket.IO connection to the core process (127.0.0.1:7788). This bypasses the SocketProvider auth requirement and ensures dictation:toggle and dictation:transcription events are received regardless of login state. Dispatches the existing dictation://insert-text DOM event to bridge transcribed text into the Conversations chat input. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(voice): fallback to default audio config when preferred config fails macOS may advertise a 16kHz F32 config as supported but reject it at stream creation time (known cpal quirk). Add a fallback path that retries with the device's default_input_config(), handling F32, I16, and U16 sample formats with proper resampling. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(voice): add 3s timeout on Ollama LLM cleanup Wrap the Ollama inference call in a 3-second tokio::time::timeout so dictation feels responsive. If cleanup doesn't complete in time, fall back to raw Whisper text immediately instead of blocking for 2+ minutes. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
f0118ab674 |
fix(autocomplete): graceful shutdown on app exit (#460)
* fix(autocomplete): graceful shutdown on app exit Stop the autocomplete engine and quit the Swift overlay helper process when the core receives a shutdown signal, preventing orphan processes. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor(core): extract shutdown logic into generic facility Move domain-specific autocomplete cleanup out of the jsonrpc adapter into a new `core::shutdown` module with a hook registry. Adds SIGTERM handling alongside SIGINT via `tokio::select!` on Unix platforms. Addresses CodeRabbit review feedback on PR #460: - jsonrpc.rs stays transport-only (no domain logic) - Process responds to both SIGINT and SIGTERM gracefully Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(autocomplete): suppress duplicate error notifications and auto-stop after repeated failures The autocomplete polling loop was showing an error notification badge on every single refresh cycle when a dependency was unavailable (e.g. Ollama not running, macOS Automation permission denied). This caused floods of identical macOS notifications. Changes: - Track last notified error message; skip badge if identical to previous - Count consecutive errors; auto-stop engine after 5 failures with a clear log message, preventing endless notification spam - Reset counters on successful refresh or engine restart Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
dc5e7adeb6 |
fix(auth): update RPC method names for authentication calls (#463)
* fix(auth): update RPC method names for authentication calls Refactor authentication-related RPC method names to use underscores instead of dots for consistency. Updated methods include `get_state`, `get_session_token`, `clear_session`, and `store_session`. chore: update OpenHuman version to 0.51.19 style: standardize string formatting in quickjs_libs/bootstrap.js and other files - Replace single quotes with double quotes for string literals in various functions. - Ensure consistent formatting across console logging and error handling. fix(config): improve token retrieval logic in ops_core.rs - Enhance the logic for retrieving the active session token from the credentials store, accommodating user-specific directories. * test: align auth and OAuth assertions with current behavior Update stale test expectations for underscore-style auth RPC methods and light-theme OAuth button classes, and make the bypass-login E2E assertion resilient to the current auth persistence model. Made-with: Cursor * chore: apply formatter output for login flow spec Include Prettier formatting adjustments produced by the pre-push hook so the branch can pass repository push checks cleanly. Made-with: Cursor |
||
|
|
9764a875e8 |
fix(subconscious): seed defaults into per-user workspace + fix Intelligence page stale log (#462)
* fix(subconscious): seed defaults and spawn heartbeat on startup The subconscious engine was only constructed lazily on the first engine-routed RPC (trigger, tasks_add, status). Because handle_tasks_list bypasses the engine and reads the store directly, a fresh install showed an empty Subconscious panel until the user clicked "Run now", even though SubconsciousEngine::new() seeds the 3 default system tasks on construction. Separately, HeartbeatEngine::run() — the periodic tick loop — was never spawned in production code. The only callers of HeartbeatEngine were tests, so ticks never fired automatically; users had to trigger each evaluation manually. Both issues are fixed together in run_server_inner, following the existing start_if_enabled pattern used by voice, screen_intelligence, and autocomplete: 1. Call get_or_init_engine() at startup to construct the SubconsciousEngine eagerly, which runs seed_default_tasks via from_heartbeat_config. Construction is idempotent via OnceLock; seeding is idempotent by title match, so repeat startups do not duplicate the defaults. 2. Construct HeartbeatEngine with the heartbeat config and workspace_dir, then tokio::spawn heartbeat.run() so the periodic tick loop runs for the process lifetime. The loop re-acquires the shared engine via get_or_init_engine() on each tick. Guarded by config.heartbeat.enabled so users who disable the heartbeat get neither startup seeding nor the background loop. Add engine_construction_seeds_default_tasks integration test that locks in the invariant: constructing SubconsciousEngine on a fresh workspace_dir must leave the 3 default system tasks in the store, with no tick, trigger, or explicit seed call. Also asserts that reconstructing the engine on the same workspace does not duplicate the defaults. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(subconscious): defer engine bootstrap until after login Default system tasks seeded at sidecar startup into the pre-login global workspace (`~/.openhuman/workspace/`) instead of the per-user workspace (`~/.openhuman/users/<id>/workspace/`) the UI reads from after login. The engine singleton is built lazily via `get_or_init_engine()` and cached in a `OnceLock`. `Config::load_or_init` resolves `workspace_dir` from `active_user.toml` — which does not exist until after login. When the engine was constructed on startup it therefore seeded into the global default, then the frozen singleton kept pointing at that path for the rest of the session while RPC handlers like `tasks_list` re-loaded config per call and read from the correct per-user path, silently returning an empty list. Fix: - `subconscious/global.rs`: add `bootstrap_after_login()` (idempotent via `BOOTSTRAPPED: AtomicBool`) which builds the engine against the now-correct per-user workspace and spawns the heartbeat loop. Track the heartbeat `JoinHandle` in a static so it can be aborted cleanly. Add `reset_engine_for_user_switch()` that aborts the heartbeat, clears the engine option, and resets the bootstrap flag. - `core/jsonrpc.rs`: replace the unconditional eager init on startup with a conditional one that only bootstraps if `active_user.toml` already exists (so a user logged in from a previous session still gets the engine up immediately after restart). - `credentials/ops.rs`: call `bootstrap_after_login()` at the end of `verify_and_store_session` so a fresh login triggers seeding against the per-user workspace. Call `reset_engine_for_user_switch()` in `clear_session` so logout tears down the engine + heartbeat loop and a subsequent login rebuilds them against the new user. Verified locally: sidecar restart with no `active_user.toml` logs "bootstrap deferred — waiting for login"; post-login logs "seeded 3 tasks on init" + "heartbeat periodic loop spawned"; and `subconscious.tasks_list` returns the 3 system defaults from the per-user DB. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(subconscious): bound config load + guard frontend poll Two related fixes for the Intelligence page freezing on a stale subconscious activity-log snapshot while ticks kept progressing in the sidecar. Root cause (backend): the subconscious RPC handlers were the only outlier in the entire JSON-RPC surface that called the raw `Config::load_or_init()` instead of the shared `load_config_with_timeout()` wrapper that every other domain schemas.rs uses (cron, webhooks, voice, team, skills, service, referral, doctor, …). `load_or_init` constructs a fresh `SecretStore` and runs a chain of `decrypt_optional_secret` calls on every invocation, which may IPC to the OS keychain — slow, unbounded, no caching. Under the Intelligence page's 3-second poll (4 parallel RPCs × ~7 keychain round-trips each = ~28 keychain calls every 3s), this pileup was enough to pin the frontend's `Promise.all` past the poll interval. Root cause (frontend): `useSubconscious.refresh()` uses `fetchingRef` as an in-flight guard. The ref is only cleared inside the `finally` block that runs after `Promise.all` settles. With no per-RPC timeout on the client side either, a single slow backend call would leave the ref stuck `true`, and every subsequent 3s `setInterval` tick would silently early-return at the top of `refresh`. The poller kept firing, but every call was a no-op — so the UI froze on whatever snapshot it last successfully fetched, even though the backend was still ticking through new decisions. Backend fix (`src/openhuman/subconscious/schemas.rs`): - Replace the local `load_config()` helper body to delegate to `crate::openhuman::config::load_config_with_timeout()`. Matches the 28 other domain schemas.rs files and brings subconscious handlers under the same 30s bound used everywhere else. Frontend fix (`app/src/hooks/useSubconscious.ts`): - Add a `withTimeout` helper (2.5s per-RPC, strictly less than the 3s poll interval) that races each of the 4 parallel RPCs against a timeout and resolves `null` on timeout — matching the existing `.catch(() => null)` contract so downstream setState logic is unchanged. - Clear `fetchingRef.current = false` in the useEffect cleanup so a late-returning request or a React Strict Mode double-mount in dev can't leave the ref stuck `true` for the next mount. Defense in depth: the backend bound prevents a permanent hang and matches repo conventions, while the frontend bound guarantees the 3s poll loop can never be pinned beyond one tick regardless of server-side latency. Verified locally — `cargo check` clean, `tsc --noEmit` clean, all 18 pre-existing warnings in unrelated modules. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style(jsonrpc): cargo fmt the startup bootstrap block CI ran `cargo fmt --all -- --check` and flagged the conditional bootstrap block in `run_server_inner` — `let already_logged_in` should fold onto one line, the `.and_then` closure body should inline, the `match ... .await` chain should fold, and the short log!() calls should not break across lines. No behavior change. Fixes three jobs on PR #462 that were all failing at the same `cargo fmt --all -- --check` step (Rust Quality, Rust Tests, Type Check TypeScript — the last one chains cargo fmt after its prettier check). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
fa5f822f95 |
feat(subconscious): stabilize heartbeat + subconscious loop (#392) (#437)
* feat(subconscious): stabilize heartbeat + subconscious loop (#392) - Enable heartbeat by default (enabled=true, inference_enabled=true, 5min interval) - Seed system tasks on engine init, not first tick - SQLite-backed task/log/escalation persistence - Overlap guard with generation counter — stale ticks are cancelled - Single log entry per task per tick, updated in place (in_progress → act/noop/escalate/failed/cancelled) - Rate-limit retry (429 only) for agentic-v1 cloud model calls - Approval gate: unsolicited write actions on read-only tasks require user approval - Analysis-only mode for agentic-v1 on read-only escalations - Non-blocking status RPC — reads from DB, never blocks on engine mutex - Frontend: system vs user task distinction, toggle switches, expandable activity log - Frontend: 3s auto-poll on Subconscious tab, skill-related escalation navigation - Consecutive failure counter in status (resets on success) - last_tick_at only advances on successful evaluation - Missing LLM evaluation fallback — unevaluated tasks default to noop - Docs: subconscious.md architecture guide, memory-sync-functions.md reference Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: fix Prettier formatting for subconscious frontend files Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * ci: retrigger checks Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(heartbeat): use disabled config in run_returns_immediately_when_disabled test HeartbeatConfig::default() has enabled: true, so run() entered the infinite loop and never returned — hanging the test (and CI) indefinitely. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(subconscious): remove HEARTBEAT.md task import, use SQLite as sole task source Tasks are now managed exclusively in SQLite via the Subconscious UI. HEARTBEAT.md is retained for instructions/context only, not as a task list. Situation report now reads pending tasks from SQLite instead of HEARTBEAT.md. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: cargo fmt on subconscious engine Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
0f578382d1 |
fix(autocomplete): auto-start engine when config enabled (#412) (#442)
* fix(autocomplete): add start_if_enabled for engine auto-start at boot (#412) The autocomplete engine was never started automatically when config had autocomplete.enabled = true. Add start_if_enabled() that checks config and starts the global engine singleton during core process startup. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(autocomplete): wire engine startup into core server init (#412) Call start_if_enabled() after config load in the JSON-RPC server so the autocomplete engine runs automatically when the core process boots with autocomplete enabled. Remove stale E2E test assertions that conflicted with the new startup path. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(autocomplete): auto-start engine when enabled via set_style RPC (#412) When the frontend enables autocomplete through set_style(enabled=true), automatically start the engine so suggestions begin immediately without requiring a separate start call or app restart. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
371bcd34ea |
Feat/chat issue (#441)
* feat: display app version in settings panel * fix(onboarding): auto-refresh accessibility state after grant (#351) * style(onboarding): apply formatter for issue #351 fix * fix(onboarding): ESLint + typed mock for ScreenPermissionsStep; clear flag in handler Consolidate tauriCommands imports and drop redundant mock cast. Handle granted accessibility in focus/visibility callback instead of a follow-up effect. Made-with: Cursor * feat(env): add support for custom dotenv path and update dependencies - Introduced an optional environment variable `OPENHUMAN_DOTENV_PATH` to specify a custom path for dotenv files, enhancing configuration flexibility. - Updated `Cargo.toml` to include the `dotenvy` dependency for improved dotenv file handling. - Enhanced the `.env.example` file with a new comment for the custom dotenv path. - Added data-testid attributes and button types in `SkillDebugModal` and `Skills` components for better testability. - Created new tests for Gmail and Notion third-party skills to ensure proper functionality of sync and debug tools. - Added documentation for memory sync functions to clarify usage patterns and function details. * fix: address CodeRabbit review on PR #441 - Dotenv: treat empty OPENHUMAN_DOTENV_PATH as unset; propagate from_path errors - Document OPENHUMAN_DOTENV_PATH parent-env requirement in .env.example - Memory docs: MD040 fence language; clarify skill namespace vs integration id - QuickJS bootstrap: modern helpers, generic platform.notify log, template URLs - Skills UI: type=button on close/settings; async waitFor in sync tests - Gmail OAuth e2e: workspace env matches MemoryClient; env/engine drop guards; redact secrets from logs - Add replace_global_engine for test teardown Made-with: Cursor |
||
|
|
90fac95d7a |
Fix: telegram replies (#436)
* feat(config): enhance world-readable config warning mechanism - Introduced a new static variable to track previously warned world-readable config files, preventing duplicate warnings. - Updated the warning logic to only log a warning for each unique world-readable config file, improving log clarity and reducing noise. - Added new `ChannelReactionReceived` and `ChannelReactionSent` events to the DomainEvent enum, expanding event handling capabilities in the event bus. - Included tests for the new reaction events to ensure proper functionality and integration. * feat(logging): add log file constraints and event filtering - Introduced functions to parse log file constraints from environment variables and filter log events based on these constraints. - Enhanced the `init_for_cli_run` function to apply the new filtering logic, improving log management and clarity. - Updated the `conversation_history_key` function to include thread context for Telegram, ensuring accurate message targeting. - Added a new trait method `supports_reactions` to the `Channel` trait, indicating support for emoji reactions. - Implemented integration tests for Telegram channel features, including reaction handling and thread message forwarding. * feat(telegram): enhance message handling with reactions and typing indicators - Added support for emoji reactions in Telegram responses, allowing for contextual acknowledgment of user messages. - Implemented a decision heuristic for when to use reactions, improving user interaction quality. - Introduced a typing indicator that activates immediately upon receiving a message, providing instant feedback to users. - Updated the channel delivery instructions to include new reaction syntax and guidelines for usage. - Enhanced tests to cover new reaction handling and message acknowledgment features, ensuring robust functionality. * fix(tests): update route key for Telegram message handling tests - Changed the route key in tests from `telegram_alice` to `telegram_alice_chat-1` to match the updated `conversation_history_key` format for Telegram. - This adjustment ensures accurate routing and consistency in message handling tests. * refactor(tests): streamline message handling in runtime tool calls - Refactored the message handling tests to utilize a `ChannelMessage` struct for improved clarity and maintainability. - Updated the route key generation to use the `conversation_history_key` function, ensuring consistency in message routing. - Simplified the invocation of `process_channel_message` by directly passing the constructed message, enhancing readability. * fix(telegram): enhance finalize_draft method to support thread context - Updated the `finalize_draft` method in the `Channel` trait and its implementation for `TelegramChannel` to accept an optional `thread_ts` parameter, allowing for message threading. - Adjusted related message handling functions to utilize the new parameter, ensuring proper message context during sending. - Modified tests to reflect changes in the `finalize_draft` method signature, enhancing the robustness of message handling in threaded conversations. * refactor(tests): ran format |
||
|
|
45a821645a |
feat(refer) : Implement referral system with UI components and API integration (#430)
* Implement referral system with UI components and API integration - Added to document the referral system, including reward structure, rules, data model, migration, core services, and API endpoints. - Created component to display referral stats and allow users to apply referral codes. - Integrated referral functionality into the onboarding process with for applying referral codes. - Updated page to include the new . - Implemented API calls in for fetching referral stats and applying referral codes. - Added tests for the referral API to ensure proper functionality and data normalization. - Introduced device fingerprinting for referral code application to prevent abuse. This commit establishes a comprehensive referral system, enhancing user engagement and incentivizing referrals. * Update OLLAMA_BASE_URL to local development address * Enhance referral system and onboarding process - Added a new function to format reward rates from basis points to percentage for better display in the ReferralRewardsSection. - Improved loading state management in the referral stats loading process to prevent race conditions. - Updated the Onboarding component to handle referral step skipping more effectively, ensuring a smoother user experience. - Fixed a typo in the WelcomeStep component's button label for clarity. - Enhanced error handling in the referral API to provide clearer feedback on failures. These changes improve the usability and reliability of the referral system and onboarding experience. |
||
|
|
4ee518cf31 |
feat(tree-summarizer): hierarchical summary tree module with CLI (#423)
* feat(tree-summarizer): implement hierarchical summarization engine and event handling - Introduced a new `tree_summarizer` module to manage hierarchical time-based summaries, organizing data into a tree structure (root → year → month → day → hour). - Added functionality to ingest raw content, summarize it into hour leaves, and propagate summaries upward through the tree. - Implemented event handling for summarization completion and tree rebuild events, enhancing observability and modularity. - Created RPC operations for ingesting content, triggering summarization, querying the tree, and retrieving tree status. - Added comprehensive tests to ensure the reliability of the summarization process and event handling. This update significantly enhances the summarization capabilities of the system, allowing for efficient data organization and retrieval. * feat(tree-summarizer): add CLI support for tree summarization commands - Introduced a new `tree-summarizer` command to the CLI, allowing users to ingest content, run summarization jobs, query the summary tree, check status, and rebuild the tree. - Updated the CLI help documentation to include the new command and its subcommands. - Added a new module `tree_summarizer_cli` to encapsulate the tree summarization functionality. This enhancement improves the usability of the summarization features, providing a streamlined interface for managing hierarchical summaries directly from the command line. * style: apply cargo fmt to tree_summarizer module * feat(tree-summarizer): implement TreeSummarizerEventSubscriber for observability logging - Added a new `TreeSummarizerEventSubscriber` to log events related to tree summarization, enhancing observability. - Updated the `start_channels` function to register the new subscriber. - Refactored the `run_summarization` function to group buffered entries by hour and publish events upon completion of summarization. - Improved documentation and added tests for the new subscriber functionality. This update aims to provide better insights into the summarization process and facilitate future cross-module workflows. * refactor(tree_summarizer): streamline buffer backup and function signature - Simplified the buffer backup process by consolidating the rename operation with context handling for better error reporting. - Cleaned up the function signature of `derive_node_ids_from_hour_id` for improved readability. These changes enhance code clarity and maintainability within the tree summarization engine. * feat(tree_summarizer): enhance tree summarization with metadata support - Updated the `tree_summarizer_ingest` function to accept an optional metadata parameter, allowing users to include additional context during content ingestion. - Refactored related functions to validate and handle metadata, improving the overall robustness of the summarization process. - Adjusted the buffer write functionality to store metadata alongside content, enhancing the data structure for future retrieval and processing. These changes aim to enrich the summarization capabilities and provide more context for ingested content. * refactor(tree_summarizer): improve error message formatting in node ID validation - Enhanced the formatting of error messages in the `validate_node_id` function for better readability and consistency. - Adjusted the string formatting to use multi-line syntax, improving clarity in error reporting. - Minor formatting changes in the `strip_buffer_frontmatter` function to enhance code readability. These changes aim to improve the maintainability and clarity of error handling within the tree summarization module. * refactor(tree_summarizer): enhance buffer management and summarization process - Replaced the buffer draining mechanism with a non-destructive read approach, allowing for safer data handling during summarization. - Introduced a new `buffer_delete` function to explicitly manage the deletion of buffer entries after successful processing. - Updated the `run_summarization` function to reflect these changes, ensuring that buffer entries are only deleted after durable writes are confirmed. - Improved the backup process for the buffer directory during tree rebuilds, ensuring it is preserved outside the tree structure. These modifications aim to improve data integrity and clarity in the summarization workflow. * refactor(tree_summarizer): improve markdown parsing and timestamp handling - Updated the `parse_node_markdown` function to trim trailing whitespace from the body after splitting frontmatter, enhancing data cleanliness. - Modified test cases to use specific timestamps instead of the current time, ensuring consistent and predictable test results. - Adjusted assertions in tests to reflect the new timestamp-based ordering of entries. These changes aim to improve the robustness of markdown parsing and the reliability of test outcomes in the tree summarization module. |
||
|
|
0609493e1a | Add RAM-tiered local AI presets (#425) | ||
|
|
cd2a4a9a87 |
feat(screen-intelligence): OCR-only mode without vision model (#424)
* feat(settings): add 'Use Vision Model' option to Screen Intelligence Panel - Introduced a new checkbox in the Screen Intelligence Panel to toggle the use of a vision model for richer context extraction from screenshots. - Updated state management to handle the new option and integrated it into the configuration and processing logic. - Adjusted related tests and configurations to support the new feature, ensuring compatibility across the application. * feat(cli): add --no-vision-model option for screen intelligence - Introduced a new command-line option `--no-vision-model` to allow users to skip the vision model and use OCR and text LLM only. - Updated the CLI options parsing to handle the new flag and modified the bootstrap logic to respect this setting. - Enhanced usage documentation to reflect the new option and its alias `--ocr-only` for clarity. * fix(screen-intelligence): read use_vision_model from engine runtime config The processing worker was reading use_vision_model from the persisted config file (Config::load_or_init), so the CLI --no-vision-model flag had no effect. Now reads from the engine's in-memory runtime config which the CLI correctly overrides via apply_config(). Also moves image compression before OCR pass. * fix: add use_vision_model to test fixtures and fix rustfmt Add the new use_vision_model field to all AccessibilityConfig test fixtures so TypeScript compilation passes. Also includes rustfmt auto-fix for screen_intelligence_cli.rs. |
||
|
|
94f7531d69 |
Make QuickJS runtime fully synchronous to fix async deadlock (#422)
* feat(sync): implement background sync handling in event loop - Introduced a mechanism to track background sync status with a new `sync_in_flight` flag. - Updated the event loop to check for sync completion and persist state to memory upon completion. - Modified the `handle_sync` function to fire the `onSync` handler in the JS runtime asynchronously, allowing for immediate return while the sync process runs in the background. - Enhanced logging to provide feedback on sync status and completion. * fix(net): disable connection pooling in HTTP client to prevent hanging on POST requests through staging proxy * Enhance promise handling in JS call processing - Added logging to track promise resolution progress and timeout events. - Implemented a polling mechanism to drive the QuickJS job queue until promises resolve or timeout occurs. - Introduced debug information logging for stalled promises to aid in troubleshooting. - Improved feedback during polling to indicate ongoing operations. This update aims to improve the reliability and debuggability of asynchronous JavaScript calls within the application. * Refactor JS fetch implementation for synchronous behavior - Updated the `fetch` function in the QuickJS library to operate synchronously from the JavaScript perspective, blocking the thread until the HTTP request completes. - Enhanced error handling and logging for HTTP requests, including timeouts and response details. - Removed the previous asynchronous implementation to streamline the fetch operation. - Cleaned up the code by removing unused comments and improving readability. This change aims to improve the consistency and reliability of network operations within the application. * Fix Rust formatting in ops_net.rs Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
ca4eb39e9b | fix(oauth): enhance logging to include Notion-Version in fetch requests | ||
|
|
91996a1dd0 |
fix(voice): reduce dictation hallucinations and improve Fn/focus reliability (#385) (#409)
* fix(voice): add per-segment confidence validation in whisper engine (#385) Reject whisper segments with avg token log-probability below -0.7 or entropy above 2.4. Return TranscriptionResult with confidence metadata instead of plain String. Update callers in speech.rs and streaming.rs. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(voice): upgrade default STT model from tiny to base (#385) Base model produces significantly fewer hallucinations than tiny, especially in noisy/quiet conditions. User can still override via config. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(voice): add real-time silence gating in audio capture (#385) Gate sustained silence (>500ms) from being sent to whisper to prevent hallucinations. Maintain 100ms look-ahead ring buffer so speech onset after pauses is not clipped. Thresholds adapt to source sample rate. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(voice): fix Fn key timing race condition in hotkey event loop (#385) start_recording() blocks 1-7s on cpal device init but macOS fires Fn Release almost immediately, causing skipped cycles. Move recording start to spawn_blocking so the event loop stays responsive. Buffer Release events during setup and ensure minimum 1.5s recording duration when release arrives before recording handle is ready. Also includes: capture focused app on hotkey press, pass through pipeline for focus validation before paste. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(voice): validate and restore focus before paste, attempt regardless (#385) Add expected_app parameter to insert_text(). Before Cmd+V, validate focus via accessibility API and restore via AppleScript if shifted. Don't abort paste on focus validation failure — attempt insertion regardless so text is never silently lost. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(voice-ui): align voice server RPC response shape in settings panel (#385) * style(voice): apply rustfmt formatting in text_input * fix(voice): address CodeRabbit regressions in server, streaming, and settings polling --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
fb8987bcad |
Improve inline autocomplete reliability, sanitization, and debug logging (#407)
* Enhance autocomplete functionality and logging. Increased debounce time for autocomplete suggestions and added minimum context character requirement. Improved inline suggestion handling with new cleanup logic for tab acceptance. Introduced a new logging option for autocomplete-only logs in CLI. Updated various components to support these changes, including sanitization and error handling in the autocomplete engine. * Add autocomplete CLI adapter for improved argument handling This commit introduces a new module, , which encapsulates the argument parsing and logging logic specific to the autocomplete namespace in the CLI. Key features include extraction of leading verbose flags, handling of the flag, and improved help message printing. The existing CLI command handling has been refactored to utilize this new adapter, enhancing code organization and maintainability. * Refactor inline completion sanitization and enhance context handling |
||
|
|
000b40bf43 |
fix(local_ai): Windows Ollama discovery + DirectML GPU acceleration for GLiNER RelEx (#416)
Ollama was not found on Windows because find_system_ollama_binary lacked common Windows install paths (%LOCALAPPDATA%\Programs\Ollama). The server spawn also silently swallowed errors, and the NSIS installer fallback didn't check system paths after install. GLiNER RelEx ONNX sessions were CPU-only — no execution providers were configured. Now offers DirectML (Windows), CoreML (macOS), and CUDA as GPU backends with automatic fallback. Updated the release to v0.5-onnx.2 with a DirectML-enabled onnxruntime.dll. Bundle completeness now requires the platform DLL and verifies checksums to trigger re-download on update. Changes: - Add Windows common paths to find_system_ollama_binary (install.rs) - Log and return spawn errors in start_and_wait_for_server (ollama_admin.rs) - Fall back to find_system_ollama_binary after Windows installer (ollama_admin.rs) - Add platform_execution_providers() with DirectML/CoreML/CUDA (relex.rs) - Require ORT DLL in bundle_complete check (relex.rs) - Verify platform DLL checksums in managed_bundle_complete (relex.rs) - Update release URL and SHA256 hashes for v0.5-onnx.2 (relex.rs) Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
4afa751024 |
feat(memory): global singleton, CLI, graph extraction fixes & light storage (#383)
* feat(memory): add CLI support for memory commands - Introduced a new `memory` subcommand in the CLI for memory ingestion, graph inspection, and debugging. - Implemented various subcommands including `ingest`, `docs`, `graph`, `query`, and `namespaces` for comprehensive memory management. - Updated the CLI entry point to route the `memory` command appropriately, enhancing the command-line interface functionality. * refactor(memory_cli): streamline memory command ingestion and improve error handling - Simplified the ingestion process by removing unnecessary workspace directory creation and embedding logic. - Updated the ingestion function to utilize the `create_memory_client` for better client management. - Changed the limit parameter type from `usize` to `u32` for consistency and improved error handling in command arguments. - Enhanced logging for ingestion start to focus on model name only, removing redundant extraction mode information. * refactor(memory): implement global memory client singleton for improved resource management - Introduced a new `global.rs` module to manage a process-global memory client singleton, ensuring consistent access across subsystems. - Updated `create_memory_client` to utilize the global client, enhancing memory management and reducing resource contention. - Refactored various modules to replace local memory client instances with the global singleton, improving performance and reliability. - Adjusted CLI and screen intelligence components to leverage the global memory client for document persistence and ingestion operations. This refactor enhances the architecture by centralizing memory client management, leading to better resource utilization and simplified code structure. * feat(memory): add put_doc_light for screen-intelligence, skip vectors/graph Screen-intelligence captures are too frequent and ephemeral to justify vector embedding and GLiNER graph extraction per frame. Adds a lightweight storage path (put_doc_light) that persists the document row and markdown file without chunking, embedding, or graph extraction. Three-tier storage: - put_doc_light: DB + markdown only (screen-intelligence) - put_doc: DB + markdown + vectors + background graph (skill sync) - ingest_doc: full synchronous pipeline (CLI, debugging) * style: apply cargo fmt formatting * fix: add missing window_id field in AppContext test helper * fix(test): fall back to per-call MemoryClient when global not initialized In tests with isolated OPENHUMAN_WORKSPACE, the process-global singleton may not be initialized or may point at the wrong directory. Fall back to creating a client from Config (which respects env vars) when the global is not ready. * fix(test): use per-call MemoryClient for screen-intelligence persistence put_doc_light does no background work (no vectors, no graph), so a per-call client created from Config is safe and avoids the global singleton which may point at a different workspace in test suites. |
||
|
|
a98817917c |
feat(screen-intelligence): standalone server with OCR + vision pipeline (#382)
* feat(screen-intelligence): add new commands for diagnostics and vision processing - Introduced `doctor` command for system readiness diagnostics, checking permissions and platform support. - Added `vision` command to analyze frames and persist vision summaries. - Updated CLI usage documentation to reflect new command options and improved verbosity handling. - Enhanced the `run_server` function to provide detailed endpoint information for better user guidance. This update improves the functionality and usability of the screen intelligence CLI, enabling better diagnostics and vision processing capabilities. * feat(screen-intelligence): update CLI endpoints for status monitoring - Added `tower-http` dependency for enhanced HTTP capabilities. - Updated endpoint documentation in `run_server` to reflect changes from SSE to long-polling for status updates. - Renamed `/events` endpoint to `/watch` with a query parameter for interval control, improving clarity and usability. This update enhances the screen intelligence CLI by providing more flexible status monitoring options. * feat(screen-intelligence): integrate standalone server for screen intelligence - Added a new `server` module to handle the standalone screen intelligence server functionality. - Updated CLI commands to include an `--auto-start` option for initiating capture sessions on server boot. - Enhanced the `run_server` function to provide detailed endpoint information and improved logging for server status. - Integrated the screen intelligence engine with JSON-RPC and REST endpoints for better debugging and usability. This update significantly enhances the screen intelligence capabilities by allowing it to run independently and providing more flexible configuration options. * refactor(screen-intelligence): simplify CLI options and server configuration - Removed the `--port` and `--auto-start` options from the CLI for the `screen-intelligence run` command, streamlining the command usage. - Updated the server configuration to focus on session duration and logging, enhancing clarity and usability. - Adjusted the documentation to reflect the new command structure and improved logging details for the screen intelligence server. This refactor improves the user experience by simplifying command options and enhancing the clarity of server operations. * feat(screen-intelligence): enhance screenshot management and CLI options - Updated the CLI usage documentation to include the new `--keep` option for retaining screenshots after processing. - Modified the server configuration to support the `keep_screenshots` flag, allowing for immediate saving of screenshots to disk. - Enhanced the `run_server` function to reflect the updated configuration and logging details regarding screenshot retention. - Improved the capture logic to prioritize window ID for more reliable screenshot capturing on macOS. These changes improve the usability and functionality of the screen intelligence feature by providing better control over screenshot management. * refactor(accessibility): improve capture logic and window ID handling - Updated the capture mode logic to prioritize reliable window ID capture on macOS, falling back to fullscreen capture to avoid issues with region-based capture. - Refactored the method for resolving the frontmost window ID to utilize Swift for better performance and reliability, replacing the previous JavaScript-based approach. - Enhanced the frame processing in the accessibility engine to track processed timestamps, preventing re-analysis of the same screenshot and improving efficiency. - Adjusted the vision summary parsing to support both JSON and plain text formats, ensuring better compatibility and usability. These changes enhance the reliability and performance of the accessibility features, particularly in macOS environments. * feat(accessibility): integrate Apple Vision OCR and enhance LLM context analysis - Implemented Apple Vision OCR for extracting text from screenshots, improving accuracy and reducing hallucination. - Refactored the vision processing flow to include structured prompts for LLM context analysis, enhancing the quality of the output summary. - Updated the vision summary structure to include detailed fields such as APP, DOING, FOCUS, and MOOD, providing clearer insights into the captured content. - Improved error handling and logging for image processing and OCR operations, ensuring better reliability and debugging capabilities. These changes significantly enhance the accessibility engine's ability to analyze and summarize screen content effectively. * refactor(screen-intelligence): enhance capture logic and summary structure - Improved capture mode handling to prioritize window ID availability, with a fallback to bounds-based capture before defaulting to fullscreen. - Updated the vision summary structure to clearly separate synthesized summaries, visual context, and raw OCR text for better clarity and usability. - Enhanced logging to provide more informative output during capture operations, aiding in debugging and performance monitoring. These changes improve the reliability and clarity of the screen intelligence features, particularly in managing capture contexts and summarizing visual information. * refactor(accessibility): enhance capture logic to reject fullscreen fallback - Updated the screen capture logic to prevent falling back to fullscreen mode when no valid window ID or bounds are available, improving privacy and user intent. - Added detailed logging for cases where fullscreen capture is refused, aiding in debugging and understanding capture decisions. - Introduced tests to ensure that the new logic correctly rejects fullscreen capture under invalid conditions, enhancing reliability. These changes improve the accessibility engine's handling of screen captures, ensuring that only relevant content is captured. * feat(screen-intelligence): implement capture and processing workers for enhanced vision analysis - Introduced a dedicated `capture_worker` to manage screenshot capturing from the foreground context, ensuring efficient frame handling and immediate saving of screenshots when configured. - Added a `processing_worker` to analyze captured frames using Apple Vision OCR and LLM, synthesizing insights and persisting results to unified memory. - Refactored the `AccessibilityEngine` to integrate these workers, improving the overall architecture and separation of concerns in the screen intelligence module. - Enhanced logging throughout the capture and processing workflows for better debugging and performance monitoring. These changes significantly improve the screen intelligence capabilities by enabling real-time capture and analysis of visual data, enhancing user experience and functionality. * Merge remote-tracking branch 'upstream/main' into fix/screen-intgellignce-2 * refactor(screen-intelligence): improve type handling and visibility in capture and processing modules - Updated the calculation of baseline milliseconds in `capture_worker` to ensure proper type handling with floating-point division. - Made several fields in `SessionRuntime` and `EngineState` public for better accessibility within the crate. - Changed the visibility of the `analyze_frame` function in `processing_worker` to public within the crate, allowing it to be called from `engine.rs`. These changes enhance type safety and improve the modularity of the screen intelligence components, facilitating better integration and usage across the codebase. * refactor(screen-intelligence): update visibility and helper function usage in engine and processing modules - Changed the `SessionRuntime` struct to be public within the crate, enhancing accessibility. - Updated the `analyze_frame` function parameter to use a reference instead of a direct reference to the `AccessibilityEngine`, improving clarity. - Refactored calls to `persist_vision_summary` to utilize the helper function from the `super::helpers` module, promoting better organization and code reuse. These changes streamline the code structure and improve the modularity of the screen intelligence components. * feat(screen-intelligence): introduce input handling and autocomplete features - Added a new `input.rs` module to manage input actions, including keyboard/mouse automation and predictive text functionalities. - Implemented methods for handling input actions, generating autocomplete suggestions, and committing selected suggestions within the `AccessibilityEngine`. - Created a `state.rs` module to encapsulate engine state management, including session runtime and engine state structures. - Enhanced the `engine.rs` module to support session lifecycle management and integrate new input functionalities. - Introduced a `vision.rs` module for vision-related query methods, improving the overall architecture and modularity of the screen intelligence components. These changes significantly enhance user interaction capabilities and streamline the management of input actions and vision processing. * fix(screen-intelligence): update worker imports to use state module Workers imported AccessibilityEngine from engine.rs but it was moved to state.rs during the refactor. Fix the import paths. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(screen-intelligence): clean up formatting and improve readability in various modules - Reformatted print statements in `screen_intelligence_cli.rs` for better readability. - Simplified assertions in tests within `capture.rs` to enhance clarity. - Streamlined function definitions and calls in `focus.rs`, `engine.rs`, and `processing_worker.rs` for improved code organization. - Updated import statements in `mod.rs` and `server.rs` to maintain consistency and clarity. These changes enhance the overall readability and maintainability of the codebase, promoting better coding practices across the screen intelligence components. * refactor(screen-intelligence): update import paths for AccessibilityEngine and EngineState - Changed the import statements in `tests.rs` to source `AccessibilityEngine` and `EngineState` from the `state` module instead of the `engine` module. - This adjustment aligns with recent refactoring efforts to improve module organization and maintainability. These changes enhance code clarity and ensure consistency in module usage across the screen intelligence components. * fix(screen-intelligence): fix test compilation and assertions - Update test imports from engine to state module - Add mock env var path to processing_worker::analyze_frame for test support - Update parser tests for plain-text mode (non-JSON fallback) - Handle missing window_id gracefully in capture_scheduler test Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(screen-intelligence): enhance code clarity and performance in various modules - Updated command-line usage documentation in `screen_intelligence_cli.rs` to reflect new options. - Improved string handling for truncation in `run_start_session` and `run_vision` functions to use character counts instead of byte lengths, ensuring accurate truncation for multi-byte characters. - Added intentional blocking sleep in `resolve_frontmost_window_id` to minimize impact on the Tokio runtime during app switches. - Implemented a consistent default confidence score in vision processing and improved YAML escaping in `persist_vision_summary`. - Refactored session management in `capture_worker` and `engine` modules to reduce lock contention and improve performance during I/O operations. These changes enhance the overall performance, maintainability, and user experience of the screen intelligence components. * fix(tests): update confidence assertion in vision summary test - Adjusted the expected confidence value in the `parse_vision_missing_fields` test to reflect the new default confidence score of 0.8, ensuring consistency across JSON and plain-text branches. This change improves the accuracy of the test and aligns it with recent updates in the vision processing logic. * refactor(screen-intelligence): improve code formatting and readability in various modules - Enhanced string formatting for truncation in `run_vision` to improve clarity. - Reformatted variable declarations in `capture_worker` for better readability. - Updated YAML escaping logic in `helpers.rs` for consistency and clarity. These changes contribute to improved maintainability and readability of the codebase. * refactor(screen-intelligence): reorganize analyze_frame function for improved flow - Moved configuration validation to the beginning of the `analyze_frame` function to ensure proper setup before processing. - Reintroduced the Apple Vision OCR and Vision LLM steps in the correct order, enhancing the logical flow of the analysis process. These changes improve the readability and maintainability of the code, ensuring a clearer structure for future modifications. --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
8627cee960 |
Feat/overlay (#378)
* feat(autocomplete): add overlay TTL configuration to AutocompletePanel - Introduced `overlay_ttl_ms` parameter to the Autocomplete configuration, allowing users to set the overlay display duration. - Updated AutocompletePanel to include a new input field for adjusting the overlay TTL in milliseconds. - Enhanced parsing and saving logic to handle the new configuration parameter. - Added corresponding tests to ensure functionality and validate the new overlay TTL feature. This update improves user control over the autocomplete overlay behavior, enhancing the overall user experience. * Refactor accessibility code for improved readability and consistency - Simplified log statements in `precompile_helper_background` for better clarity. - Reformatted `detect_input_monitoring_permission` check in `keys.rs` for enhanced readability. - Rearranged imports in `mod.rs` to maintain consistent structure. - Improved formatting of `ElementBounds` initialization across multiple test cases in `overlay.rs` and `types.rs` for better visual alignment. - Enhanced test context creation in `types.rs` for improved clarity. These changes enhance code maintainability and readability across the accessibility module. * fix(overlay): parent core RPC, voice toggle, and debug for #342 - Pass OPENHUMAN_OVERLAY_PARENT_RPC_URL from sidecar spawn and strip inherited OPENHUMAN_CORE_PORT so the overlay no longer fights for the parent listen port. - Overlay UI uses HTTP JSON-RPC to the parent sidecar for globe, debug, and voice STT so state matches the main app; add parentCoreRpc helper mirroring legacy method aliases. - Skip embedded JSON-RPC server when parent URL is set; use OPENHUMAN_OVERLAY_EMBEDDED_CORE_PORT (default 7799) for standalone dev. - Fix screen intelligence status method name; add voice_status polling, STT section, collapsible debug summary, and connection banner when core is unreachable. - Voice capture switch gates the mic with real STT availability from voice_status. Closes #342 Made-with: Cursor * fix: address CodeRabbit review on overlay/autocomplete (PR #378) - helper: correlate JSON-RPC replies with monotonic request ids; discard mismatched lines until deadline (fixes stale response after timeout). - helper: serialize Swift compile via HELPER_COMPILE_LOCK (precompile vs first use). - Overlay Tab hint: pass tab_hint from Rust from accept_with_tab; Swift hides hint when empty. - keys: re-check Input Monitoring on an interval when denied so grant without restart works. - engine: use non-zero confidence placeholder (0.75) until inline_complete returns scores. - overlay dedupe: suppress identical badge only within 400ms, not for process lifetime. Co-authored-by: Code review feedback <noreply@github.com> Made-with: Cursor * fix(ci): resolve clippy and warning issues for Rust gates - Use match on anchor_bounds in autocomplete overlay (avoid unnecessary unwrap) - Drop unused test imports in registry_ops and rpc dispatch - Prefix unused notion_doc_id in subconscious integration test Made-with: Cursor * fix(overlay): improve parent RPC URL handling in App component - Updated the useEffect hook to manage the parent RPC URL more robustly by introducing a mounted flag to prevent state updates on unmounted components. - Added error handling to set the parent RPC URL to null in case of invocation failure, enhancing the reliability of the component's behavior. * feat(overlay): implement timeout handling for parent core RPC requests - Introduced a default timeout for parent core RPC requests, enhancing reliability by preventing indefinite waiting for responses. - Added an AbortController to manage request timeouts, throwing a specific error message when a timeout occurs. - Updated the `callParentCoreRpc` function to accept a customizable timeout parameter, improving flexibility for RPC calls. * fix(overlay): allow stopping active recording regardless of config state - Updated the main button handler in the App component to always permit stopping an active recording when the status is "listening", improving user experience and control over the recording process. - Removed redundant code that previously checked the status before stopping the recording, streamlining the logic. * feat(overlay): update Cargo.lock with new dependencies and versions - Added new packages including `alsa`, `alsa-sys`, `arboard`, `block`, `cocoa`, `core-foundation`, `core-graphics`, `coreaudio-rs`, `coreaudio-sys`, `cpal`, `crunchy`, and `dasp_sample` to enhance functionality and support for audio processing and system interactions. - Updated existing dependencies to their latest versions for improved performance and compatibility. - Modified the `show_overlay` function in `ops.rs` to include an additional parameter, enhancing the overlay display functionality. * feat(autocomplete): add overlay_ttl_ms parameter to Autocomplete interfaces - Introduced a new optional parameter `overlay_ttl_ms` to both `AutocompleteSetStyleParams` and `AutocompleteConfig` interfaces, allowing for customizable overlay timeout settings. - This enhancement improves the flexibility of the autocomplete feature by enabling developers to specify how long the overlay should remain visible. --------- Co-authored-by: Code review feedback <noreply@github.com> Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai> |
||
|
|
e8fd08c800 |
feat(memory): use timestamp-prefixed document IDs for chronological sorting (#381)
Replace random UUID document IDs with `{unix_seconds}_{8char_hex}` format
so memory documents sort chronologically by filename and ID.
|
||
|
|
3851d1ef67 |
fix(voice): anti-hallucination, clipboard paste, Fn key reliability (#380)
* fix(dictation): update hotkey default value and documentation - Changed the default global hotkey for dictation from "CmdOrCtrl+Shift+D" to "Fn" in both the configuration schema and the associated documentation. - Updated the hotkey parsing function to recognize "Fn" as a valid key, enhancing the flexibility of hotkey configurations. - Added a test case to ensure the "Fn" key can be parsed correctly, improving the robustness of the hotkey handling functionality. * fix(voice): update default activation mode and hotkey in configuration - Changed the default activation mode for voice and dictation from "tap" to "push" in the respective configuration schemas. - Updated the default hotkey for voice commands from "ctrl+shift+space" to "Fn" across various modules and documentation. - Adjusted related tests to reflect the new defaults, ensuring consistency in behavior and expectations. * feat(voice): integrate embedded global voice server startup - Added a new asynchronous function `start_if_enabled` to the voice server module, which initializes the embedded voice server based on configuration settings. - Updated the server run logic to check if the voice server should auto-start, enhancing the startup process for the core application. - Integrated the new server startup function into the main server run logic, ensuring the voice server is launched if enabled in the configuration. * feat(voice): add VoicePanel for managing voice server settings - Introduced a new `VoicePanel` component to handle voice server configurations, including startup options, hotkeys, and runtime controls. - Updated routing in the settings page to include the new voice settings section. - Enhanced the `useSettingsNavigation` hook to support navigation to the voice settings. - Added tests for the `VoicePanel` to ensure functionality and reliability of the voice server management features. * refactor(dictation): update documentation and improve component initialization - Revised comments in `DictationHotkeyManager` to clarify the component's mounting process within the app tree. - Removed unused imports and unnecessary state management from `ServiceBlockingGate`, streamlining the component's logic. - Updated tests for `ServiceBlockingGate` to reflect changes in behavior, ensuring accurate rendering of child components based on service status. - Enhanced the `Cargo.lock` file by updating dependencies to their latest versions for improved stability and security. * fix(voice): update default skip_cleanup setting and enhance VoicePanel options - Changed the default value of `skip_cleanup` in the voice server configuration from `false` to `true` to improve transcription handling. - Reordered options in the `VoicePanel` component to ensure "Natural cleanup" is displayed alongside "Verbatim transcription" for better user clarity. - Updated tests to reflect the new default settings and ensure proper functionality of the VoicePanel component. * feat(window): add window management commands for Tauri application - Introduced a new module `window.ts` containing functions for managing window visibility and state in a Tauri application. - Implemented commands to show, hide, toggle visibility, minimize, maximize, close, and set the title of the main window. - Added checks to ensure commands are only executed in a Tauri environment, enhancing compatibility with web contexts. * feat(tauriCommands): add comprehensive Tauri command modules - Introduced multiple new modules for Tauri commands, including `accessibility`, `autocomplete`, `config`, `conscious`, `core`, `cron`, `hardware`, `localAi`, and `window`. - Each module contains functions for managing specific functionalities such as accessibility permissions, autocomplete suggestions, configuration settings, and hardware interactions. - Implemented checks to ensure commands are executed only in a Tauri environment, enhancing compatibility and reliability. - This addition significantly expands the command capabilities of the Tauri application, providing a robust framework for future development. * feat(voice): enhance audio transcription with initial prompt support - Added functionality to transcribe audio with an optional initial prompt, allowing for vocabulary bias and improved conversational continuity. - Updated the `transcribe_pcm_f32` and `transcribe_wav_file` functions to accept an `initial_prompt` parameter, enhancing recognition of specific vocabulary. - Implemented peak RMS energy tracking during audio recording for silence detection, ensuring recordings below a defined threshold are skipped. - Enhanced the voice server configuration to include a silence threshold and custom dictionary for better transcription context. - Introduced methods to build and manage recent transcripts for improved continuity across consecutive recordings. - Updated tests to validate new features and ensure proper functionality of the transcription process. * feat(voice): enhance VoiceServerConfig with silence detection and custom dictionary - Added `silence_threshold` to the `VoiceServerConfig` for improved silence detection, allowing recordings with low RMS energy to be skipped. - Introduced `custom_dictionary` to bias transcription towards specific vocabulary, enhancing recognition of names and technical terms. - Updated the `voice_transcribe` and `voice_transcribe_bytes` functions to utilize the new `initial_prompt` parameter for better context during transcription. - Adjusted the `transcribe_pcm_i16` function to accept additional parameters for improved flexibility in handling audio input. * feat(voice): add silence threshold and custom dictionary features to VoicePanel - Implemented a new input for setting the silence threshold, allowing recordings with low RMS energy to be skipped. - Added functionality for a custom dictionary, enabling users to add specific vocabulary words to improve transcription accuracy. - Updated the VoiceServerSettings interface and related functions to support the new features, ensuring seamless integration with existing settings. - Enhanced the UI in the VoicePanel to facilitate user interaction with the new settings. * feat(voice): propagate silence threshold and custom dictionary to voice server command - Added `silence_threshold` and `custom_dictionary` parameters to the `run_voice_server_command` function, ensuring these settings are utilized during voice server operations. - Enhanced integration with the existing voice server configuration to support improved transcription accuracy and silence detection. * fix(tauriCommands): update import paths for coreRpcClient - Adjusted import paths for `callCoreRpc` in multiple Tauri command modules to ensure correct referencing from the updated directory structure. - This change enhances module organization and maintains consistency across the codebase. * feat(dependencies): update Cargo.lock and Cargo.toml for new packages and versions - Added new dependencies including `arboard`, `fax`, `fax_derive`, `gethostname`, `half`, `quick-error`, `tiff`, and `x11rb` to enhance functionality and support for clipboard operations, transcription improvements, and system interactions. - Updated existing dependencies to their latest versions for better performance and compatibility. - Modified `VoicePanel` to utilize the updated settings and ensure proper handling of voice server configurations. - Enhanced the text input mechanism to use clipboard-paste for improved reliability in text insertion. * fix(voice): update skip_cleanup default value and enhance logging - Changed the default value of `skip_cleanup` in `VoiceServerConfig` from `true` to `false` to align with expected behavior and improve transcription handling. - Added detailed logging in the transcription cleanup process to provide better insights into the LLM state and cleanup decisions. - Removed unused functions related to unreliable key releases in hotkey handling to simplify the codebase. - Updated tests to reflect the new default settings for `skip_cleanup` and ensure proper functionality across components. * style: apply linter formatting fixes * fix(voice): remove unused warn import in hotkey module * test(voice): add silence threshold and custom dictionary to VoicePanel tests - Updated tests for the VoicePanel component to include new parameters: `silence_threshold` and `custom_dictionary`. - Ensured that the tests reflect the latest configuration settings for improved transcription accuracy and functionality. |
||
|
|
11f718d8bc |
feat(voice): standalone voice dictation server with hotkey support (#368)
* feat: add standalone voice dictation server with hotkey support - Introduced a new `voice` subcommand to the CLI for running a standalone voice dictation server that listens for a hotkey, records audio, transcribes it using Whisper, and inserts the result into the active text field. - Implemented configuration options for the voice server, including hotkey combination, activation mode (tap or push), and an option to skip LLM post-processing. - Added audio capture functionality using the `cpal` crate and integrated hotkey listening with the `rdev` crate for global key event handling. - Enhanced the configuration schema to include voice server settings and updated the main configuration structure accordingly. - Updated relevant modules and tests to ensure consistent behavior and functionality across the application. This feature enhances user interaction by allowing voice dictation directly into any active text field, improving accessibility and usability. * feat: add voice dictation server with hotkey support - Introduced a standalone voice dictation server that listens for a configurable hotkey to start recording audio, transcribes it using whisper, and inserts the transcribed text into the active text field. - Added CLI support for the `voice` command, allowing users to manage the voice server's configuration, including hotkey and activation mode settings. - Implemented configuration structures for the voice server, including options for automatic start, hotkey combination, activation mode, and cleanup behavior. - Enhanced audio capture functionality using the `cpal` library for microphone input and integrated text insertion using the `enigo` library for simulating keyboard input. - Updated relevant modules and schemas to support the new voice server features, ensuring a cohesive integration within the OpenHuman platform. * refactor: streamline voice server command and enhance audio capture functionality - Updated the `run_voice_server_command` function to initialize the configuration with environment overrides instead of loading from a file, improving performance and flexibility. - Refactored the audio capture logic in `start_recording` to enhance thread management and error handling, ensuring a more robust audio stream setup. - Improved the handling of audio stream creation and playback, ensuring that all cpal objects are managed on the same thread as required, enhancing stability during recording operations. * fix: remove unused import in voice server module - Eliminated the `HotkeyListenerHandle` import from the `server.rs` file, streamlining the code and improving clarity by removing unnecessary dependencies. * feat(voice): auto-enable LLM cleanup when local model is ready The postprocessor now checks the local LLM state and automatically enables transcription cleanup when the model is downloaded and ready, even if not explicitly configured. Falls back gracefully to raw text when the LLM is unavailable. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: apply cargo fmt + prettier formatting Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(voice): dictation config, hotkey lifecycle, and WebSocket streaming (#332) Add the foundational infrastructure for voice dictation (EPIC #332): **Rust core:** - New `DictationConfig` schema with serde defaults and env var overrides (enabled, hotkey, activation_mode, llm_refinement, streaming, interval) - RPC controllers: `config_get_dictation_settings` / `config_update_dictation_settings` - WebSocket endpoint `/ws/dictation` for streaming PCM16 transcription with periodic partial inference and final LLM refinement - Microphone permission declaration (`NSMicrophoneUsageDescription`) in Tauri macOS bundle config **Frontend:** - `useDictationHotkey` hook: fetches config from core RPC, auto-registers global hotkey, listens for `dictation://toggle` events - `DictationHotkeyManager` headless component mounted in App.tsx - Fix voice RPC response type mismatch: voice handlers return flat results (no `{result, logs}` wrapper), so remove incorrect `CommandResponse<T>` wrapping from `openhumanVoiceStatus`, `openhumanVoiceTranscribe`, `openhumanVoiceTranscribeBytes`, and `openhumanVoiceTts` Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(tauri): remove invalid infoPlist config that breaks tauri dev The `infoPlist` field in tauri.conf.json expects a string path, not an inline object. Remove it for now — microphone permission will be added via a proper Info.plist supplement in the production build pipeline. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * style: apply Prettier formatting to dictation files Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * format files * feat(dictation): integrate dictation listener and event broadcasting - Added a global dictation hotkey listener that activates based on configuration. - Implemented a web channel bridge to handle dictation events and broadcast them to connected clients. - Updated the voice module to include the new dictation listener functionality. This enhances the voice dictation capabilities by ensuring real-time event handling and client communication. * update code * format * feat(voice): enhance voice server configuration and functionality - Updated `Cargo.toml` to mark voice-related dependencies as optional. - Introduced `VoiceActivationMode` enum for better control over voice server activation. - Refactored voice server command handling and dictation event broadcasting to support new features. - Added conditional compilation for voice features across various modules, ensuring they are only included when enabled. This commit improves the modularity and configurability of the voice server, allowing for more flexible integration and usage. * refactor: clean up whitespace and formatting in core and voice modules - Removed unnecessary blank lines in `cli.rs`, `jsonrpc.rs`, `schemas.rs`, and `socketio.rs` to improve code readability. - Adjusted import order in `mod.rs` for better organization. This commit enhances the overall code quality by ensuring consistent formatting across multiple files. * chore: update Dockerfile and test workflow to install additional system dependencies - Added installation of system dependencies (cmake, ALSA, X11) in the Dockerfile for improved build support. - Updated the GitHub Actions workflow to reflect the new dependencies, ensuring consistent environment setup for testing. This commit enhances the build environment by including necessary libraries for audio and GUI support. * format * fix claude * format --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: oxoxDev <nikhil@tinyhumans.ai> |
||
|
|
07b1df4f24 |
feat(event_bus): wire webhooks, channels & skills through the event bus (#379)
* feat(event_bus): enhance domain event handling across modules - Added new `DomainEvent` variants for channel and skill events, including `ChannelMessageReceived`, `ChannelMessageProcessed`, `ChannelConnected`, `ChannelDisconnected`, `SkillLoaded`, `SkillStopped`, and `SkillStartFailed`. - Implemented event publishing in the channels and skills modules to track message processing and skill lifecycle events. - Created dedicated event bus handler files for the skills and webhooks domains, preparing for future subscriber implementations. - Updated documentation in `CLAUDE.md` to reflect the new domain events and their usage. These changes improve the observability and modularity of the system by leveraging an event-driven architecture for cross-module communication. * feat(event_bus): implement channel and webhook event handling - Introduced `ChannelInboundSubscriber` to handle inbound channel messages, triggering the agent inference loop and sending replies via the backend REST API. - Added `WebhookRequestSubscriber` to manage incoming webhook requests, routing them to the appropriate skill and handling responses. - Updated the global event bus initialization in `bootstrap_skill_runtime` to register both channel and webhook subscribers. - Enhanced `DomainEvent` with new variants for channel inbound messages and webhook requests, improving event-driven communication across modules. These changes enhance the modularity and responsiveness of the system by leveraging an event-driven architecture for channel and webhook interactions. * refactor(event_bus): update domain event documentation and subscriber initialization - Revised the documentation in `CLAUDE.md` to provide a concise overview of domain events and their associated subscriber files, enhancing clarity for future development. - Updated the `start_channels` function to initialize `WebhookRequestSubscriber` and `ChannelInboundSubscriber`, ensuring proper event handling for webhooks and channel messages. - Streamlined the event bus subscriber registration process, reinforcing the modular architecture of the system. These changes improve the maintainability and usability of the event bus framework, facilitating better cross-module communication. * style: apply cargo fmt formatting Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(event_bus): remove duplicate subscriber registration in start_channels WebhookRequestSubscriber and ChannelInboundSubscriber were registered in both bootstrap_skill_runtime() and start_channels(), causing events to be handled twice when both paths run in the same process. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(event_bus): prevent subscriber handles from being dropped on function exit SubscriptionHandle::drop aborts the background task. Since bootstrap_skill_runtime() returns immediately after setup, the local handles were dropped, cancelling both subscribers. Use std::mem::forget to leak the handles so the tasks live for the entire process. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(event_bus): ensure subscriber handles persist beyond function exit Modified the handling of subscriber registration to prevent premature dropping of handles in `bootstrap_skill_runtime()`. This change ensures that the background tasks for subscribers remain active for the entire process lifecycle, enhancing event handling reliability. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(webhooks): use proper JSON serialization for error response bodies Hand-escaped JSON strings only handled double quotes, not backslashes, newlines, or other control chars. Replaced with serde_json serialization via an error_body() helper. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(webhooks): use proper JSON serialization for error response bodies Hand-escaped JSON strings only handled double quotes, not backslashes, newlines, or other control chars. Replaced with serde_json serialization via an error_body() helper. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
5affb1e75f |
feat(integrations): add agent integration tools for Twilio, Google Places, and Parallel (#375)
* feat(integrations): add agent integration tools for Twilio, Google Places, and Parallel Add a new `src/openhuman/integrations/` module with 5 backend-proxied tools that give agents access to phone calls, location search, and web search/extraction. Each tool calls the backend API which handles external API keys, billing, rate limiting, and markup — no pricing logic on the client side. Tools added: - TwilioCallTool (scope: CLI/RPC only — requires explicit user action) - GooglePlacesSearchTool, GooglePlacesDetailsTool (scope: All) - ParallelSearchTool, ParallelExtractTool (scope: All) Includes IntegrationsConfig with per-integration toggles, shared IntegrationClient with pricing cache (fetched from backend GET /agent-integrations/pricing), ToolScope enum for future scope-based filtering, and 33 unit tests. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(integrations): add agent integration tools for Twilio, Google Places, and Parallel Add a new `src/openhuman/integrations/` module with 5 backend-proxied tools that give agents access to phone calls, location search, and web search/extraction. Each tool calls the backend API which handles external API keys, billing, rate limiting, and markup — no pricing logic on the client side. Tools added: - TwilioCallTool (scope: CLI/RPC only — requires explicit user action) - GooglePlacesSearchTool, GooglePlacesDetailsTool (scope: All) - ParallelSearchTool, ParallelExtractTool (scope: All) Includes IntegrationsConfig with per-integration toggles, shared IntegrationClient with pricing cache (fetched from backend GET /agent-integrations/pricing), ToolScope enum for future scope-based filtering, and 33 unit tests. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix code * format code * Refactor tool integration code to improve clarity and error handling - Updated `tool_loop.rs` to cache the result of `find_tool` for efficiency and clarity in logging. - Removed unused `scope` methods from `GooglePlacesSearchTool`, `ParallelSearchTool`, and `TwilioCallTool`, simplifying the codebase. - Modified `BackendResponse` structure in `mod.rs` to use `Option<T>` for `data`, enhancing error handling when no data is returned from the backend. - Updated related tests to reflect changes in the response structure and ensure proper handling of cases with no data. These changes enhance code maintainability and improve the robustness of the integration tools. * Remove comments * Add shared HTTP client for integration tools - Introduced `IntegrationClient` to manage backend interactions, including methods for POST and GET requests with error handling. - Implemented caching for pricing information fetched from the backend. - Created a new `types` module to define shared types such as `BackendResponse` and `IntegrationPricing`, enhancing code organization and clarity. - Updated `mod.rs` to include the new client and types, ensuring proper module structure. These changes improve the maintainability and functionality of the integration tools by providing a robust client for backend communication. --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |