# Marty -- Complete AI Knowledge Base & Technical Specification > Marty is an autonomous, Jarvis-class AI desktop companion designed natively for macOS (Apple Silicon M1-M4 and Intel). It combines real-time screen vision via ScreenCaptureKit, sub-100ms bidirectional voice intelligence via Google Gemini Live, and hands-free desktop automation to help professionals analyze code, spreadsheets, documents, and workflows at the speed of thought. Website: https://www.martys.space License / Pricing: 100% Free software (Bring-Your-Own-Key model via Google AI Studio) Target Platform: macOS 14 Sonoma, macOS 15 Sequoia, and later Hardware Support: Apple Silicon (M1/M2/M3/M4, Pro, Max, Ultra) with Neural Engine acceleration, and Intel Macs Local Data Directory: ~/Library/Application Support/com.honeysinghbisht.gemini/ --- ## Table of Contents 1. What is Marty? 2. Architectural Overview & Hardware Integration 3. Four Primary Interaction Surfaces 4. Complete Feature Breakdown 5. Autonomous Multi-Step Agentic Workflows 6. Connected Apps & Integration Tool Catalog 7. Tri-Tier Cognitive Memory Engine 8. macOS Permissions & Security Architecture 9. Voice Commands & Practical Prompts Cheatsheet 10. Frequently Asked Questions (FAQ) 11. Getting Started & Onboarding Guide --- ## 1. What is Marty? Marty is an autonomous native macOS companion that lives directly on your Mac desktop rather than inside a sandboxed web browser tab. Unlike conventional web chatbots, static screen-capture wrappers, or basic dictation tools, Marty operates as a persistent, ultra-low-latency voice and vision companion. ### Core Value Proposition - No Manual Screenshots or Copy-Pasting: Marty continuously sees your screen (1 FPS stream via ScreenCaptureKit) or inspects regions via native crosshair snip. - Zero Context Switching: Marty works directly over your current active window (Xcode, Figma, Terminal, Safari, Preview, Numbers). - Sub-100ms Bidirectional Audio: Full-duplex streaming audio with Apple VoiceProcessingIO acoustic echo cancellation allows natural conversational interruptions without waiting for beeps. - Multi-App Agentic Automation: Executes complex multi-step chains across Apple Mail, Apple Notes, Apple Pages, Google Chrome, Safari, and macOS system settings. - 100% Privacy & BYOK: Bring your own Gemini API key from Google AI Studio. Stored locally with zero training on your personal files or queries. --- ## 2. Architectural Overview & Hardware Integration - Screen Perception Pipeline: Built on Apple ScreenCaptureKit. Streams frames at 1 FPS (1280x720) during active sessions. Employs on-device Apple Vision framework OCR to parse text, numbers, formulas, and visual boundaries before generating optimized multimodal payloads. - Audio Engineering Pipeline: Full-duplex audio stream with 16 kHz Int16 PCM microphone input and 24 kHz Float32 high-fidelity audio output. Features Apple VoiceProcessingIO Audio Unit for hardware Acoustic Echo Cancellation (AECAudioStream) so incoming speaker output does not bleed into the mic, enabling smooth voice barge-in. - Supported Voice Personas: 7 expressive Gemini Live voices: Capella, Puck, Charon, Kore, Fenrir, Aoede, Pegasus. - Supported Gemini Models: * Gemini 3.7 Flash (Default -- Optimal balance of speed, multimodal vision reasoning, and agentic tool use) * Gemini 3.5 Flash-Lite (Fastest possible conversational response times) * Gemini 3.1 Pro (Deepest reasoning for complex multi-file coding and algorithmic tasks) - On-Device Wake Word: Implemented via Apple Speech.framework for continuous, low-power, zero-network wake-word detection ("Hey Marty" or custom aliases like "Jarvis", "Friday", "Nova"). - Native macOS Automation Runtime: Interfaces with macOS Accessibility APIs (AXUIElement), Apple Events / AppleScript, and Apple Shortcuts to drive native applications without web-extension overhead. - Storage & Encryption: Facts, preferences, and vector memory embeddings are stored in encrypted local storage. --- ## 3. Four Primary Interaction Surfaces Marty adapts to diverse workflows with four distinct surfaces: 1. Floating Bubble HUD (Dynamic Island for macOS) - Always-on-top translucent glass pill floating over fullscreen IDEs, browsers, or keynotes. - Interactive 9-bar reactive audio waveform showing live mic input and AI voice synthesis. - Expandable mini-console featuring real-time subtitles, quick region-snip trigger, mute button, and instant dismiss. 2. On-Device Voice Wake Word ("Hey Marty") - 100% local wake word detection with zero network transmission during standby. - Responds to "Hey Marty" (or custom aliases such as "Jarvis", "Friday", "Nova"). - Triggers an instant acoustic feedback chime and immediate session activation. 3. Full Canvas Workspace (Main Window) - Comprehensive desktop window with multi-session conversational history. - Input composer supporting typed text, screenshots, and file attachments (PDFs, images, code files). - Live tool execution visualizer showing real-time background actions (e.g. searching the web, reading Safari DOM, drafting Mail). - Integrated Memory Inspector for viewing and editing stored facts. 4. Global Keyboard Shortcuts & Menu Bar Tray - Default hotkey: Option + Space (⌥ Space) or Command + Option + Space (⌘⌥ Space). - Menu bar item shows connection health, active microphone level, and one-click quick settings. --- ## 4. Complete Feature Breakdown ### A. Live Screen Perception & Document Understanding - Focused Window Awareness: Automatically detects the active application, window coordinates, and active document title. - Visual OCR & Layout Comprehension: Reads financial spreadsheets (Numbers/Excel), code errors (Xcode/VS Code), long PDF documents (Preview), and vector designs (Figma). - Multi-Monitor Support: Reads content across multiple external monitors connected to MacBook, Mac mini, Mac Studio, or Mac Pro. - Native Region Snip: Crosshair tool (screencapture -i -r -x) to select rectangular screen areas for high-resolution inspection. - Instant Diagnostics: Identifies calculation mistakes in spreadsheets or syntax errors in code and explains the resolution aloud in seconds. ### B. Autonomous Browser Agents (Google Chrome & Apple Safari) - Semantic DOM Extraction: Parses web pages into structured outlines and maps interactive elements to numbered badges: [1], [2], [3]. - Precision Element Clicking: Clicks links, buttons, and dropdowns by element ID or CSS selector. - Intelligent Input: Types text into search boxes and textareas via synthetic events without capturing your hardware keyboard. - Multi-Field Form Autofill: Fills complex forms in a single step using key-value mappings. - Tab Management: Lists, switches, creates, reloads, and closes tabs across browser windows. ### C. Apple Productivity Suite Agent - Apple Mail: Reads inboxes, retrieves full message bodies, drafts replies, and sends emails upon verbal confirmation. - Apple Notes: Searches notes by keyword, reads note contents, creates formatted notes in specific folders, and appends action items. - Apple Pages: Inspects active document text and word counts, creates documents, appends paragraphs, and exports to PDF. ### D. Silent Background Web Agent - Headless Search: Queries DuckDuckGo and Wikipedia in the background without opening disruptive browser windows. - Clean Content Extraction: Fetches clean Markdown text from web articles and speaks distilled takeaways. ### E. macOS System Control & Window Snapping - Application Control: Opens, switches, and gracefully quits installed macOS apps. - Window Management: Verbal window snapping side-by-side or into split-screen layouts. - Hardware Settings: Controls system volume, mute, display brightness, Dark Mode, and screen lock. --- ## 5. Autonomous Multi-Step Agentic Workflows Marty excels at compound instructions that require chaining multiple tools across different applications: ### Workflow 1: Cross-App Research & Briefing Pipeline - User Prompt: "Marty, look up the top 3 open-source vector databases in the background, create a comparison note in Apple Notes with bullet points, and draft an email to the engineering team with the summary." - Step 1: Silently queries search engines via the background web agent. - Step 2: Parses and summarizes the clean content. - Step 3: Creates and formats a new note in Apple Notes (notes_create_note). - Step 4: Opens Apple Mail and creates a formatted draft (mail_send_email). - Step 5: Speaks the completed status and asks for verbal confirmation to send. ### Workflow 2: Visual Bug Triage & Fix Workflow - User Prompt: "Marty, look at the error on my screen, search DuckDuckGo for the solution, navigate Safari to the GitHub pull request, and paste the suggested fix into the comment box." - Execution: ScreenCaptureKit frame -> Vision OCR / error analysis -> background web search -> Safari DOM element identification -> synthetic form input. ### Workflow 3: Meeting Preparation & Action Items - User Prompt: "Check my unread emails from the design team, extract their key feedback into my 'Design Review' note, and open Pages with our spec document." - Execution: Fetches unread emails -> parses action items -> appends to Apple Notes -> launches Pages. --- ## 6. Connected Apps & Integration Tool Catalog | Application | Category | Supported Tools | Capabilities | |---|---|---|---| | Apple Mail | Productivity | mail_get_recent_or_unread, mail_read_message, mail_send_email | Read inboxes, read bodies, draft replies, send emails with verbal confirmation. | | Apple Notes | Productivity | notes_list_or_search, notes_read_note, notes_create_note, notes_append_text | Keyword search, read notes, create formatted notes, append items. | | Apple Pages | Productivity | pages_get_active_document, pages_create_document, pages_append_text, pages_export_pdf | Inspect active document text and word count, draft documents, export PDF. | | Google Chrome | Browsers | chrome_read_page, chrome_click_element, chrome_type_text, chrome_fill_form, chrome_navigate, chrome_search_web, chrome_manage_tabs | Semantic DOM reading, numbered badge element clicking, typing, form fill, tab management. | | Apple Safari | Browsers | safari_read_page, safari_click_element, safari_type_text, safari_fill_form, safari_navigate, safari_search_web, safari_manage_tabs | Autonomous DOM reading, element clicking, smart input, form fill, tab control. | | macOS System | System | open_app, close_app, system_control, type_text, press_keyboard | App launch/quit, accessibility typing, volume/mute control, lock screen, Shortcuts execution. | --- ## 7. Tri-Tier Cognitive Memory Engine Marty features a persistent cognitive memory engine managed locally on your Mac: 1. Tier 1: In-Memory Short-Term Conversation Buffer - Maintains the last 6 turns of live dialogue for immediate conversational context. 2. Tier 2: Declarative Explicit Memory (JSON on Disk) - Stores key-value facts and preferences across Personal, Work, and Habit categories. - Triggered when you say "Marty, remember that..." - Tools: save_declarative_memory(key, value, category), get_declarative_memory(key). 3. Tier 3: Long-Term Semantic Vector Memory - Semantic embeddings generated via Apple NLEmbedding framework. - Stored locally and queried using Accelerate framework cosine similarity. - Tools: store_long_term_memory(content, memory_type), search_long_term_memory(query). ### Built-in Memory Inspector Users maintain full transparency and sovereignty over their data. Open the Memory Inspector from the main canvas to view, edit, search, or delete stored memories anytime. --- ## 8. macOS Permissions & Security Architecture Marty is engineered to comply strictly with macOS security sandboxing: - Local-First Audio: Wake word detection operates 100% on-device. Audio only streams to Gemini over encrypted WebSocket (wss://) while an active session is live. - Local Storage: Declarative facts and vector memories reside locally in ~/Library/Application Support/com.honeysinghbisht.gemini/ - Safety Guardrail: Marty will never send an email, close unsaved work, or execute high-impact actions without asking for your explicit verbal confirmation. - No Arbitrary Shell Execution: Arbitrary terminal shell commands are disabled by design for maximum system safety. - Automatic App Blacklisting: Automatically blinds perception when password vaults (1Password, Bitwarden, Keychain), banking websites, or private browsing tabs are active. ### Required macOS Permissions 1. Microphone: For voice input and echo-cancelled conversation. 2. Speech Recognition: For on-device wake-word detection (Speech.framework). 3. Accessibility: For typing into apps and navigating UI elements. 4. Screen Recording: For live screen perception via ScreenCaptureKit. 5. Automation (Chrome / Safari / Mail / Notes / Pages): For autonomous app driving. 6. "Allow JavaScript from Apple Events": Enabled in Chrome (View -> Developer) and Safari (Develop menu). --- ## 9. Voice Commands & Practical Prompts Cheatsheet ### Screen & Document Analysis - "Marty, look at my screen--what error is this compiler showing?" - "Take a snip of this diagram and extract the text." - "What is the hex color code of the button on my screen?" - "Read through this 5-page PDF in Preview and give me a 3-bullet summary." - "Look at the table in my spreadsheet and tell me which month had the highest expenses." ### General Research & Knowledge - "Hey Marty, what's the latest news on NASA's Artemis mission?" - "Marty, what was Apple's Q3 services revenue?" - "Summarize the article currently open in my Safari tab." - "Who is the CEO of the company mentioned in this press release?" ### Communications & Writing - "Draft an email to Alex saying the designs look great and confirm our Tuesday 10 AM sync." - "Do I have any urgent unread emails in Mail today?" - "Create a note called Workout Plan with 4 days of splits." - "Add 'Buy oat milk' to my Grocery note." - "Create a new document in Pages titled Q4 Product Strategy and add an outline for our mobile launch." ### Multi-Step Agentic Workflows - "Marty, summarize the open article in Safari, create a note titled 'AI Trends 2026', and draft an email to the product team with the summary." - "Look at the stack trace on my screen, search for the fix in the background, and append the solution to my 'Debugging Notes'." - "Check my unread emails from Sarah, draft a reply confirming tomorrow's meeting, and save a reminder to my todo note." ### Desktop & System Controls - "Open Visual Studio Code." - "Snap Safari to the left half and Notes to the right half." - "Turn the volume down 20%." - "Turn on Dark Mode and mute system volume." - "Lock my screen." - "Pause screen perception while I log into my password vault." - "Thanks Marty, goodbye!" --- ## 10. Frequently Asked Questions (FAQ) ### What is Marty? Marty is an autonomous AI assistant built natively for macOS (Apple Silicon and Intel). It floats seamlessly over your desktop, sees what is on your screen in real time, listens and speaks with zero delay, and performs multi-step actions across your native Mac apps hands-free. ### Is Marty free to use? Yes, Marty is completely free. It operates on a Bring-Your-Own-Key model where you connect your own Gemini API key (which can be obtained for free from Google AI Studio at aistudio.google.com). This gives you full ownership over your AI usage and data privacy with zero subscription fees or hidden costs. ### How do I get a Gemini API key? You can generate a free Gemini API key in less than a minute at Google AI Studio (aistudio.google.com). Once you install and open Marty on your Mac, simply paste your key into the onboarding screen or Settings. ### How does Screen Vision work? When you ask a question or summon Marty, it inspects your active focused window or entire multi-monitor setup in real time via ScreenCaptureKit. It reads text, data tables, spreadsheet formulas, code errors, PDFs, and design mockups directly on screen without manual screenshots or copy-pasting. ### How does Real-Time Voice work? Marty uses a low-latency conversational voice engine with full interruptibility powered by Gemini Live API over WebSockets. With Apple VoiceProcessingIO acoustic echo cancellation, you can speak naturally in full sentences without waiting for beeps, pause mid-thought, or speak over Marty at any moment. ### Can Marty control and automate my Mac apps? Yes. Marty connects directly with system events, AppleScript, Accessibility APIs, and macOS Shortcuts. You can issue cross-app commands like researching details online, drafting an email in Apple Mail, saving takeaways into Apple Notes, or rearranging windows on your screen in one continuous flow. ### Is my screen and voice data secure? Yes. Marty is built around on-device security principles. Audio listening and screen perception only activate when you explicitly summon Marty. Sensitive applications (password managers, banking sites, private browsing tabs) are automatically blacklisted, and personal data is never used to train public AI models. ### Does Marty work offline? Core desktop controls--including window management, Apple Shortcuts execution, local memory lookup, and keyboard commands--function locally on your Mac without internet access. Complex multimodal reasoning and vision analysis use your private Gemini API key over encrypted connections. ### What is Smart Memory? Marty retains your preferences, project context, and formatting guidelines across sessions in an on-device encrypted database using Apple NLEmbedding and Accelerate cosine similarity. If you tell Marty to always draft emails under 100 words in bullet points, it remembers and applies that style automatically. ### What macOS versions and hardware are supported? Marty requires macOS 14 (Sonoma), macOS 15 (Sequoia), or later. It is natively compiled for Apple Silicon (M1, M2, M3, M4, Pro, Max, and Ultra) with Neural Engine acceleration, and also supports Intel-based Macs. ### How do I summon or dismiss Marty on my Mac? You can summon Marty anytime by pressing Option + Space (⌥ Space) anywhere in macOS, by saying the optional wake phrase "Hey Marty", or by clicking the capsule in your menu bar. Press Escape (Esc) to dismiss it instantly. ### How do I get Early Access? Visit https://www.martys.space/early-access and enter your name and email. Early access invites with direct .dmg installer downloads and setup guides are distributed to waitlist members on a weekly basis. --- ## 11. Getting Started & Onboarding Guide 1. Download & Launch: Open Marty.app. The Floating Bubble HUD appears in the top-right corner, and the Menu Bar item appears in your menu bar. 2. Grant Permissions: Follow the Permissions Setup Sheet to grant Microphone, Speech Recognition, Accessibility, Screen Recording, and Automation for supported apps. 3. Add Gemini API Key: Open Settings (gear icon), paste your free Gemini API key from Google AI Studio, and select your preferred model (Gemini 3.7 Flash recommended) and voice persona. 4. Start Interacting: Press Option + Space, say "Hey Marty", or click the microphone to speak naturally.