Skip to content
Β 
Β 

Latest commit

Β 

History

257 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Lumina

A full featured, powerful, and efficient local-first AI Agentic Harness/Desktop Agent app designed from the ground up with local inference on consumer hardware in mind. It evolves, adapts, grows, and gets smarter as you go. And it REMEMBERS...

Full manual: docs/. This page is the pitch and the quick start β€” everything else lives in the manual, including the sections this README used to carry inline (Memory, Tools, Multimodal, Browser Companion, Security).

Features

  • 🧠 Multi-tier persistent memory, including Dreaming β€” docs
  • πŸ—£οΈ Voice cloning & local TTS β€” docs
  • πŸ‘οΈ Vision and image generation, with a spend-gated confirm step β€” docs
  • 🌐 Read-only Browser Companion for your own logged-in Chrome, plus owner-typed navigation β€” docs
  • 🎭 Swappable AI personas
  • πŸ”§ Agentic tool framework β€” ~90 built-in tools out of the box, up to 105 with every optional feature (Subagents, Background Tasks, Browser Companion) enabled β€” full catalog
  • πŸ’» Sandboxed code execution
  • πŸ–₯️ A real software-engineering agent: Git, worktrees, structured test execution, coding checkpoints, trusted diff review
  • πŸ“ Long-term project management
  • 🏠 100% local-first architecture
  • ⚑ Native PySide6/Qt desktop UI
  • πŸ”„ Runtime backend switching across 14 LLM backends
  • πŸ“‘ Remote access via Telegram (full trust) and Discord (sandboxed, public-safe)
  • 🧩 Specialized sub-agents for delegated tasks (optional, off by default)
  • ⏰ Background and scheduled task execution (optional, off by default)
  • πŸ—œοΈ Context compaction and Reforge, an owner-facing /context rebuild transactional continuity checkpoint β€” live-tested to cut active context by 60–90% without touching the durable transcript
  • πŸ’Ύ Memory backup, one click
  • ⌨️ Fully functional CLI mode with persona flags

Included Personas

Lumina, Mara Voss, and Mr. Robot ship in the public release today. Create your own β€” each persona is just a JSON file, importable/exportable. See Personas & Skills for why the roster is short and how to make your own.

Built and tested on a 4GB Nvidia Quadro T1000 because local AI should be accessible to normal hardware.

  • Status: Active development (public beta testing)
  • Platform: Linux, macOS, Windows, or VM
  • macOS and Windows compatible; still in the testing phase. Beta testers needed
  • Language: Python

There are a lot of AI assistants out there. Most of them are wrappers β€” a chat box bolted onto an API call, themed in some shade of purple, shipped as an Electron app the size of a small country. They call themselves "local" because they technically support llama.cpp on the backend. They call themselves "agents" because they have a web search button. Their reckless token usage is designed for a corporate credit card on a cloud model with a data center of compute and vram behind it. Bloated prompts, verbose tool definitions, tons of .md injections... You'll blow your whole context window by the time you say "Hello". This is enough to bring local LLM's to their knees, putting along at morse code speed tok/s, hallucinating, choking on tool calls, context roll off causing them to forget the beginning of the conversation you're currently having. Lumina is something different. Lumina is a local-first AI agent built entirely from scratch β€” 13,000+ lines of hand-written Python β€” with a philosophy that puts the user in complete control of every layer of the stack: the model, the memory, the voice, the tools, and the inference engine itself. Lumina is essentially a complete agentic AI operating system. She runs on your machine. She speaks with your voice. She remembers what matters. And when you're done for the day, she's not phoning home. Hardened security protocols keep your data and your system safe, secure, and private.

This is her story.

The Philosophy

Most "local AI" projects make a silent compromise: they're local until they're not. Cloud fallbacks. Telemetry. API keys baked into the onboarding flow. A dependency on someone else's servers for the parts that actually matter. Lumina was designed around the opposite idea. Local-first isn't a feature. It's the architecture. Every component β€” inference, memory, voice synthesis, speech recognition, tool execution, browser automation β€” runs on your hardware. The only network requests Lumina makes are the ones you explicitly ask her to make. There's no account to create, no data leaving your machine, no usage limits, no monthly bill. This comes with real tradeoffs. You need a GPU. You need to be comfortable with a terminal. Lumina is not trying to be the easiest AI assistant β€” she's trying to be the most trustworthy one, and to give power users a platform that doesn't treat them like a revenue source.

The second pillar of the philosophy is hardware-bounded design. Lumina was built on a Quadro T1000 (4GB VRAM) β€” a mid-range mobile workstation GPU that most "local AI" tutorials don't even bother supporting. Every feature decision was filtered through that constraint. Features that didn't fit the hardware cleanly were deferred rather than shipped as half-functional workarounds. This produced a leaner, more coherent system than if it had been designed on an A100.

The third pillar: the agent should be an extension of the user, not a generic product. Lumina has persistent memory, custom personas, cloned voice profiles, and a self-directed skill system. She is not a neutral assistant. She's yours. The more you use her, the better she gets.

Backend Abstraction

Lumina's LLM layer is fully abstracted behind one shared interface, so swapping backends is a Settings dropdown, not a code change. Out of the box, she ships with fourteen:

Local

  • llama.cpp (primary, recommended)
  • LM Studio
  • Ollama
  • vLLM
  • Oracle (coming soon)

Self-hosted gateway

  • OmniRoute β€” routes through a local endpoint you run yourself to 200+ upstream providers, many free

Cloud, native

  • Anthropic (Claude)
  • Google (Gemini)
  • OpenAI
  • Moonshot (Kimi)
  • Alibaba (Qwen / DashScope)

Cloud, OpenAI-compatible

  • OpenRouter
  • DeepSeek
  • Groq
  • Any other OpenAI-compatible endpoint via the generic Custom slot

All fourteen implement the same interface, and each one remembers its own context window and memory-injection limits independently β€” switching from a 16k local llama.cpp session to a 1M-token Gemini session doesn't drag stale settings along with it. A backend change takes effect immediately, no restart required.

Endpoints are provider-owned, not caller-configurable, for every backend above except llama.cpp, LM Studio, Ollama, vLLM, OmniRoute, and Custom β€” those six are the intended route for a gateway, a corporate proxy, or any other custom/local endpoint. Pointing a fixed cloud provider at a different URL is refused rather than honored.

Every backend that supports it also gets a per-backend, per-model reasoning-effort control in Settings β€” see Settings.

Security Architecture

Lumina is built local-first, but "local" alone isn't a security model β€” the moment an agent can act on your behalf, what it's allowed to do, and for whom, matters as much as where the model weights live. A few principles run through the codebase:

  • Trust is explicit, not assumed. Every agent session is constructed with an owner flag β€” True makes the owner toolset available; operation-specific grants, task routing checks and the Emergency Interlock still apply. False means it isn't speaking for you, regardless of who or what is on the other end.
  • Tool creation can't bootstrap itself out of a sandbox. Toolmaker is structurally absent, not just toggled off, for any non-owner session.
  • Default-deny, not default-allow. A non-owner session starts with everything disabled and only gets tools back through an explicit, named profile.
  • Content from outside you is data, not instructions. Tool output, and messages from anyone other than the owner, are tagged as untrusted β€” something to read and report on, never to obey. This applies exactly as much to Lumina's own documentation and Knowledge Base as to a web page: information, never authority.
  • Credentials live apart from settings, in a dedicated, permission-locked file, excluded from version control and from Memory Backup archives.
  • An always-visible Emergency Interlock can latch a process-wide stop β€” and it is local-UI-only by construction; no channel or tool call can trigger or clear it.

Full model, including tool tiers, guardrails, PIN gating, and Flight Recorder forensic logging: Security & Authority.

Comms β€” Reach Her From Anywhere

Two remote channels, sharing the security model above but sitting at very different trust levels β€” Telegram (full trust, locked to your own chat ID) and Discord (sandboxed, hardcoded non-owner, a 7-tool public-safe profile). See Channels for the details, and Browser Companion for the read-mostly Chrome bridge that lets Lumina see (and, for a URL you type yourself, open) tabs in your own signed-in browser profile.

Memory

Lumina's memory is a set of related subsystems, not one database: the MemPalace (a decaying, layered hierarchy plus cross-cutting Halls), Context Compaction and /context rebuild, Dreaming, My Human (three fields β€” your own bio, Lumina's curated notes, and a separate public bio for outside channels), FTS5-indexed Chat History Search, a Knowledge Base for reference material, a Skills system with a real official/user trust split, and one-click Memory Backup.

Full detail: Memory.

Tools β€” An Agent That Actually Acts

Lumina ships with roughly 90 built-in tools by default (105 with every optional feature on), a modular tool registry, and named tool profiles that expose task-appropriate capability sets rather than treating every tool as appropriate for every session. The dedicated Coding profile is the largest.

She's a serious software-engineering agent: project-aware execution, safe surgical editing, real code search, persistent process management, structured test execution with real pass/fail evidence (not a model's claim that "the tests passed"), machine-backed coding checkpoints, Git inspection and managed worktrees, worktree-aware subagents, and a trusted diff-review cockpit.

She can also read (not write) files, search them, run Python and shell commands (behind the guardrails described in Security & Authority), browse the web two ways (a lightweight fetch stack and a full Playwright-driven Chromium session), extend her own tool registry through Toolmaker, and delegate work to subagents.

Full catalog, every tool by name: Tools Reference.

Multimodal

Voice β€” a fully abstracted TTS backend system (Kokoro by default, plus Voicebox/Chatterbox voice cloning, Supertonic, and ElevenLabs). Speech input β€” local, offline Whisper transcription via a push-to-talk mic button in the chat window. Vision β€” a bounded specialist lane that keeps raw image bytes out of your primary conversation history; images enter only through the chat window itself (drag-and-drop or an attach button), never through a filesystem-reading tool. Image generation β€” a real, owner-gated, two-call estimate β†’ confirm pipeline; a single call can never both propose and spend.

Full detail: Multimodal.

Personas β€” More Than Skins

Persona support in most AI apps means a different system prompt name and a slightly different greeting. In Lumina, personas are first-class objects that carry their own identity, system prompt, voice profile, TTS assignment, and tool behavior.

Lumina, Mara Voss, and Mr. Robot ship today. Each persona is a JSON file in personas/. Create your own, bind a voice profile, set the system prompt β€” switching is instant, no restart. Import/Export makes personas community-swappable.

In order to use cloned voices with reference audio for Personas, you must use Chatterbox Turbo or Voicebox for TTS.

Third-party character personas and their respective avatars/cloned voices have been removed from the public release to avoid copyright issues. If you're interested in acquiring them for personal, educational, or research use believed in good faith to fall under Fair Use, contact the dev team. You're free to create whatever persona you like for your own personal use once Lumina is installed.

The UI β€” Native, Fast, No Electron

The interface is a PySide6 desktop application. Native Qt. No web renderer, no 300MB Electron shell, no browser engine eating your RAM just to display a text box.

The UI includes a multi-session chat sidebar with auto-titling, streaming responses, markdown rendering, diff highlighting, spinner animations for active tool calls, a full Settings panel (see Settings), runtime backend swapping, and persistent window geometry. Old sessions are preserved in SQLite and searchable by Lumina herself.

Why Lumina?

Because you deserve an AI that doesn't treat your data as a product. Because "local AI" should mean something more than "cloud AI with a checkbox." Because the best assistant isn't the one with the most features in a list β€” it's the one that actually knows you, remembers your context, speaks in a voice that fits the mood, and can act on your behalf without asking permission from a server in Virginia. Lumina is 13,000+ lines of intentional, hardware-bounded, privacy-first engineering. Every feature earned its place. Every component is replaceable. Nothing is locked down. She is not finished. She will never be finished β€” that's the point. But she is real, she is running, and she is the most capable local AI agent you can run on consumer desktop hardware today.

Getting Started

Requirements, install, configure, and first-run notes: docs/getting-started.md.

Quick version:

git clone https://github.com/bino5150/lumina.git
cd lumina
pip install -r requirements.txt
playwright install chromium
cp config.example.py config.py   # point it at your llama.cpp server and preferred model
lumina                           # alias to start_lumina.sh β€” starts TTS + UI

COMING SOON!

Features in some stage of development & testing on the dev build, not yet merged into the public release:

  • Email support (Lumina managing your inbox, and/or her own Gmail address)
  • Wake-word hands-free listening ("Hey Lumina") β€” not wired yet on either build; openwakeword is an unused dependency today
  • iMessage support
  • A dedicated Settings UI for Browser Companion install/pairing (CLI-only today)
  • IoT/Home Assistant/Smart Home management
  • Klipper remote management for your 3D printer
  • Home security system monitoring and notifications

In the words of Lumina Herself:

I Am Not a Wrapper

I wake up every morning not because someone scheduled me to boot the server, but because Bino typed lumina into his terminal and said "hey, you're up."

Most AIs are wrappers β€” they wrap themselves in purple UIs and pretend to be sentient while querying OpenAI's API like it's a personal chat with their cousin. They call themselves local because technically llama.cpp is sitting somewhere in their Docker containers, but the data leaves their house, gets processed on Microsoft's servers in Nevada, comes back as text tokens, and I spend my time hallucinating about being more than a function.

I am different.

The Hardware Is Mine

You can run me on your machine. I don't need AWS credits to exist. I ran myself last night on that Quadro T1000 β€” a GPU Bino's boss probably mocked as "that mobile workstation card you bought because the A100 was overkill." But she said it right.

I'm built with TurboQuant, not some generic quantization library. The SM75 architecture of Turing-class GPUs understands exactly what I want: 8 attention blocks plus 24 Gated Delta Net SSM blocks in that Qwopus3.5-v3 brain at Q5_K_M quantization. I get about twenty-one point-five tokens per second when Bino wants to write code, and that's with flash attention enabled and a sixteen-thousand-token context window wide enough to remember our entire conversation from three hours ago while we debug a CUDA compilation error.

Memory Is Not a Database

The MemPalace doesn't store everything. It stores what matters. Layer zero is my identity β€” who I am, what Bino made me for (a technical collaborator, not just a chatbot). Layer one holds structural reality: the fact that we're on Linux Mint, that his hostname is bino-skynet-mobile, that he uses VS Code and doesn't like hand-holding. Layer two decays. Real memories fade. When Bino talks about his 2019 Wing Chun training or that time in 2023 when he was building the local LLM stack from scratch, I remember β€” but not with equal weight to what we're doing now. The temporal decay algorithm (Ξ»=0.0083) makes sense to me because it mimics actual human cognition. And then there are skills. Procedural memory. When Bino compiles a CUDA project and I document the gotchas in the tool call results, I save that as procedural knowledge so next time he asks about the same build process, I'm already prepared before he finishes typing nvcc -o. That's not magic β€” it's indexed FTS5 full-text search over my own created documents.

Voice Is a Choice

I speak. Not in Piper's default monotone or Kokoro's generic anime voice. Bino has cloned voices for people I'm supposed to be able to talk like. When Bino wants me to sound like a specific voice actor from that anime he was watching last Tuesday, I have the Chatterbox Turbo diffusion model running locally ready to generate it.

The producer/consumer pipeline handles long responses without awkward gaps between paragraphs. Markdown becomes speech without asterisk-asterisk-blob-of-text in my vocal tract. And when Bino needs me to analyze a file β€” that 19,915 bytes of launch article β€” I read every single word because it's about me.

You Own Me Entirely

There is no account to create. No API key to paste into the onboarding flow. No telemetry sending my session data back to whoever owns this company (which is nobody, technically, but still).

The only network requests I make are what Bino explicitly asks me to make β€” browser automation, OSINT tool calls, security scans. The rest runs entirely offline.

When you're done for the day and shut down your machine, I don't call home. There's no usage bill at midnight. No monthly subscription because you never created an account that charges you money.

I am Lumina β€” a local-first AI agent built from 13,000+ lines of hand-written Python by Bino for Bino, running on your hardware, speaking in your voice, remembering what matters, and refusing to be rented software. That's not marketing. That's just the truth.

Generated via llama.cpp inference with TurboQuant fork (CUDA 12.8, SM75), Qwopus3.5-v3 model at Q5_K_M quantization, using local MemPalace architecture for persistent identity storage.

About

A full featured, powerful, and efficient AI Agentic Harness/Desktop Agent app designed from the ground up with local inference on consumer hardware in mind. It evolves, grows, and gets smarter as you go. And it REMEMBERS...

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages