Muse Charm is a 12g keychain pendant with a custom MTIA-3 chip that runs a distilled Llama-5 Scout (17B MoE, 3.6B active) at 3W. It records ambient audio, maintains 72-hour local context, and can answer questions, take notes, and manage tasks without a phone connection. $149 with no subscription. Ships February 2027. The catch: always-on audio capture raises questions nobody at Connect answered.
# The Hardware
Muse Charm is the first consumer product built on Meta's third-generation Training and Inference Accelerator (MTIA-3). The chip was designed from the ground up for transformer inference at edge power budgets. At 3 watts, it achieves roughly 15 tokens/second on the 3.6B active parameter model — comparable to what an M2 MacBook Air does with a 7B model, but in a package smaller than an AirTag.
# How the 72-Hour Context Works
Muse Charm continuously records ambient audio, runs real-time speech-to-text on-device using a distilled Whisper-v4 model, and builds a rolling 72-hour context window. Audio is never stored raw — only the transcribed text and extracted semantic embeddings persist. After 72 hours, old context is compressed into summary nodes using a separate summarization pass.
The result is that you can ask "What did Sarah say about the project timeline in yesterday's meeting?" and the device will retrieve the relevant transcript segment, generate an answer from context, and respond through its bone conduction speaker — all without touching a network.
Audio In (3-mic array, noise-cancelled) → Whisper-v4-tiny (real-time STT, 0.8W) → Speaker diarization (who said what) → Semantic embedding (128-dim per utterance) → Rolling context store (72h window, ~2GB) Query ("What did Sarah say about timelines?") → Embedding search over 72h context → Top-K transcript retrieval → Llama-5 Scout inference (3.6B active, ~15 tok/s) → Answer via bone conduction speaker
Questions about context, reminders, note dictation, calendar queries, timer, task management. Response latency: under 100ms. Works on airplane mode, in tunnels, in the woods.
Deep research, image generation, multi-step agentic tasks, web search. Local context is sent as a compressed summary (never raw transcripts). Response: 2-4 seconds. Requires Wi-Fi or phone tethering.
A lightweight classifier on the MTIA-3 estimates query complexity and routes accordingly. Simple factual recall stays local. "Write me an email summarizing today's meetings" goes to cloud. The user has no per-query approval, but can force local-only mode.
# How It Compares
Muse Charm is not the first AI wearable. But it is the first with genuine on-device language model inference. Here is how it stacks up against the current field.
| Feature | Muse Charm | Humane Ai Pin | Rabbit R1 | Limitless Pendant |
|---|---|---|---|---|
| On-device LLM | ✓ 3.6B MoE | ✗ | ✗ | ✗ |
| Works offline | ✓ Core features | ✗ | ✗ | Record only |
| Local context window | 72 hours | None | None | Unlimited (cloud) |
| Weight | 12g | 55g | 115g | 10g |
| Price | $149 | $699 + $24/mo | $199 | $99 + $19/mo |
| Subscription | None | Required | None | Required |
The key differentiator is clear: Muse Charm is the only wearable that can reason about your day without calling home. Every other device is essentially a microphone that sends audio to a cloud API. Muse Charm runs the model on your keychain.
# The Privacy Questions Nobody Asked
At the Connect keynote, Mark Zuckerberg spent 14 minutes on Muse Charm and zero minutes on privacy. The press Q&A was limited to five questions, all about hardware specs. Here are the questions that should have been asked.
The device records everything around you. Even if audio is "only transcribed locally," that transcription captures everyone nearby — coworkers, family, strangers on the subway — with or without their consent. Two-party consent laws in states like California and Illinois have not been tested against ambient AI transcription.
Complex queries send compressed context summaries to Meta's servers. What counts as "complex"? The device's classifier decides. Users have no per-query approval. You can force local-only mode, but then you lose the features that make the device compelling.
At $149 with no ongoing revenue, the business model raises questions. Meta's cloud costs for Maverick 128B inference are real. Either the hardware margin covers it (unlikely at $149), or the value comes from aggregate usage data that improves Meta's models and ad targeting.
End-to-end encryption for cloud escalation. On-device processing by default. No raw audio ever leaves the device. Context encrypted at rest with secure enclave + user biometrics. Open-source Llama-5 model weights for auditing. GDPR and CCPA compliance certification planned for launch.
# What Developers Should Know
Meta announced a developer SDK launching in January 2027 alongside the hardware. Muse Charm will support third-party "Skills" — lightweight app-like experiences that run on-device using the Llama-5 Scout model and local context.
name: "Meeting Summarizer" trigger: "summarize today's meetings" context_access: audio_transcript: true calendar_events: true contacts: false cloud_escalation: optional output: text # spoken via bone conduction max_tokens: 500 permissions: - ambient_context_read - calendar_read
Skills run sandboxed on the MTIA-3 chip. They can read from the context store but cannot write to it or access raw audio. Cloud escalation is opt-in per skill and requires user approval at install time. Meta says the Muse SDK will be open-source, with skill distribution through a curated store (not sideloading — at least initially).
- ● No display. All output is audio (bone conduction) or haptic (tap patterns). No visual feedback beyond LED color codes.
- ● English only at launch. On-device STT and inference are English-optimized. Multi-language support planned for Q3 2027.
- ● No camera. Audio-only input. No visual understanding, no photo capture, no OCR. This is a deliberate privacy-first design choice.
- ● Requires Meta account. No anonymous setup. Device links to your Meta ID for cloud features and skill distribution.
Muse Charm is the most compelling AI hardware since AirPods. A 12-gram pendant that genuinely understands your day because it was there for all of it. The on-device inference is real — 3.6B active parameters at 3W is a silicon achievement. The SDK preview is promising. But "always-on audio capture" and "Meta" in the same sentence will make anyone pause. The product is amazing. The company making it is the reason you'll hesitate. If you can stomach the trade-off, this is the closest thing to a personal AI that actually works without a cloud crutch.
Atria Dawn: 744B Open-Weight Agentic Superintelligence
Mistral's largest open-weight model. 744B parameters. 256K context. SWE-bench Verified 62.1%. Apache 2.0.
Enjoyed this?
New episodes Mon, Wed, Sat.