Hawky

AI agent that sees and hears in realtime

Hawky streams audio and video from your iPhone or Ray-Ban Meta glasses to a realtime model. A persistent backend agent stores context, uses tools, and continues tasks after the conversation ends.

What Hawky actually does

A look at what Hawky can do today, from recognizing people and spotting hazards to reminders, meeting mode, and visual memory.

Person identification

Live

At a party or in a meeting, Hawky recognizes who you're talking to and quietly briefs you — names, context, and what to say next.

Safety

Live

Hawky watches your surroundings and warns you about hazards in real time, before they reach you.

Reminder

In progress

Hawky notices the commitments you make in conversation and surfaces the right reminder at the right moment.

Meeting Mode

In progress

For moments when Hawky shouldn't speak. It follows along and helps without audio — a glance at your glasses display or lock screen is all it takes.

Visual Memory

Preview

"What did I do yesterday?" "What did I just do?" Hawky answers from a searchable visual record of your day — down to "how many times did my dog appear in the last 10 minutes?"

Realtime agent

Hawky connects live audio and camera input to a realtime conversation, persistent memory, and a backend agent that can use tools.

First-person vision

Stream camera input from an iPhone or Ray-Ban Meta glasses so the realtime model can respond to the scene in front of you.

Realtime conversation

Speak naturally, interrupt responses, and combine voice with current visual context.

Multiple live inputs

Combine phone audio, glasses video, typed messages, and connected-device events in one session.

Session memory

Conversation history and selected observations persist across sessions, so each interaction can start with relevant context.

Searchable visual history

Find previously captured moments using questions such as “Where did I leave this?” or “When did I last see that?”

Backend actions

Delegate longer tasks to the Hawky backend, where an agent can use tools, files, memory, and external services.

Architecture

The realtime model handles the moment. Hawky keeps context, uses tools, and continues work beyond the conversation.

1 · Surfaces

See and hear the moment

iPhone · Ray-Ban Meta · Web

cameramicrophonetext
2 · Realtime loop

Respond immediately

Natural turn-taking, interruption, and low-latency speech.

GPT LiveGemini Live
3 · Hawky bridge

Share context and delegate

Routes sessions and hands durable requests to the backend agent.

sessionsauthstreaming
4 · Backend agent

Remember and act

Uses longer reasoning and keeps working after the live turn ends.

memorytoolsfilesautomations
verified context and action results return to the live conversation
Provider capabilities differ: GPT Live focuses on realtime audio and text; Gemini Live can also use a live camera stream. Hawky provides the shared bridge to memory and actions.

What's next?

Where Hawky's vision and reach widen next.

  • 01

    Benchmark + agent gym

    Realtime-agent benchmark plus an auto-testing gym with synthetic tests.

  • 02

    Recursive self-improvement

    Agent builds agent: prompt optimization, sharper tooling, live dev.

  • 03

    CarPlay

    Hawky in the car — briefing and acting from the dashboard.

  • 04

    Latency-adaptive sampling

    Sample perception at a rate tuned to live latency.

  • 05

    Device support

    Android, Apple Watch, and more glasses beyond Ray-Ban Meta.

  • 06

    Perception pipeline

    Face recognition and speaker diarization to map audio to who's talking.

  • 07

    Personality & memory

    Tunable soul/identity and shareable memory presets.

Citation

Hawky is open source on GitHub. If it helps you, buy us more tokens or give the project a star ⭐!

github.com/hao-ai-lab/hawky
@software{hawk_ambient_agent,
  title  = {Hawky: An Ambient AI Agent},
  author = {The Hawky Team},
  year   = {2026},
  url    = {https://github.com/hao-ai-lab/hawky}
}