Launch HN: Screenpipe (YC S26) – Record how you work and turn that into agents

88 points by louis030195 ↗ HN
Hi Hacker News, I'm Louis. I built Screenpipe (https://screenpipe.com), an app that records your screen and audio locally (only!), and gives AI agents a searchable memory of what you've seen, said, and heard. This makes it easier to automate your repetitive tasks, turn them into SOPs (Standard Operating Procedure) and so on.

I made a HN-style demo video at https://www.tella.tv/video/build-your-ai-second-brain-with-s... and there’s a marketing video at https://www.youtube.com/watch?v=c1jV6E9pyug.

I’ve been obsessed with this for a long time. I’ve been maintaining a “second brain” since 2020, in which I would store journals, handwritten notes, music I listen to, projects I'm working on, conversations I have with people, personal CRM etc. I experimented a lot of RAG in the early days with ParlAI, hundreds of fine-tuned GPT2 models, and GPT3 (https://forum.obsidian.md/t/fine-tuning-openai-api-gpt3-on-y...). Later I built Ava, the first Obsidian AI plugin, which grew to a few thousands of users quickly. It then became Embedbase, an API to make it easier to build AI apps powered by RAG.

What I learned from all this is how important it is for the models to have context about what you’re doing on your computer, in order to get them to do what you want.

In the early days there was fine tuning but it was too much pain, then there was tool calling so that AI can access software you use but still kinda not autonomous enough. needing micro management. Then MCP came, but it felt too static, and non technical users struggled to build and use MCP. Then we got skills. Most recently we’ve seen Karpathy’s LLM-maintained wiki, Garry's GBrain, etc., where an agent incrementally maintains a persistent collection of Markdown pages. New sources update entity pages, strengthen or contradict existing claims, and improve a synthesis that compounds over time. I like this pattern, but it still begins with someone selecting and importing the sources. There is still no way AI can know what you and your company are doing every day, across apps, not just inside of apps.

Of course, not everyone wants this. But I do! I want AI to know what I'm doing and never lose memory ever again, and I want it to use the same software that humans do, without painful context switches.

I started building Screenpipe for myself in 2024 - a CLI to record your screen and plug this context into AI. An HN user posted it in 2024 (https://news.ycombinator.com/item?id=41695840) and that discussion influenced the product. The most useful criticism concerned recording consent, local security, CPU usage, signal-to-noise, and whether agents could act on top of the data.

The naive implementation started from continuously recording video and running OCR over every frame. But that creates duplicate data, consumes substantial resources (it basically turns your computer into a space heater!), and discards structure the operating system already knows. Screenpipe now instead listens for events such as app switches, clicks, typing pauses, scrolling, and idle fallbacks. When something meaningful changes, it pairs a screenshot with the operating system’s accessibility tree at the same timestamp. OCR is used when structured accessibility data is unavailable. We also capture audio continuously, identify speakers and transcribe locally through Parakeet/Whisper or using cloud models.

Everything is indexed in a local SQLite database, mp4 files, and sometimes md files. An AI fr...

36 comments

[ 3.4 ms ] story [ 21.7 ms ] thread
How are you planing to segregate between professional use and personal use, I dont want any agent or any llm to know all the time what i have been doing on my system, it would be privacy nightmare and most of the time the screencapture is not meaningful. People may use their work laptop or devices to checkout reddit or hackernews occasionally.
How does it run the LLMs? Or does it call a API to llama.cpp/ollama/etc. ?

During normal usage, how often does it try to parse info from the screen capture? Once a minute?

Been a user of Screenpipe to build some "Ai-buddies" for me since last year. Core of that is giving AI a look at what is happening in time space. Without Screenpipe this was rather hard to do at the performance screenpipe gives.
zero chance im trusting any cloud or third party SaaS

if this doesn't run fully local its a no go for enterprise let alone ordinary users

This sounds so ripe for abuse and so terrifying if it were that I don't even want to experiment with it. Sorry.
Implementation specifics aside, I increasingly think life and death cycles are preferable. The opposite of omnipotent context, I want cycles of clean slates so that baggage - of all kinds - does not weigh down the new.
Funny timing. I've been building something similar in my spare time called Daydream. There’s a lot of overlap: local screen/audio capture, OCR and transcription, window and activity context, SQLite, and a searchable memory of what happened.

The main difference is the product direction. Screenpipe seems focused on continuously giving agents context through APIs, MCP, and skills. Daydream is more narrowly built around answering "what did I do today?" through a timeline you can inspect, replay, search, and turn into a daily digest.

I'm also treating deletion as part of the data model. If you cut a sensitive span, its frames, audio, OCR, transcripts, embeddings, and summaries should be deleted or invalidated too.

Mine is still early and Linux-first. I'm open-sourcing it in case anyone wants to contribute, poke around, or use it as a starting point. It’s built with Tauri, a Rust backend, React/TypeScript, SQLite, GStreamer, Whisper, OCR, and VLM processing.

I genuinely didn’t know you were building this when I started. Apparently personal memory capture is becoming a SaaS category too lol.

Code is here: https://github.com/snackbit/daydream

How do you secure the database on my device?
Well, i'm not sure who your customers are (except companies in Russia, NK..), but european employers would break laws by allowing your software to be used, long even before AI became a thing and especially after the EU AI Act has passed. To put it mildly, no sane corporation would allow such software to be used which tells me a bit about the sanity of the current VC landscape.
you mean train your agents to replace myself and everybody else?
note the audio level on the tella.tv video is significantly lower (as compared to the youtube video)
(comment deleted)
I tried this ~ a week ago, tried one of the suggested automations and it pretty immediately started sending my local api keys to an endpoint… afk right now but can share more when I’m back
You had my interest until "source available". Especially when you used MIT from the start and changed it later. FOSS hackers like me feel betrayed by moves like that because your product was built on FOSS others created and you are not paying that forward.

Also, I for one would never invest a second in tools I cannot freely modify and share the code of under OSI terms. I would strongly suggest a convenience tax model. Hackers will self host and maybe contribute but those with more money than time will put in a credit card. Maybe offer end to end encrypted memory and compute to secure enclaves where additional compute on the data can be done when the laptop is closed. (Shameless plug, this is what https://caution.co enables, and 100% FOSS)

However, with "source available" you are just begging for someone to AI launder your code into a FOSS clone you will have 0 recourse on. If you FOSS it yourself then you get to capture the FOSS community destined to form around this idea. Some of that community will have bosses that will pay you.

MIT may not be the right play though. AGPL is a good middle ground as it flips the script. Corpo lawyers are allergic to AGPL and will pay for an alternative license, but for community hackers that might want to improve and recommend your code it offers no restrictions.