lifelog: Cron, Files and SQL as Memory for My Agents

Every agent session I start begins with amnesia. The model does not know that I shipped three PRs yesterday, that a client emailed about a demo, that I have a call at four, or that the last session on this repo died mid-rebase. I know all of that, so I spend the first five minutes of every session typing it in. Multiply by the number of sessions I run in a day and it is a real tax.
The data was already on my machine. I had CLIs for everything: Gmail and Calendar, Slack, Linear, GitHub, Notion, Rize, and agentsview indexing every coding-agent transcript. What was missing was one place that pulled from all of them and could answer "what happened lately?" in a form an agent could eat in one tool call.
That became lifelog. The whole architecture fits in three words: cron, files, SQL.
The rules I wrote down first
I had built a version of this before and thrown it away. The rewrite on July 4th was a greenfield restart with a short list of rules, and they are still the README:
- Own raw data, not normalized data.
data/raw/<source>/<day>.jsonlis the source of truth. If a source changes shape, I edit SQL and re-render. I never re-ingest history. - Normalize at read time. DuckDB reads the JSONL directly. There is no import step, no schema migration, no database file to corrupt.
- Keep sources boring. A source is a pull function plus an
events.sqlprojection. No plugin framework. - No ORM, no editor. The web app is read-only. Notion stays the only place I write. Runtime dependencies are the template engine and a markdown renderer.
- Checks are part of the architecture. Agents build most of this, so the pre-push gate and PR checks are not process, they are load-bearing.
Two months in, that is 17 sources, 56 MB of raw JSONL, and roughly 230 commits, most of them by agents working inside those rules.
A source is three files
Adding a source means dropping a folder under packages/sources/. Here is GitHub, the whole manifest:
// packages/sources/github/index.ts
export default defineSource({
id: "github",
label: "GitHub",
category: "activity",
rawDir: "github",
lookbackDays: 3,
pull: pullGithub,
eventsSql: `${import.meta.dir}/events.sql`,
health: githubHealth,
});
The puller shells out to gh and writes whatever comes back as one JSON line per record. The projection turns those raw records into the one shape everything downstream understands: source, kind, ts, title, body. Commits get folded into one event per repo per day so a busy day does not become 138 rows:
-- packages/sources/github/events.sql
SELECT
'github' AS source,
'commits' AS kind,
max(ts::TIMESTAMPTZ) AS ts,
repo || ' — ' || count(*) || ' commit(s)' AS title,
string_agg('- [' || title || '](' || url || ') `' || left(regexp_extract(url, '[^/]+$'), 7) || '`',
chr(10) ORDER BY ts) AS body
FROM read_json_auto('data/raw/github/*.jsonl', format = 'newline_delimited', union_by_name = true)
WHERE kind = 'commit'
GROUP BY repo, strftime(ts::TIMESTAMPTZ, '%Y-%m-%d')
union_by_name = true is the line that makes the "never re-ingest" rule work. Old files can be missing columns that new files have. DuckDB fills in nulls and moves on.
The health.ts file is source-owned too: freshness, auth, and a deep check. lifelog doctor --deep runs all of them, and a daily job posts to Slack when Google auth needs a re-login or WhatsApp is unpaired. Sources fail quietly otherwise, and a quiet failure is worse than no source at all.
What an agent actually gets
The point of all this is one command. lifelog brief is a capped, deterministic orientation: shipped work, sessions, journal, inbox, upcoming. Deterministic matters. Two agents asking on the same day get the same text, and the text is stable enough to cache.
# lifelog brief · 2026-09-08 → 2026-09-08
## shipped
- eduwass/eduwass.com — 1 shipped, +511 −1 → lifelog day 2026-09-08 --json
Post: Building a Language Server for EdgeJS in Three Days (c6204bf)
- eduwass/snap — 136 shipped, +31602 −2513 → agentsview session get f2c2d7e1-…
The rows table reads the daemon and the adapter it forgot (48b9a43); … +131 more
## sessions
- 2026-09-08 09:34 · snap — env (81m · $84.87) → agentsview session get 6eb9c706-…
Done. `swarm/multidisplay` at `2660a357b`, `make check` green, pushed, findings posted on #273 …
- +39 more sessions — lifelog day <day> --json
## journal
- 2026-09-08 12:19 · gm gm → lifelog day 2026-09-08 --json
its a new day, im better rested.. …
## upcoming
- 2026-09-15 22:00 · Goldflower Website Check-In → gws calendar events get --params '{…}'
_recall: lifelog search "<what you remember>" --json · depth: run the → command on any line_
Every line carries a drill command. That is the design in one sentence: search finds the event, drill commands get the depth. The brief never tries to be complete. It tries to be a table of contents with working links, so an agent that needs more runs agentsview session get or gws calendar events get itself.
Two more commands round it out:
lifelog search "web research sonar" --jsonruns BM25 plus vector recall over the rendered day files and returns drill commands. Vector search is qmd under the hood.lifelog explain search.ts:130-165 --jsonmaps a code range back to the commits that touched it and the agent sessions that made those commits, viagit blameplus the view below.
That last one is the feature I use most and the one that was hardest to get right.
Which session made this commit?
Nowadays 99% of the commits I push come out of an agent session. So the question "which session made this commit?" is the question "why does this code exist?", and it deserves a real answer. There is no single reliable signal, so lifelog uses an evidence ladder and takes the strongest rung that matches. Click through a few real shapes:
The strongest rung is a claim: a global git post-commit hook that writes one JSON line per commit with the agent's own session id, read straight from the environment the agent's shell exposes.
# ~/.config/git-hooks/record-claim.sh (excerpt)
if [ -n "$CLAUDE_CODE_SESSION_ID" ]; then
agent="claude"; session="$CLAUDE_CODE_SESSION_ID"
elif [ -n "$CURSOR_CONVERSATION_ID" ]; then
agent="cursor"; session="$CURSOR_CONVERSATION_ID"
elif [ -n "$CODEX_THREAD_ID" ]; then
agent="codex"; session="$CODEX_THREAD_ID"
elif [ -n "$OPENCODE" ]; then
agent="opencode" # exposes no session id to children
else
# walk /proc ancestry; if no agent is a parent, this was a human
fi
printf '{"ts":"%s","repo":"%s","sha":"%s","agent":"%s","session":"%s"}\n' … >> commit-claims.jsonl
Each agent CLI leaks its identity differently, which is why the hook is a ladder of its own. Claude Code and cursor-agent put a session id in the environment. Codex uses a thread id that happens to equal its session id. opencode only sets OPENCODE=1, so its commits fall through to the next rung. A real detail that cost an evening: the hook has to run inline, not backgrounded. A detached child reparents to init almost immediately, which severs the /proc ancestry chain and made an obviously agent-made commit read as human.
The DuckDB side ranks the rungs and unions them:
-- packages/data/views.sql (excerpt)
claim_matches AS (
SELECT c.*, s.session_id, s.agent, 'claim' AS method, -1 AS method_rank
FROM commits c
JOIN claims cl ON cl.sha = c.sha_full
JOIN sessions s ON s.session_id = cl.session OR s.session_id LIKE '%:' || cl.session
),
transcript_matches AS (
SELECT c.*, s.session_id, s.agent, 'transcript' AS method, 0 AS method_rank
FROM commits c
JOIN sessions s ON true
JOIN json_each(coalesce(s.commit_shas, '[]'::JSON)) j
ON starts_with(c.sha_full, json_extract_string(j.value, '$'))
),
time_matches AS (
SELECT c.*, s.session_id, s.agent, 'time' AS method, 1 AS method_rank
FROM commits c
JOIN sessions s
ON slug(c.repo) = slug(s.project)
AND c.commit_ts BETWEEN s.started_at AND s.ended_at + INTERVAL 10 MINUTE
WHERE NOT EXISTS (SELECT 1 FROM transcript_matches tm WHERE tm.sha_full = c.sha_full)
)
The method column survives all the way to the UI, so a time-window match renders with a hedge and a claim renders as fact. And the loose commits, the ones no rung catches, became the best debugging signal in the project. Every time I saw one under a day where I knew an agent had been working, it meant a rung was broken, and it was usually the hook.
The part I got wrong
Somewhere in July I noticed I was spending more time on the human dashboard than on the agent brief. It had swipe-between-days on mobile, favicons on time-tracking rows, a Notion-style search modal, streak badges. All nice. None of it was the goal.
The thing that pulled me back was a sentence I wrote to an agent late one night: "I started lifelog to have a good log of everything that happens around me, simple and agent-friendly, flat files and DuckDB, so that I can build very smart agents that have immediate context and can go deep as needed. Now I'm spending more time on the little web UI than on the real goal." The fix was not to delete the dashboard. It was to make the data model refuse to bend for it. Underneath, everything is still an event. Session grouping, PR cards, streaks: all presentation, all computed at read time from the same JSONL. If the dashboard vanished tomorrow, the brief would not notice.
Other things that bit:
- DuckDB children that never finish. A query can hang the CLI forever.
query()now has a hard bound and the dashboard shares in-flight queries instead of re-running them on every request. - A leaky embedding process. qmd hogged memory until another agent, seeing the hog, killed it mid-index. That one got a regression test, because a "quite serious leak" is exactly the kind of thing that returns.
- Partial pulls reading green. A source that fetched half its window used to look fine. Now progress is kept and a partial pull fails loudly.
What it is not
It is not a second brain, a note-taking app, or a place I write anything. Notion is where I write. It is not a plugin system. It is not trying to be complete; the brief is capped on purpose.
It is a memory that any agent on my machine can read in one call, backed by files I could cat, in a format I could rebuild from scratch by editing SQL. Cron, files, SQL.
Repo: github.com/eduwass/lifelog
Read other posts →