---
title: "lifelog: Cron, Files and SQL as Memory for My Agents"
date: 2026-10-06
slug: lifelog
description: "Every agent session starts with amnesia. lifelog pulls my calendar, mail, Slack, Linear, GitHub, journal and agent sessions into raw JSONL, projects them with DuckDB at read time, and hands any agent a deterministic brief of what happened lately."
---

![The lifelog dashboard: streaks, reviews, a year of activity, and per-project agent hours](/lifelog/dashboard.avif)

Every agent session I start begins with amnesia. The model does not know that I shipped three PRs yesterday, that a client emailed about a demo, that I have a call at four, or that the last session on this repo died mid-rebase. I know all of that, so I spend the first five minutes of every session typing it in. Multiply by the number of sessions I run in a day and it is a real tax.

The data was already on my machine. I had CLIs for everything: Gmail and Calendar, Slack, Linear, GitHub, Notion, Rize, and [agentsview](https://www.agentsview.io) indexing every coding-agent transcript. What was missing was one place that pulled from all of them and could answer "what happened lately?" in a form an agent could eat in one tool call.

That became [lifelog](https://github.com/eduwass/lifelog). The whole architecture fits in three words: cron, files, SQL.

## The rules I wrote down first

I had built a version of this before and thrown it away. The rewrite on July 4th was a greenfield restart with a short list of rules, and they are still the README:

- **Own raw data, not normalized data.** `data/raw/<source>/<day>.jsonl` is the source of truth. If a source changes shape, I edit SQL and re-render. I never re-ingest history.
- **Normalize at read time.** DuckDB reads the JSONL directly. There is no import step, no schema migration, no database file to corrupt.
- **Keep sources boring.** A source is a pull function plus an `events.sql` projection. No plugin framework.
- **No ORM, no editor.** The web app is read-only. Notion stays the only place I write. Runtime dependencies are the template engine and a markdown renderer.
- **Checks are part of the architecture.** Agents build most of this, so the pre-push gate and PR checks are not process, they are load-bearing.

Two months in, that is 17 sources, 56 MB of raw JSONL, and roughly 230 commits, most of them by agents working inside those rules.

```mermaid
flowchart TD
  clis["existing CLIs: gws, agent-slack, Linear, gh, ntn, agentsview, git hook"] --> pull["packages/sources/&lt;id&gt;/pull.ts"]
  pull --> raw[("data/raw/&lt;source&gt;/&lt;day&gt;.jsonl<br/>the source of truth")]
  raw --> duck[("DuckDB views, built at read time")]
  ev["packages/sources/&lt;id&gt;/events.sql<br/>per-source projection"] --> duck
  views["packages/data/views.sql<br/>cross-source joins"] --> duck
  duck --> cli["lifelog CLI: day, brief, search, explain"]
  duck --> web["read-only dashboard (PWA)"]
```

## A source is three files

Adding a source means dropping a folder under `packages/sources/`. Here is GitHub, the whole manifest:

```ts
// packages/sources/github/index.ts
export default defineSource({
  id: "github",
  label: "GitHub",
  category: "activity",
  rawDir: "github",
  lookbackDays: 3,
  pull: pullGithub,
  eventsSql: `${import.meta.dir}/events.sql`,
  health: githubHealth,
});
```

The puller shells out to `gh` and writes whatever comes back as one JSON line per record. The projection turns those raw records into the one shape everything downstream understands: `source, kind, ts, title, body`. Commits get folded into one event per repo per day so a busy day does not become 138 rows:

```sql
-- packages/sources/github/events.sql
SELECT
  'github' AS source,
  'commits' AS kind,
  max(ts::TIMESTAMPTZ) AS ts,
  repo || ' — ' || count(*) || ' commit(s)' AS title,
  string_agg('- [' || title || '](/lifelog/' || url || ') `' || left(regexp_extract(url, '[^/]+$'), 7) || '`',
             chr(10) ORDER BY ts) AS body
FROM read_json_auto('data/raw/github/*.jsonl', format = 'newline_delimited', union_by_name = true)
WHERE kind = 'commit'
GROUP BY repo, strftime(ts::TIMESTAMPTZ, '%Y-%m-%d')
```

`union_by_name = true` is the line that makes the "never re-ingest" rule work. Old files can be missing columns that new files have. DuckDB fills in nulls and moves on.

The `health.ts` file is source-owned too: freshness, auth, and a deep check. `lifelog doctor --deep` runs all of them, and a daily job posts to Slack when Google auth needs a re-login or WhatsApp is unpaired. Sources fail quietly otherwise, and a quiet failure is worse than no source at all.

## What an agent actually gets

The point of all this is one command. `lifelog brief` is a capped, deterministic orientation: shipped work, sessions, journal, inbox, upcoming. Deterministic matters. Two agents asking on the same day get the same text, and the text is stable enough to cache.

```text
# lifelog brief · 2026-09-08 → 2026-09-08

## shipped
- eduwass/eduwass.com — 1 shipped, +511 −1  →  lifelog day 2026-09-08 --json
  Post: Building a Language Server for EdgeJS in Three Days (c6204bf)
- eduwass/snap — 136 shipped, +31602 −2513  →  agentsview session get f2c2d7e1-…
  The rows table reads the daemon and the adapter it forgot (48b9a43); … +131 more

## sessions
- 2026-09-08 09:34 · snap — env (81m · $84.87)  →  agentsview session get 6eb9c706-…
  Done. `swarm/multidisplay` at `2660a357b`, `make check` green, pushed, findings posted on #273 …
- +39 more sessions — lifelog day <day> --json

## journal
- 2026-09-08 12:19 · gm gm  →  lifelog day 2026-09-08 --json
  its a new day, im better rested.. …

## upcoming
- 2026-09-15 22:00 · Goldflower Website Check-In  →  gws calendar events get --params '{…}'

_recall: lifelog search "<what you remember>" --json · depth: run the → command on any line_
```

Every line carries a drill command. That is the design in one sentence: **search finds the event, drill commands get the depth.** The brief never tries to be complete. It tries to be a table of contents with working links, so an agent that needs more runs `agentsview session get` or `gws calendar events get` itself.

Two more commands round it out:

- `lifelog search "web research sonar" --json` runs BM25 plus vector recall over the rendered day files and returns drill commands. Vector search is [qmd](https://github.com/tobi/qmd) under the hood.
- `lifelog explain search.ts:130-165 --json` maps a code range back to the commits that touched it and the agent sessions that made those commits, via `git blame` plus the view below.

That last one is the feature I use most and the one that was hardest to get right.

## Which session made this commit?

Nowadays 99% of the commits I push come out of an agent session. So the question "which session made this commit?" is the question "why does this code exist?", and it deserves a real answer. There is no single reliable signal, so lifelog uses an evidence ladder and takes the strongest rung that matches. Click through a few real shapes:

<iframe class="frame" src="evidence.html" height="420" loading="lazy" title="The session-to-commit evidence ladder, with five example commits"></iframe>

The strongest rung is a **claim**: a global git post-commit hook that writes one JSON line per commit with the agent's own session id, read straight from the environment the agent's shell exposes.

```sh
# ~/.config/git-hooks/record-claim.sh (excerpt)
if [ -n "$CLAUDE_CODE_SESSION_ID" ]; then
    agent="claude";  session="$CLAUDE_CODE_SESSION_ID"
elif [ -n "$CURSOR_CONVERSATION_ID" ]; then
    agent="cursor";  session="$CURSOR_CONVERSATION_ID"
elif [ -n "$CODEX_THREAD_ID" ]; then
    agent="codex";   session="$CODEX_THREAD_ID"
elif [ -n "$OPENCODE" ]; then
    agent="opencode"                # exposes no session id to children
else
    # walk /proc ancestry; if no agent is a parent, this was a human
fi
printf '{"ts":"%s","repo":"%s","sha":"%s","agent":"%s","session":"%s"}\n' … >> commit-claims.jsonl
```

Each agent CLI leaks its identity differently, which is why the hook is a ladder of its own. Claude Code and cursor-agent put a session id in the environment. Codex uses a thread id that happens to equal its session id. opencode only sets `OPENCODE=1`, so its commits fall through to the next rung. A real detail that cost an evening: the hook has to run inline, not backgrounded. A detached child reparents to init almost immediately, which severs the `/proc` ancestry chain and made an obviously agent-made commit read as human.

The DuckDB side ranks the rungs and unions them:

```sql
-- packages/data/views.sql (excerpt)
claim_matches AS (
  SELECT c.*, s.session_id, s.agent, 'claim' AS method, -1 AS method_rank
  FROM commits c
  JOIN claims cl ON cl.sha = c.sha_full
  JOIN sessions s ON s.session_id = cl.session OR s.session_id LIKE '%:' || cl.session
),
transcript_matches AS (
  SELECT c.*, s.session_id, s.agent, 'transcript' AS method, 0 AS method_rank
  FROM commits c
  JOIN sessions s ON true
  JOIN json_each(coalesce(s.commit_shas, '[]'::JSON)) j
    ON starts_with(c.sha_full, json_extract_string(j.value, '$'))
),
time_matches AS (
  SELECT c.*, s.session_id, s.agent, 'time' AS method, 1 AS method_rank
  FROM commits c
  JOIN sessions s
    ON slug(c.repo) = slug(s.project)
   AND c.commit_ts BETWEEN s.started_at AND s.ended_at + INTERVAL 10 MINUTE
  WHERE NOT EXISTS (SELECT 1 FROM transcript_matches tm WHERE tm.sha_full = c.sha_full)
)
```

The `method` column survives all the way to the UI, so a time-window match renders with a hedge and a claim renders as fact. And the loose commits, the ones no rung catches, became the best debugging signal in the project. Every time I saw one under a day where I knew an agent had been working, it meant a rung was broken, and it was usually the hook.

## The part I got wrong

Somewhere in July I noticed I was spending more time on the human dashboard than on the agent brief. It had swipe-between-days on mobile, favicons on time-tracking rows, a Notion-style search modal, streak badges. All nice. None of it was the goal.

The thing that pulled me back was a sentence I wrote to an agent late one night: *"I started lifelog to have a good log of everything that happens around me, simple and agent-friendly, flat files and DuckDB, so that I can build very smart agents that have immediate context and can go deep as needed. Now I'm spending more time on the little web UI than on the real goal."* The fix was not to delete the dashboard. It was to make the data model refuse to bend for it. Underneath, everything is still an event. Session grouping, PR cards, streaks: all presentation, all computed at read time from the same JSONL. If the dashboard vanished tomorrow, the brief would not notice.

Other things that bit:

- **DuckDB children that never finish.** A query can hang the CLI forever. `query()` now has a hard bound and the dashboard shares in-flight queries instead of re-running them on every request.
- **A leaky embedding process.** qmd hogged memory until another agent, seeing the hog, killed it mid-index. That one got a regression test, because a "quite serious leak" is exactly the kind of thing that returns.
- **Partial pulls reading green.** A source that fetched half its window used to look fine. Now progress is kept and a partial pull fails loudly.

## What it is not

It is not a second brain, a note-taking app, or a place I write anything. Notion is where I write. It is not a plugin system. It is not trying to be complete; the brief is capped on purpose.

It is a memory that any agent on my machine can read in one call, backed by files I could `cat`, in a format I could rebuild from scratch by editing SQL. Cron, files, SQL.

Repo: [github.com/eduwass/lifelog](https://github.com/eduwass/lifelog)
