Tier limits

Check where you stand:
For current pricing, see bestmate.io or email kaya@forever22.com. The table above is the enforced limits, which move independently of price.

What “sources” means

This trips people up, so it’s worth being precise — there are two different caps and they count different things: Connecting Granola is one connection, no matter whether it imports three meetings or three hundred. Importing a 500-note Obsidian vault is one connection and counts against the article ceiling, not the connection one.

Trials

A workspace trial covers usage — asking, drafting, the Inbox. It does not lift the storage caps, which bind from day one. Some workspace tiers (design partner, and paid plans) are trial-exempt and never expire.

Caching

Identical queries — same question, same twin, same workspace — are cached for 5 minutes. Cached responses return without re-running retrieval or the model, so repeated questions and rapid testing are cheap.

Concurrency

Requests are limited per workspace (5 concurrent). Exceeding it isn’t an error you need to handle — the request waits its turn.

Where the money actually goes

Measured over 30 days to 2026-08-24, roughly 80% of LLM spend was background work — nightly consolidation, insight grading, wiki regeneration — against 9% for user-facing generation. Background work runs on its own model tier with a single runtime dial, so the cost of the whole background pipeline can be moved without a deploy and without touching anything a person is waiting on.