Skip to content

Automation research session (planner · worker · critic)

Every hosted automation run is one research session, shaped like a person working a deal with an assistant: the planner reads the deal and the playbooks and asks for the next thing; the worker does it in a single continuous conversation; the critic checks what was saved and either says "that's fine" or sends a follow-up ("good — but we forgot the HOA fee"). Requires GEMINI_API_KEY on the automation worker.

Trigger fields:

Field Meaning
goal import | research | refresh | underwrite
kind refresh only — the deal-tab request (see Goals)
comps_kinds Resolved comp kinds; gates paid STR tools
_engine research_session

playbook_key stays for quotas/UI (analyze_address vs deal_tab_refresh).

Roles

Role Who Tools Decides
Planner Gemini agent with its own conversation for the run Read-only deal tools + docs (get_deal, find_deals_by_address, list_deal_*, get_agent_manual, search_docs, …) The next request to the worker, done, or pause — on import also use_deal / import_listing. Reads the playbooks once and remembers them.
Worker Gemini agent, one conversation for the whole run Full MCP catalog minus guardrails (below) How to do each request; saves through the same MCP write tools external agents use. Remembers what it already fetched.
Critic Gemini agent with its own conversation for the run Same read-only set as the planner accept, follow_up (sent straight to the worker, max 2 per request), or pause. Also gates done. Verifies saved records instead of trusting summaries.
Hooks Code — Platform follow-ups keyed on what the worker touched (below).
Brief Gemini, once at the end — Human-facing summary → result_summary. Names open risks; never mutates money fields.

Hard rules: the planner never researches; the critic never picks the roadmap; guardrails are code, not prompts.

Control loop

flowchart LR
  planner[Planner reads deal + playbooks] -->|request| worker[Worker session]
  worker --> hooks[Hooks: platform follow-ups]
  hooks --> critic[Critic verifies saved records]
  critic -->|follow_up| worker
  critic -->|accept| planner
  planner -->|done| rails[Code rails]
  rails -->|missing bar| worker
  rails --> gate[Critic done gate]
  gate -->|follow_up| worker
  gate -->|accept| brief[Brief]
  planner -->|pause| human[Human]
  critic -->|pause| human
  1. Import opening (planner): with no deal linked, the planner first calls find_deals_by_address for the run's address / listing URL (loose addresses are fine — ZIP and commas optional; match: address = same street + unit + ZIP, or + city/state when no ZIP was typed, listing_url, or possible = same house number + street nearby — the planner confirms unit/street). Existing deal → use_deal (deal_id): code links the run, runs the same setup as a fresh import (restore it from the archive if archived, reopen new/passed/watch/blocked → researching, seed focus strategies, attach the requested listing URL when it is for this house, strip other-address URLs, auction pause) and the run continues as research — a re-import asks for a fresh look, so the planner re-checks the live listing and anything stale rather than stopping because fields are filled. No deal → import_listing: the platform listing fetch (portal order by URL host) runs once; if it finds nothing, the planner asks the worker to find the listing and add it with add_manual_lead. Several plausible deals → pause. Ingest dedupe on address / URL still backs this up if the planner imports anyway.
  2. Planner gets the goal brief, the user's request (refresh), run settings, the deal snapshot (data_gaps, qa_flags, research context), what happened since its last decision, and the remaining budget.
  3. Worker handles the request as a new user message in its conversation. attach_photos puts gallery images in that message for condition / gallery work.
  4. Hooks run after every worker turn (see below).
  5. Critic reviews the turn against the request; a follow-up goes straight back to the worker.
  6. On planner done, code rails run first (each at most once): foreign listing URLs → remove them; disclosed repairs with no fix-up costs → estimate CapEx (with photos); import/research with soft gap missing_growth_drivers → research a growth driver; distressed-sale conditions (qa_flags.distress_conditions: short sale, pre-foreclosure / foreclosure, REO, probate, auction) with no matching open handoff → record an action_needed handoff of what a person must verify; underwrite with those conditions and status ranked → re-check against the distress override. Then the critic's done gate.

Hooks (deterministic)

Worker touched Code runs
add_manual_lead (import) Link the run to the deal, new → researching, seed focus strategies from the trigger, redevelop detection, strip other-address listing URLs, vision gallery clean-up, pause for auction / auction date ("Auction needs title review")
add_deal_attachment_url(s) Vision keep/drop on the listing gallery (drop faces, other homes, maps, logos)
Money inputs (update_deal, comps, appraisal/tax, repair items, OpEx items, update_deal_property) reconcile_uw_after_fill: clear unsupported rent/ARV, sync derived inputs, rescore

Worker guardrails (code)

  • Denied everywhere: automation-run tools (no recursion), workspace settings (UW defaults, financing profiles), cross-deal bulk writes, market-snapshot overrides, contact-directory admin (merge/delete/promote), buyer outreach (buyer CRUD, links, shares, messages).
  • Firecrawl: scrape, search, map, parse only.
  • import only: add_manual_lead.
  • underwrite only: update_deal_status, advance_deal, scenario create/fork/select/apply, match_deal_buyers.
  • STR only (str in comps_kinds): AirROI, VRBO, AirDNA — paid.

Goals

Goal Entry Done when
import Auto import / create_automation_run(address\|listing_url) Go/no-go ready package for human review (listing facts, photos, trends, tax, comps, strategy money inputs, ≥1 active growth driver)
research Deal Research / create_automation_run(goal=research, deal_id) Same ready-package bar (no ingest)
refresh Deal-tab research buttons / create_automation_run(deal_id, kind) The tab's request is satisfied, or what remains needs a human
underwrite Deal Underwrite / create_automation_run(goal=underwrite, deal_id) Underwrite + scenarios playbooks satisfied (thesis, dry-run advance_deal, buyer-fit gate), or pause

Refresh kind → request: listing_property (price, status, facts, write-up, attributes, HOA, agent) · listing_photos (full gallery) · listing_details (both, plus every portal URL) · listing_history (timeline events) · trends_ensure · appraisal_tax · comps · uw_fill · growth_drivers · capex · opex.

import / research / refresh never set deal status or call advance_deal — those tools are not even available outside underwrite.

Comps kinds (cost)

comps_kinds resolves from (in order) the request, income_strategies, then deal strategy_intent / income_strategy: LTR → rent · STR → str · flip/wholesale/redevelop → sale · primary → all three. STR-only paid tools are removed from the worker unless str is in scope.

Config

Env Default Meaning
AUTOMATION_PLANNER_MAX_STEPS 12 Max worker turns per run — includes critic follow-ups
AUTOMATION_MAX_PLANNER_TURNS 24 Max planner decisions per run
AUTOMATION_RUN_TIMEOUT_SECS 2400 Wall-clock cap per run
AUTOMATION_PHASE_TIMEOUT_SECS 900 Time limit for one worker / reviewer turn

Hitting a run cap pauses the run (with a brief) instead of failing it; review, then restart.

Context & token use: on long runs, older tool results are replaced with a placeholder and older history is summarized once so the run can keep going. Playbooks (get_agent_manual, read_doc) are never cleared. Each worker turn and each planner/critic decision has a model-call cap; hitting it ends the turn cleanly, and a planner/critic that hits it without deciding is asked once more to decide. Tool errors go back to the model as text.

Nothing the models read is truncated: tool results, tool and field descriptions, the worker's reports to the planner, and the deal snapshot (write-up, all comments, handoffs, listing links, prior runs) arrive whole — a cut result silently hides data (run a4fd8dd2 cut get_deal and the underwrite playbook mid-JSON). Context size is managed only after use, by clearing and summarization above. Tool schemas are only made Gemini-compatible (e.g. field titles dropped) and are sent first on every call, so they are cache hits after the first. metrics.llm.cached_input_tokens shows how much input the cache served; set AUTOMATION_LLM_USD_PER_MTOK_CACHED_INPUT for the discounted estimate. The full transcript stays in steps_json; workspace-wide lists (list_deals, list_markets, list_buyers) are not given to the worker.

Exactly-once execution: a worker must claim a run (lease_owner / lease_expires_at) before executing it and renews the lease on a heartbeat (AUTOMATION_LEASE_TTL_SECS, default 300). A redelivered queue message for a run with a live lease is skipped, so a run longer than the queue visibility timeout is never executed twice. Each worker sweeps every AUTOMATION_SWEEP_INTERVAL_SECS: runs whose lease expired (worker died) fail with "stopped responding — restart"; queued rows untouched for AUTOMATION_REQUEUE_AFTER_SECS (and <24h old) are re-enqueued.

Run metrics: automation_runs.metrics (REST/MCP metrics) rolls up LLM calls, input/output tokens and latency per role (worker, planner, critic, brief, vision), MCP tool calls/errors/latency per tool, and paid calls (apify, airroi, firecrawl). Each worker turn in steps_json carries its own usage. est_llm_cost_usd appears only when AUTOMATION_LLM_USD_PER_MTOK_INPUT / _OUTPUT are set. started_at − created_at = queue wait. Paid calls made by platform code outside agent tools (import bootstrap fetch, trends refresh) are not counted yet.

Steps in steps_json

Step Meaning
existing_deal Import linked to a deal the workspace already had (planner use_deal); the run continues as research
import_listing Platform listing fetch (import) — one child row per portal tried (running → completed / failed / skipped + what it found), updated live
plan_N Planner decision — reasoning + read-only checks in agent_log, then Next: <label> — <request>, Stop — …, Pause — …, or on import Use the existing deal — … / Look up the listing — …
turn_N One worker turn: label, instruction, agent_log (full transcript), tool_calls, usage
critic_N Critic verdict — reasoning + verification tool calls in agent_log, then Pass — …, Follow-up — …, Pause — …

Assistant entries in agent_log carry thinking — Gemini's thought summary for that model call (why it is about to call a tool, what it concluded) — alongside content (the visible reply). The UI shows it above the tool rows so a run reads as reasoning → action → result. Planner, critic and brief steps are opened as running before they decide, and their thoughts / read-only checks stream in live (each event publishes an AutomationRunUpdated), then the decision line completes the same step. | brief_1 | End-of-run brief |

MCP / REST

One entrypoint for every goal — MCP create_automation_run / REST POST /api/automation-runs — then poll get_automation_run. Restart (restart_automation_run) rebuilds the same request from the stored run.

Goal Body
import exactly one of address / listing_url; optional income_strategies, deal_purpose
research / underwrite deal_id; optional income_strategies, comps_kinds
refresh deal_id + kind

Reading runs over MCP: a finished run's steps_json is often 1M+ characters, so MCP returns it trimmed. list_deal_automation_runs defaults to detail="summary" (run fields + one line per step, no steps_json); get_automation_run defaults to detail="steps" (steps without agent_log / tool_calls). Read a transcript one step at a time: get_automation_run(run_id, detail="log", steps="turn_3") (drops thinking and tool outputs) or detail="full". REST (the web UI) still returns the full run.

goal may be omitted: kind implies refresh, a locator implies import. Common: aggressiveness. Quotas: import counts against the Quick Analyze cap; deal goals against the tab-refresh cap.