Automation research session (planner · worker · critic)¶
Every hosted automation run is one research session, shaped like a person working a deal with an assistant: the planner reads the deal and the playbooks and asks for the next thing; the worker does it in a single continuous conversation; the critic checks what was saved and either says "that's fine" or sends a follow-up ("good — but we forgot the HOA fee"). Requires GEMINI_API_KEY on the automation worker.
Trigger fields:
| Field | Meaning |
|---|---|
goal |
import | research | refresh | underwrite |
kind |
refresh only — the deal-tab request (see Goals) |
comps_kinds |
Resolved comp kinds; gates paid STR tools |
_engine |
research_session |
playbook_key stays for quotas/UI (analyze_address vs deal_tab_refresh).
Roles¶
| Role | Who | Tools | Decides |
|---|---|---|---|
| Planner | Gemini agent with its own conversation for the run | Read-only deal tools + docs (get_deal, find_deals_by_address, list_deal_*, get_agent_manual, search_docs, …) |
The next request to the worker, done, or pause — on import also use_deal / import_listing. Reads the playbooks once and remembers them. |
| Worker | Gemini agent, one conversation for the whole run | Full MCP catalog minus guardrails (below) | How to do each request; saves through the same MCP write tools external agents use. Remembers what it already fetched. |
| Critic | Gemini agent with its own conversation for the run | Same read-only set as the planner | accept, follow_up (sent straight to the worker, max 2 per request), or pause. Also gates done. Verifies saved records instead of trusting summaries. |
| Hooks | Code | — | Platform follow-ups keyed on what the worker touched (below). |
| Brief | Gemini, once at the end | — | Human-facing summary → result_summary. Names open risks; never mutates money fields. |
Hard rules: the planner never researches; the critic never picks the roadmap; guardrails are code, not prompts.
Control loop¶
flowchart LR
planner[Planner reads deal + playbooks] -->|request| worker[Worker session]
worker --> hooks[Hooks: platform follow-ups]
hooks --> critic[Critic verifies saved records]
critic -->|follow_up| worker
critic -->|accept| planner
planner -->|done| rails[Code rails]
rails -->|missing bar| worker
rails --> gate[Critic done gate]
gate -->|follow_up| worker
gate -->|accept| brief[Brief]
planner -->|pause| human[Human]
critic -->|pause| human
- Import opening (planner): with no deal linked, the planner first calls
find_deals_by_addressfor the run's address / listing URL (loose addresses are fine — ZIP and commas optional;match:address= same street + unit + ZIP, or + city/state when no ZIP was typed,listing_url, orpossible= same house number + street nearby — the planner confirms unit/street). Existing deal →use_deal(deal_id): code links the run, runs the same setup as a fresh import (restore it from the archive if archived, reopennew/passed/watch/blocked→researching, seed focus strategies, attach the requested listing URL when it is for this house, strip other-address URLs, auction pause) and the run continues asresearch— a re-import asks for a fresh look, so the planner re-checks the live listing and anything stale rather than stopping because fields are filled. No deal →import_listing: the platform listing fetch (portal order by URL host) runs once; if it finds nothing, the planner asks the worker to find the listing and add it withadd_manual_lead. Several plausible deals →pause. Ingest dedupe on address / URL still backs this up if the planner imports anyway. - Planner gets the goal brief, the user's request (
refresh), run settings, the deal snapshot (data_gaps,qa_flags, research context), what happened since its last decision, and the remaining budget. - Worker handles the request as a new user message in its conversation.
attach_photosputs gallery images in that message for condition / gallery work. - Hooks run after every worker turn (see below).
- Critic reviews the turn against the request; a follow-up goes straight back to the worker.
- On planner
done, code rails run first (each at most once): foreign listing URLs → remove them; disclosed repairs with no fix-up costs → estimate CapEx (with photos);import/researchwith soft gapmissing_growth_drivers→ research a growth driver; distressed-sale conditions (qa_flags.distress_conditions: short sale, pre-foreclosure / foreclosure, REO, probate, auction) with no matching open handoff → record anaction_neededhandoff of what a person must verify;underwritewith those conditions and statusranked→ re-check against the distress override. Then the critic's done gate.
Hooks (deterministic)¶
| Worker touched | Code runs |
|---|---|
add_manual_lead (import) |
Link the run to the deal, new → researching, seed focus strategies from the trigger, redevelop detection, strip other-address listing URLs, vision gallery clean-up, pause for auction / auction date ("Auction needs title review") |
add_deal_attachment_url(s) |
Vision keep/drop on the listing gallery (drop faces, other homes, maps, logos) |
Money inputs (update_deal, comps, appraisal/tax, repair items, OpEx items, update_deal_property) |
reconcile_uw_after_fill: clear unsupported rent/ARV, sync derived inputs, rescore |
Worker guardrails (code)¶
- Denied everywhere: automation-run tools (no recursion), workspace settings (UW defaults, financing profiles), cross-deal bulk writes, market-snapshot overrides, contact-directory admin (merge/delete/promote), buyer outreach (buyer CRUD, links, shares, messages).
- Firecrawl:
scrape,search,map,parseonly. importonly:add_manual_lead.underwriteonly:update_deal_status,advance_deal, scenario create/fork/select/apply,match_deal_buyers.- STR only (
strincomps_kinds): AirROI, VRBO, AirDNA — paid.
Goals¶
| Goal | Entry | Done when |
|---|---|---|
import |
Auto import / create_automation_run(address\|listing_url) |
Go/no-go ready package for human review (listing facts, photos, trends, tax, comps, strategy money inputs, ≥1 active growth driver) |
research |
Deal Research / create_automation_run(goal=research, deal_id) |
Same ready-package bar (no ingest) |
refresh |
Deal-tab research buttons / create_automation_run(deal_id, kind) |
The tab's request is satisfied, or what remains needs a human |
underwrite |
Deal Underwrite / create_automation_run(goal=underwrite, deal_id) |
Underwrite + scenarios playbooks satisfied (thesis, dry-run advance_deal, buyer-fit gate), or pause |
Refresh kind → request: listing_property (price, status, facts, write-up, attributes, HOA, agent) · listing_photos (full gallery) · listing_details (both, plus every portal URL) · listing_history (timeline events) · trends_ensure · appraisal_tax · comps · uw_fill · growth_drivers · capex · opex.
import / research / refresh never set deal status or call advance_deal — those tools are not even available outside underwrite.
Comps kinds (cost)¶
comps_kinds resolves from (in order) the request, income_strategies, then deal strategy_intent / income_strategy: LTR → rent · STR → str · flip/wholesale/redevelop → sale · primary → all three. STR-only paid tools are removed from the worker unless str is in scope.
Config¶
| Env | Default | Meaning |
|---|---|---|
AUTOMATION_PLANNER_MAX_STEPS |
12 | Max worker turns per run — includes critic follow-ups |
AUTOMATION_MAX_PLANNER_TURNS |
24 | Max planner decisions per run |
AUTOMATION_RUN_TIMEOUT_SECS |
2400 | Wall-clock cap per run |
AUTOMATION_PHASE_TIMEOUT_SECS |
900 | Time limit for one worker / reviewer turn |
Hitting a run cap pauses the run (with a brief) instead of failing it; review, then restart.
Context & token use: on long runs, older tool results are replaced with a placeholder and older history is summarized once so the run can keep going. Playbooks (get_agent_manual, read_doc) are never cleared. Each worker turn and each planner/critic decision has a model-call cap; hitting it ends the turn cleanly, and a planner/critic that hits it without deciding is asked once more to decide. Tool errors go back to the model as text.
Nothing the models read is truncated: tool results, tool and field descriptions, the worker's reports to the planner, and the deal snapshot (write-up, all comments, handoffs, listing links, prior runs) arrive whole — a cut result silently hides data (run a4fd8dd2 cut get_deal and the underwrite playbook mid-JSON). Context size is managed only after use, by clearing and summarization above. Tool schemas are only made Gemini-compatible (e.g. field titles dropped) and are sent first on every call, so they are cache hits after the first. metrics.llm.cached_input_tokens shows how much input the cache served; set AUTOMATION_LLM_USD_PER_MTOK_CACHED_INPUT for the discounted estimate. The full transcript stays in steps_json; workspace-wide lists (list_deals, list_markets, list_buyers) are not given to the worker.
Exactly-once execution: a worker must claim a run (lease_owner / lease_expires_at) before executing it and renews the lease on a heartbeat (AUTOMATION_LEASE_TTL_SECS, default 300). A redelivered queue message for a run with a live lease is skipped, so a run longer than the queue visibility timeout is never executed twice. Each worker sweeps every AUTOMATION_SWEEP_INTERVAL_SECS: runs whose lease expired (worker died) fail with "stopped responding — restart"; queued rows untouched for AUTOMATION_REQUEUE_AFTER_SECS (and <24h old) are re-enqueued.
Run metrics: automation_runs.metrics (REST/MCP metrics) rolls up LLM calls, input/output tokens and latency per role (worker, planner, critic, brief, vision), MCP tool calls/errors/latency per tool, and paid calls (apify, airroi, firecrawl). Each worker turn in steps_json carries its own usage. est_llm_cost_usd appears only when AUTOMATION_LLM_USD_PER_MTOK_INPUT / _OUTPUT are set. started_at − created_at = queue wait. Paid calls made by platform code outside agent tools (import bootstrap fetch, trends refresh) are not counted yet.
Steps in steps_json¶
| Step | Meaning |
|---|---|
existing_deal |
Import linked to a deal the workspace already had (planner use_deal); the run continues as research |
import_listing |
Platform listing fetch (import) — one child row per portal tried (running → completed / failed / skipped + what it found), updated live |
plan_N |
Planner decision — reasoning + read-only checks in agent_log, then Next: <label> — <request>, Stop — …, Pause — …, or on import Use the existing deal — … / Look up the listing — … |
turn_N |
One worker turn: label, instruction, agent_log (full transcript), tool_calls, usage |
critic_N |
Critic verdict — reasoning + verification tool calls in agent_log, then Pass — …, Follow-up — …, Pause — … |
Assistant entries in agent_log carry thinking — Gemini's thought summary for that model call (why it is about to call a tool, what it concluded) — alongside content (the visible reply). The UI shows it above the tool rows so a run reads as reasoning → action → result. Planner, critic and brief steps are opened as running before they decide, and their thoughts / read-only checks stream in live (each event publishes an AutomationRunUpdated), then the decision line completes the same step.
| brief_1 | End-of-run brief |
MCP / REST¶
One entrypoint for every goal — MCP create_automation_run / REST POST /api/automation-runs — then poll get_automation_run. Restart (restart_automation_run) rebuilds the same request from the stored run.
| Goal | Body |
|---|---|
import |
exactly one of address / listing_url; optional income_strategies, deal_purpose |
research / underwrite |
deal_id; optional income_strategies, comps_kinds |
refresh |
deal_id + kind |
Reading runs over MCP: a finished run's steps_json is often 1M+ characters, so MCP returns it trimmed. list_deal_automation_runs defaults to detail="summary" (run fields + one line per step, no steps_json); get_automation_run defaults to detail="steps" (steps without agent_log / tool_calls). Read a transcript one step at a time: get_automation_run(run_id, detail="log", steps="turn_3") (drops thinking and tool outputs) or detail="full". REST (the web UI) still returns the full run.
goal may be omitted: kind implies refresh, a locator implies import. Common: aggressiveness. Quotas: import counts against the Quick Analyze cap; deal goals against the tab-refresh cap.