7 accounts
Bounded public-source set
Give an agent a specific sales job: review an account list, find a relevant signal or inspect a campaign. Start with the Codex example below, then use the framework to decide what your agent should own.
Open the Codex example
CodexOne folder. A reviewable result.
Start with our seven-company packet. Run its local checker, then inspect the decisions and the missing facts.
01 Get the example
Unzip it, then open the codex-company-review-packet folder in your terminal or local Codex project.
codex-company-review-packet/ ├── ledger.json account decisions ├── source-manifest.json sources and dates ├── contract.json review criteria ├── prompts/ next-step recipes └── review_packet.py offline checker
02 Run the offline check
Python 3.9+ recommended, standard library only. The checker makes no model or network calls.
Open the extracted folder in your existing local Codex workspace, then use this tutorial recipe.
03 Inspect the result
out/review-summary.jsonCounts, structural-check status and the next review boundary.
out/reviewed-ledger.csvCompany, decision, reason, source references and next action.
These are the preserved September 23, 2026 decisions. The checker validates structure and exports files. It does not research companies, verify source facts or change decisions. Human review remains pending.
Start a separate exercise in an existing Codex workspace with suitable permissions. This prompt is a new recipe, not the original run transcript.
Start with a small company-research task and a file you can check. In our September 23, 2026 Codex desktop exercise, 12 search queries produced seven company records. After agent review, none was retained: four were rejected and three needed research. The useful output was a sourced decision ledger, including what prevented a match.
Download the Codex packet (ZIP)
The packet contains the reviewed JSON and CSV, source links, a cleaned run receipt, proposed prompts and a Python checker. You can reproduce the offline check without running a model or searching for another company.
Bounded public-source set
Criterion contradicted
Required evidence missing
No contact-ready list
The observed surface was an existing local task in the Codex desktop application, with built-in web search, primary-page reads and local Python exports. It was not a Codex CLI prospecting run. The app build, model ID, effective configuration, elapsed time and cost were not retained, so this is not an exact environment replay or a speed comparison.
For a separate future discovery run, record the search mode first. OpenAI documents cached and live local search through web_search; hosted web search and local command networking have separate controls. A search-index date is not the date an event happened. See the web-search documentation and configuration basics. This packet does not establish which setting the original task used.
Our contract required an independent privately held B2B software company, UK or Ireland headquarters, an English product site, and a company-reported headcount of 50 to 1,000 dated within the previous year. It also required an official sales-hiring or commercial-expansion announcement between June 25 and September 23, 2026.
Those dates belong to this historical example. For a new run, set your own criteria and dates in a new folder before searching. A London office is not proof of UK headquarters; a future workforce target is not current headcount. Use research for unresolved facts, reject for an evidenced contradiction, and retain only when every account criterion is supported. Retain still does not identify a recipient or authorize outreach.
Proposed prompt, prepared after the run: the full original discovery prompt was not retained. The instructions below are a reusable recipe, not a transcript or another measured model run. The ZIP also includes an explicitly unexecuted prompt for future discovery.
Work only in this extracted packet folder. Do not search the web,
call APIs, install anything, enrich contacts or send messages.
Read README.md, contract.json, observed-run.json, ledger.json
and source-manifest.json. Preserve unknown values and decisions.
Explain why each account is reject or research.
Run python3 review_packet.py and inspect both files in out/.
Report the counts, structural errors and unresolved facts.
A structural pass is not factual validation or contact approval.
Ask a second review pass to challenge entity scope, event dates and missing ownership evidence. Store correction suggestions separately so you can see what changed. In our recorded review, WrxFlo moved from retain to research because its independent private-company status was not established. Funding alone did not satisfy that criterion.
These are agent-reviewed company decisions. Human review remains pending before operational use.
On a narrow screen, scroll the table sideways. Keyboard users can focus the table area and use the arrow keys.
| Company | Decision | What prevented retain |
|---|---|---|
| WrxFlo | Research | July 7 announcement supports dated size and expansion; independent private-company status remains unknown. |
| Monterro | Reject | Stockholm-based investment firm; a London office does not make it a UK software vendor. |
| Bechtle UK | Reject | Reseller and IT-services group, outside the independent software-vendor criterion. |
| Spectro Cloud | Reject | June 17 announcement precedes the June 25 window. This does not rule out other events. |
| Endra | Reject | Stockholm headquarters. A future 100-person target cannot fill the present headcount field. |
| One Identity | Research | Primary page returned 403. The indexed June 24 date is not a qualifying current event. |
| Vizzia | Research | Official ATS returned empty text. Posting date and company fit remain unresolved. |
Keep an access failure as an access failure. Do not turn the One Identity search snippet into a successfully retrieved source or borrow a later third-party date. The downloaded source manifest records the distinction. These are company-level observations, not a market census or verified buying intent.
With Python 3 available, run this from the extracted folder, either in a terminal or through your local Codex task:
python3 review_packet.py
# Expected counts in out/review-summary.json:
# accounts: 7
# retain: 0, reject: 4, research: 3
The new checker uses only Python's standard library. It checks required fields, duplicate IDs/domains, source references, date formats and evidence requirements for any retained record. It preserves the decisions and blank contact fields, then writes a CSV and summary under out/. We executed this offline utility against the packet and checked duplicate, missing-source, stale-event and unsupported-retain failures with synthetic mutations. Running it twice produced identical outputs without changing the input files.
The checker cannot confirm that a source is true or that a company fits your business. A failed check should lead to a documented correction, not automatic promotion. The included message examples remain HOLD: no recipient was identified and nothing was sent. The deliberately bad message in that file was constructed during editorial review, not produced in a separate Codex run.
When you want to test a real Overloop connection, use the campaign-reporting quickstart for Codex and Claude Code. It reads an existing authenticated campaign and its channel metrics with the pinned Overloop CLI. It does not import these seven accounts or start outreach. A public Max MCP catalog test likewise does not establish native Codex MCP setup.
Related-party disclosure: Overloop publishes this guide, and Max (yourmax.ai) is part of Overloop and is operated by Sortlist SA. The relevant commercial workflow is to review a signal, resolve a suitable prospect and prepare contextual email or LinkedIn outreach. This example stops before contact resolution and campaign actions.
For the companion prospecting playbook, get our existing Claude prospecting guide. The Codex packet is a technical exercise alongside it. The separate client runs used different discovery samples; their counts are not a model-quality, speed or cost benchmark.
An AI GTM agent accepts a commercial objective, gathers relevant context, decides or recommends a bounded next move, uses permitted tools, and returns enough evidence for a person to review what happened. It may research an account, map a buying group, qualify a lead, design a campaign, update a record, or trigger an approved workflow.
The word agent should describe control flow, not copywriting. If the system only fills fields in a fixed template, it is an AI-assisted workflow. If it can choose what information to seek, adapt its path, and decide when the task is complete, it behaves more like an agent. That distinction follows the practical framing in Anthropic's guide to building effective agents: workflows follow predefined paths, while agents dynamically direct their own process and tool use.
An interface that skips the final evidence step may still automate work, but it is difficult to govern or improve.
These categories overlap because vendors describe products by market position, not by a shared technical standard. The safest way to classify a product is to ignore the homepage label and watch what owns the loop.
| Category | Path through the task | Typical output | What to test |
|---|---|---|---|
| Automation | Deterministic rule: when X happens, do Y | Record update or triggered task | Reliability, retries, and edge cases |
| AI workflow | Predefined sequence with one or more model steps | Research summary, score, or draft | Prompt quality and structured output |
| Copilot | Person directs each meaningful step | Suggested copy, answer, or next action | Usefulness and edit burden |
| AI SDR | Owns a bounded sales-development loop | Prospects, outreach, follow-up, or meetings | Target quality and sending controls |
| AI GTM agent | Chooses a path across a broader GTM objective | Decision packet, action, and evidence | Reasoning boundary, permissions, and failure behavior |
An AI SDR can therefore be an AI GTM agent, but the terms are not synonyms. “SDR” tells you the organizational job. “GTM agent” should tell you that the system can reason across a goal. Read our AI SDR versus AI GTM agent guide for the detailed buying distinction.
“Improve pipeline” is not an executable objective. “Find accounts hiring their first European sales team and propose the relevant RevOps buyer” is.
ICP rules, CRM history, approved sources, product facts, exclusions, territory, and previous contact history belong in the context window.
Search, enrichment, CRM reads, drafting, record writes, sequence enrollment, and sending are separate permissions, not one “autonomy” switch.
Suppression lists, territory ownership, prohibited claims, approved channels, minimum evidence, and escalation rules must sit outside the prompt.
The agent needs to know which sources it checked, which hypothesis failed, what a person corrected, and whether the account was already contacted.
A reviewer should see the proposed target, reason, source, confidence, action, and exact changes before permission-sensitive steps run.
Suppose a sales leader asks an agent to identify accounts likely to need a new outbound process. The weak system searches for companies that raised funding and writes a congratulatory email. The credible system treats funding as one observation and tests a fuller hypothesis.
The valuable output is not the email. It is the decision, including the decision to do nothing. If an agent cannot reject an account, respect ownership, or expose conflicting evidence, it is a content generator attached to a database.
Select each connection for the job the agent must complete. A company-search result, a buying signal and a campaign action answer different questions.
Evaluate the connection with your actual client and account permissions. A vendor's documented operation is a starting point for that evaluation.
Open the Overloop developer resources for the campaign-reporting quickstart for Codex and Claude Code. It reads an existing campaign and its statistics through the CLI.
Most production systems combine deterministic software with model-driven decisions. That is a feature. Authentication, suppression, routing, sending limits, and record integrity should remain predictable; research and hypothesis formation can tolerate more flexible reasoning.
“Human review available” is meaningless unless the interface makes review possible before the risky step. The reviewer needs the source, the proposed change, the uncertainty, and a clear approve, edit, reject, or escalate choice.
| Action | Good default | What earns less oversight |
|---|---|---|
| Research public company facts | Run automatically; retain sources and timestamps | Consistently high citation coverage and low entity-matching error |
| Recommend an account or buyer | Human reviews sampled or all recommendations | Stable acceptance and low costly false-positive rates |
| Draft a message | Human reviews product claims and sensitive personalization | Low material edit rate on an approved template set |
| Update a CRM record | Preview field-level changes; preserve history | Reversible writes with monitoring and exception handling |
| Enroll or send | Explicit approval, suppression check, and rate control | Only after measured reliability on the exact segment and motion |
The NIST AI Risk Management Framework is not a sales playbook, but its emphasis on documented risk management, testing, evaluation, verification, and validation is directly useful here. Treat each new permission as a change to the system's risk profile, not as a convenience setting.
A fluent email can hide a bad target. A terse recommendation can be commercially correct. Build a test set from real situations and score the stages independently.
Include obvious fits, deceptive fits, stale signals, conflicting sources, duplicate people, existing opportunities, competitors, and people who must never be contacted. Track acceptance, correction, rejection, citation coverage, false positives, cost per reviewed recommendation, and time saved. Reply rate alone arrives too late and mixes targeting, copy, offer, channel, and deliverability.
Our Claude prospecting exercise shows this review process on real company discovery: zero accounts retained, four rejected and one held for more research.
The agent observes the same inputs as the team and records recommendations, but cannot change records or contact anyone. Compare its decisions with what people actually did.
The agent proposes an account, buyer, reason, and next action. A person approves or rejects every packet and records why.
Allow reversible, low-risk actions such as creating a research note or task. Keep enrollment, sending, deletion, and ownership changes behind approval.
Reduce review only for segments and actions with enough evidence. Keep sampling, alerts, rollback paths, and a kill switch.
Do not advance because the pilot feels impressive. Define the threshold in advance: which errors matter, how many reviewed cases are enough, who owns incidents, and what result sends the system back one stage.
If the team cannot state its ICP, exclusions, owner rules, and approved offers, an agent has no stable policy to execute. Document the motion first.
If every new demo request should create the same CRM task, use automation. Model-driven decision making adds cost and uncertainty without adding value.
An agent cannot repair missing ownership, duplicate accounts, or an obsolete product narrative through confidence. Fix the source of truth.
Pricing changes, contract commitments, large sends, and sensitive record deletion should stay deterministic and explicitly authorized.
If the category is right, compare actual products in our documentation-led guide to the best AI GTM agents. It separates research agents, signal-to-action systems, ABM assistants, and autonomous SDRs instead of pretending they do the same job.
The Codex desktop recipe and downloadable packet were added on September 23, 2026 from the retained run and reviewed files. The setup documentation was checked on that date. No new company search or product call was performed for this addition.
This guide is a category framework, not a benchmark or a claim that one architecture fits every sales team. We reviewed the sources below on Aug. 20, 2026 and separated their general engineering guidance from our GTM-specific interpretation.
Product names do not establish technical architecture. Validate each vendor against your own data, permissions, and edge cases.
An AI GTM agent accepts a go-to-market objective, gathers relevant context, chooses or recommends a bounded action, uses permitted tools, and returns evidence that a person can review.
No. An AI SDR owns a narrower sales-development job such as research, outreach, or follow-up. An AI GTM agent may reason across a broader decision that can include targeting, account research, buying groups, campaign planning, or orchestration.
No. Recommendation mode is often the safest and most useful starting point. An agent can assemble context and propose an action while a person approves the target, message, and permission-sensitive steps.
Test it on representative and adversarial cases, score the evidence and decisions separately from the writing, inspect permissions and failure behavior, and measure human acceptance, correction, and rejection rates.