Campaigns can look clear and convincing to the people who created them while important questions about audience fit, objections and likely action remain unanswered. I wanted a faster way to pressure-test a finished campaign before launch and make the reasoning behind the feedback easy to inspect.

I built Campaign Simulation Lab, a local-first prototype that tests campaigns with a synthetic audience of AI agents. A marketing manager enters the campaign, defines the intended audience, market, buying stage, channel and objective, and starts the simulation.

The starting panel contains ten respondent agents. Each receives its own role within the target audience and approaches the campaign from that perspective.

The tool combines a headline score with each agent’s reaction, likely action, objections and assessment of relevance, clarity and credibility. It then turns the findings into a readable report. Each run retains the campaign brief, audience assumptions, agent instructions, model details and individual responses, making it possible to inspect how the result was formed.

Separating audience construction from campaign evaluation

The system uses agents in two distinct roles.

First, an audience-building agent creates the panel from the manager’s brief. It defines the perspectives needed to represent different parts of the intended audience.

Separate respondent agents then evaluate the campaign from those perspectives.

I deliberately kept the audience-building agent blind to the campaign offer, copy and creative. Giving it access to the campaign could allow it to quietly reshape the audience to make the offer appear more relevant than it really was.

This separation was added after early testing exposed audience drift. Some respondent agents behaved like marketing professionals regardless of the audience entered. Keeping audience construction independent from campaign evaluation helped address that drift and gave every agent a clearer role.

Fixing the scoring logic

Testing also uncovered a more serious problem in the scoring.

In a deliberately mismatched example, a campaign about tractor tyres was shown to a panel of agents representing marketing managers. The first scoring method returned 36. Clear presentation had partly compensated for the fact that the offer was irrelevant to the audience, overstating the campaign’s apparent fit.

I rebuilt the scoring so that relevance and likely action became hard gates. The same agent panel and campaign mismatch then scored 10.

That test confirmed the correction of a specific scoring problem. Predictive accuracy against real customers will require separate validation, but the revised system became much better at recognising an obvious failure of audience fit.

Local execution and usable progress

The agents can run through a local AI model, keeping campaign data and results in the local environment during those runs. Local execution also avoids per-run model charges, although it still uses the machine’s processing time. An external model can be connected when a comparison is useful.

Across 14 recorded development runs, a ten-agent local panel completed in just over two minutes on average. The timing varied depending on the campaign and the level of detail requested.

Longer agent simulations also need to be understandable to the user. I added progress tracking, estimated completion time, completed-agent counts, clear diagnostics and reconnection. A temporary browser interruption would therefore not automatically create a duplicate test.

Controlled preview access allowed the prototype to be tested externally, while PDF export made the agents’ findings easier to review and share.

The practical value

The current prototype provides directional evidence about a written campaign’s fit with its intended audience. The agents can expose obvious mismatches, surface likely objections and show which assumptions deserve real customer validation before budget is committed.

I treat the results from the agent panel as one structured input to the wider campaign decision. Real customer research and live campaign performance still provide the final validation. The current version also evaluates the written campaign brief without genuinely interpreting uploaded visuals, making visual analysis part of the next development stage.

The tool already creates value by giving teams a quick, consistent and documented first review before launch. Instead of receiving one general AI answer, the manager can inspect how several independently instructed agents react to the same campaign.

How much more value can it create?

Its value can grow through reusable agent panels, larger simulations, genuine visual analysis and calibration against human research and live campaign results.

Feeding those outcomes back into the system could create historical benchmarks, sharpen the agent scoring and show where synthetic feedback is most useful. Over time, the lab can become a more calibrated input to campaign decisions and help teams direct human research towards the questions that matter most.