A useful weekly marketing review needs more than a fluent answer. Someone must retrieve the inputs, calculate the numbers, save a report, and explain which conclusions the evidence supports.
We tested that workflow with Pi 1.0 inside Teamday. A live Kimi K3 model read two documents through Teamday's native MCP connection, calculated campaign metrics in code, and saved a Markdown report. The numbers checked out. The integration also revealed missing tool activity in the chat and an overlap bug in how concurrent runs received their tools. We fixed both.
The campaign data was deliberately synthetic. This was a real model and a real production execution path, using controlled inputs rather than customer records. It demonstrates that the workflow runs; it is not a benchmark of every model or a claim about marketing results.
What changed in Pi 1.0
Earendil released Pi 1.0 on October 1, 2026. Pi is an agent harness: the software that runs the model's conversation, exposes tools, and executes the work the model requests. The model supplies the reasoning; the harness supplies the machinery.
The release brings together native MCP support, codemode, deferred tool loading, virtual models, cache warming, and changes to instructions and tools during a conversation. Some arrived in the releases leading up to 1.0. The 1.0 release notes also describe leaner codemode instructions, image generation from scripts, and stronger MCP authentication handling.
For business use, the important question is which of those capabilities turns a request into completed work.
The test: turn two inputs into one review
We created a temporary marketing analyst with access to a dedicated test directory and two Knowledge files:
- Campaign metrics for four channels, with spend, leads, and qualified leads.
- A brief defining the review and prohibiting invented revenue or conversion claims.
The task was to read both files through MCP, fetch the independent inputs together with Promise.allSettled, calculate the metrics, and save a report. It did not authorize changing advertising accounts or contacting anyone.
The model used codemode to discover the Knowledge tool and read the inputs. Its first path attempt returned a file-not-found response; it recovered by supplying the Knowledge directory's working context and relative file paths. It then calculated the results and wrote the report.
Here are the controlled inputs and independently checked outputs:
| Campaign | Spend | Leads | Qualified leads | Cost per qualified lead |
|---|---|---|---|---|
| Search Brand | $1,800 | 60 | 36 | $50.00 |
| Search Nonbrand | $2,400 | 80 | 32 | $75.00 |
| $1,600 | 40 | 16 | $100.00 | |
| Newsletter | $200 | 50 | 30 | $6.67 |
| Total | $6,000 | 230 | 114 | $52.63 |
Overall cost per qualified lead increased from the fixture's previous-week value of $50.00 to $52.63: a 5.3% increase. The saved report included the calculation, campaign comparisons, and embedded table and chart data for Teamday's report viewer.
The paid-only view tells a different story: $5,800 spend and 84 qualified leads give $69.05 per qualified lead, up 15.1% from the $60.00 paid-only baseline. Including the low-cost owned Newsletter makes the combined number look better, so the two summaries should stay separate.
These numbers came from the input documents and code. They were not supplied as an expected answer in the agent's task.
The arithmetic passed; the causal claim needed review
The report correctly identified LinkedIn's highest cost per qualified lead. It also suggested that targeting or creative fatigue might explain weaker paid efficiency.
The fixture did not contain creative history, audience changes, or conversion data. It could establish the metric change, but not its cause. That distinction matters in business work: a correct calculation does not automatically make the explanation correct.
We asked for a revision that expanded the paid-only summary and removed unsupported explanations. The model re-read the campaign input through MCP, read the saved review with the native file tool because it lived outside the Knowledge folder, and rewrote the report. It explicitly stated that the causes could not be determined from these inputs and listed the evidence needed to investigate LinkedIn.
A report can usefully point to an investigation. A budget change needs additional evidence. Both live turns left external campaigns untouched.
Why codemode matters beyond coding
MCP connects Pi to configured tools. A business agent's tools might retrieve campaign statistics, search Knowledge, inspect CRM records, or read support tickets. The connection still needs the appropriate account access and permissions.
Codemode lets the agent combine those calls in a short JavaScript program. It can fetch independent inputs together, calculate ratios, filter rows, and return the useful result rather than sending every intermediate record back through the model.
For example, a marketing analyst could retrieve yesterday's spend and qualified leads, join them by campaign, and return the five largest changes. A support analyst could collect ticket counts and response times, then return the queues that need attention. Those are possible workflows when the required tools and data are connected; our live test covered the Knowledge-to-report path.
Earendil reports that its default GPT-5.6 prompt with codemode shrank from about 5,300 to 3,300 tokens. That is an upstream prompt-size example, not a measured reduction in Teamday's total usage or bill. Task length, retries, returned data, and model choice still affect the result. Source: Pi 1.0 release notes.
What we had to make work in Teamday
Installing a newer CLI was only part of the integration.
Teamday translates the active agent's managed MCP configuration into Pi's native format. It preserves server selection, tool include/exclude lists, headers, environments, and request timeouts. Configured stdio and streamable HTTP servers are supported; Pi does not support legacy SSE transport.
Our overlap test found that writing those servers into a shared provider configuration could give one run another run's tool list. We changed the integration to load each run's servers through a private extension using Pi's native API. Two simultaneous Pi processes now pass the isolation test while sharing login and conversation storage. Removed servers are not carried into the next turn.
We also fixed the parts users and operators depend on:
- Visible work: Pi's native tool execution events, including MCP calls nested inside codemode, now map to Teamday's chat tool timeline.
- Readable replies: text streams into the chat, with repeated terminal copies suppressed so the answer is not duplicated.
- Usage accounting: assistant token usage is counted once per model call, including cache hits, rather than dropped or counted again in repeated end events.
- Failure handling: a provider error fails the run even when Pi's JSON-mode process exits successfully. A successful retry clears the transient failure.
- Conversation continuity: the live follow-up exposed a runner metadata overwrite that incorrectly marked the earlier session as stale. Finalization now preserves the stored metadata required for native session resume.
Our repeatable local CLI test exercises codemode, an MCP tool call, chat activity, streaming, usage, conversation resume, an authentication failure, and concurrent tool isolation. The live marketing test exercises the separate Teamday path with an actual model and persisted output.
The other features, translated into business needs
Large tool collections should not overwhelm every request
Deferred tools can be discovered when needed instead of declaring a company's entire tool catalog up front. Pi's MCP tools can also stay behind codemode's tool search. This is relevant to an employee with several connected apps: the agent needs to find the right operation for the current task. Our integration uses codemode exposure for managed tools. Pi MCP documentation.
Different work can justify different models
Pi's virtual models let an extension choose a physical model for each request. A future business workflow could route routine classification to a smaller model and a difficult review to a stronger one.
That is a direction to evaluate, not a routing system we shipped in this test. Routing also needs quality checks and correct accounting for the models actually used.
Long tool waits can make cache behavior matter
Pi's cache-warming settings can refresh eligible prompt caches during active runs. Refreshes consume usage, so the decision should depend on whether preserving the cache is worthwhile. This may help long workflows with substantial shared context; our Kimi test did not establish a cache-warming benefit.
Creative work can share the same execution loop
Pi 1.0 adds image generation to codemode. That could let a workflow analyze campaign performance and then prepare a new creative concept. A generated image still needs to be saved and made available in the application's asset workflow. We did not test that handoff or ship a new image-generation route here. Pi codemode documentation.
Pi Durable deserves a separate evaluation
Earendil also released Pi Durable as an experimental package. It stores task checkpoints and distinguishes operations that can safely run again after a crash from operations that should be reported as interrupted.
That addresses a serious business requirement. Reading a report again and sending a customer email again have different consequences. A persistent conversation alone does not solve repeated external actions.
Teamday's tested integration uses the Pi 1.0 CLI and our existing runner. We have not replaced that runner with Pi Durable or demonstrated crash-safe replay of external business actions in this test.
Start with a report whose answer you can check
A good first task has known inputs, a saved output, and a clear success criterion: a campaign review with checked arithmetic, a support summary with verifiable counts, or a research brief with cited sources.
In Teamday, choose a Pi-backed model in the chat's model controls and connect the compatible provider credentials. Give the agent the relevant files or connected tools, and name the output you want saved. Connected apps need their own access; selecting Pi does not connect them automatically.
If you are starting with marketing work, our small-business workflow guide describes how to choose a bounded first task. Pi adds another execution option for doing that work inside Teamday.