The finance systems & AI brief for controllers and CFOs
Bespoke consulting for finance tech stack strategy, selection, & implementation
Get a personalized roadmap for your finance tech stack. CFOLAYER partners with CFOs and controllers to design, select, and implement the right systems for your scale and business model - eliminating vendor confusion and slow time-to-value.
🏛️ Systems & Stack
Ramp fully launches in Canada alongside new Toronto office (4 minute read)
Ramp opened its platform broadly to Canadian-headquartered businesses outside Saskatchewan and Quebec, with cards in CAD and USD, CAD bill pay and reimbursements, and automated GST, HST, PST and QST coding. The Toronto office opened in early August with about a dozen hires planned to double within six months, on top of roughly 100 existing Canadian employees. This is Ramp's first market outside the US, arriving with more than 70,000 customers and a $44 billion valuation, and it lands directly on Float at 7,500 Canadian business customers and Venn at more than 15,000. What this means for your stack: if you run a Canadian entity on a US spend platform and your team is hand-coding provincial tax, this is the quarter to re-evaluate, and the incumbents' built-for-Canada positioning is about to be tested rather than assumed. Discuss with ChatGPT →
Ambrook raises $30M to take vertical accounting beyond the farm (8 minute read)
Ambrook grew from 2,500 to more than 8,000 business customers in a year and is pushing out of agriculture into trucking, contracting and real estate, with over 1,000 trucking businesses already on the platform. The wedge was tax and entity complexity rather than price, since farms file differently and frequently run five or six interconnected businesses on one property, which broke QuickBooks. At least half of Ambrook's customers had no accounting software at all before signing up. What this means for your stack: the detail worth noting is that a 650-acre farm in Vermont connected Ambrook's MCP server to Claude and now handles bookkeeping and ledger reconciliation without paying a third-party service, which nobody sold him. When customers wire your MCP server into their own agents unprompted, that is a different product signal than a launch post. Discuss with ChatGPT →
Clear Street opens a pre-IPO platform, starting with Databricks (3 minute read)
Clear Street Private Markets launched with pre-IPO exposure to Databricks, recently valued at $188 billion, with up to 30 more private companies planned by year end, mostly $5 billion to $20 billion names roughly six to 24 months from an IPO. The distinguishing feature is margin lending against private shares, which is rare in this market, alongside dedicated private company equity research. Participation is limited to qualified purchasers and accredited investors. What this means for your stack: read the structure before the pitch, because participants are not buying Databricks shares, they obtain an interest in an SPV that holds an interest in a fund that owns the stock. If pre-IPO exposure turns into a normal treasury or executive comp line item, someone on your team inherits valuation, liquidity classification and disclosure questions for an asset with no daily mark. Discuss with ChatGPT →
🤖 AI in Finance
Sarah Friar set two ambitions when she built OpenAI's finance function: a zero-day close, and a forecast that updates continuously rather than on a calendar. Neither is finished, and what she published is the working method rather than a victory lap. Her fourth lesson is the one for controllers: AI drafts, people sign off, every output connects to a reliable source, every forecast carries an explanation, and every change to an approved baseline requires finance authorization. She also notes that tokenmaxxing came and went, and that usage limits, budget controls, role-based access and approval thresholds are now straightforward to set. What this means for your stack: manage AI spend like any other variable expense and keep the signature with a human who can be asked why. Her own research found 40% of finance professionals' specialized AI use is work outside traditional finance and 22% is engineering-related, which tells you what your team will start doing once you hand them the tools. Discuss with ChatGPT →
Approval fatigue, finally measured (5 minute read)
Anthropic is making auto mode the default in Claude Code from August 14, meaning a classifier reviews each action instead of prompting the user every time, and to justify the change it published the data. In a controlled study with 1,053 paid professional testers, one clearly dangerous command was swapped into an otherwise normal session. Humans caught it 13.6% of the time. The classifier caught 89%. Human performance also decayed inside a single session, from about 17% early on to about 5% after 50 or more prior prompts, while the classifier held flat. Separately, users approve 97% of permission prompts but reject 39% of plans. What this means for your stack: people engage with what asks for judgment and click through what asks for permission, and that is the same shape as your invoice approval queue, your journal entry review and your quarterly access recertification. This is the first published, controlled measurement of how badly click-through review actually performs. Discuss with ChatGPT →

Does AI being good at math mean it is good at your math? (6 minute read)
A research version of Claude improved a longstanding lower bound on the Riemann zeta function from 41.6% to 67.2%, working across two sessions, 31 million output tokens and roughly 60 coordinated subagents. A reader asked whether this transfers to finance. Directly, no, and it is worth being clear about that: arithmetic across large datasets is still a database job, not a model job. What does transfer is the verification method, because the result only counted after subagents hunted for counterexamples, 54 arXiv papers were checked for prior art, the proof was independently re-derived from scratch, two human mathematicians validated it, and a machine-checkable version was produced in Lean. What this means for your stack: roughly a quarter of those 60 subagents existed only to check the others' work. The useful question is not whether the model got it right, it is whether you have built anything capable of independently confirming that it did. Discuss with ChatGPT →
Freehand raises $75M to automate enterprise spending (3 minute read)
Co-led by Battery Ventures and NewRoad Capital Partners, with AI agents already running enterprise spend and finance workflows at Meta, Unilever, Johnson & Johnson, Pfizer, Dunkin' and Cardinal Health. What this means for your stack: the logos are the signal rather than the round size, because enterprise procurement organizations are slow and conservative, and these ones signed. Discuss with ChatGPT →
⚖️ Regulation & Reporting
SEC launches a new enforcement unit aimed at accounting fraud (4 minute read)
The SEC is standing up a Financial Reporting and Accounting unit inside its Division of Enforcement, staffed with both attorneys and accountants and led by Timothy Zimmerman. Enforcement director David Woodcock chaired a similar task force in 2013 and began his career as an auditor at EY. Outside counsel called the move surprising against the agency's deregulatory posture, and Foley Hoag noted the SEC may be filling a void left by thinner PCAOB enforcement and fewer DOJ financial fraud cases. What this means for your stack: Gibson Dunn's Jina Choi had the actionable read: pay attention to internal reports and complaints about financial reporting conduct, because many enforcement actions begin as whistleblower reports and those numbers are trending up this year. Fewer required filings and sharper scrutiny of the ones you do file are not contradictory policies. Discuss with ChatGPT →
France is three weeks out from mandatory B2B e-invoicing (6 minute read)
Mandatory B2B e-invoicing in France goes live September 1, 2026 for large and mid-sized enterprises, with small and micro enterprises following September 1, 2027. Note that PDP has been renamed PA, or Plateforme Agreee, so any guidance you bookmarked last year uses the wrong term. For context on where the rest of Europe sits, Belgium's domestic mandate has been live since January and passed one million registered Peppol receivers within weeks, and Germany's issuing obligation starts January 2027 for companies over 800,000 euros in turnover. What this means for your stack: the obligation to receive compliant e-invoices applies to every business from September 1, 2026 regardless of size, so a small French entity cannot wait until 2027 and the receiving side is what bites first. If you sell into or buy from France and have not connected to an approved platform, that is this month's problem, not next quarter's. Discuss with ChatGPT →
🛠️ The Practitioner
Agentic payments, explained for the person who signs the wire (7 minute read)
A LEDGER original. A reader asked what agentic payments actually means for a CFO worried about agents moving large sums. Here is the answer, anchored on IMF Note 2026/004.
What it means. An agentic payment is a payment initiated by software that was given standing authority to act for you, rather than by a person clicking pay on a specific transaction. Your treasury agent notices a supplier invoice is due, checks it against the PO, picks the cheapest rail and sends it. Nobody approved that specific payment. Somebody approved the agent. Traditional payment law assumes every payment order traces back to an explicit instruction from an account holder, and the entire industry effort right now is about rebuilding that assumption in a different shape.
Nobody is proposing that an AI model move money. This is the part that gets lost in the noise. The IMF's reference architecture keeps the model away from the money entirely. Layer 1 is intent and orchestration, where the agent reasons, compares options and assembles a proposed payment, with zero authority to execute. Layer 2 is control and authorization, a deterministic rules engine, not a model, that checks whether the agent holds a valid signed mandate, whether the counterparty is allowed, whether the amount is within limit, whether the velocity cap is hit, plus sanctions screening and AML filters. Layer 3 is settlement, which executes exactly what Layer 2 approved with no reinterpretation. A model that hallucinates a vendor name in Layer 1 produces a rejected proposal, not a wire.
The mandate is the mechanism. A mandate is a cryptographically signed object stating what an agent may do, in whose name, up to what amount, with whom and for how long. Google and Coinbase's AP2 is the leading standard, with x402 as the stablecoin extension, and OAuth 2.0 and OpenID Connect are being pressed into service for what the industry is calling Know Your Agent. Practically, a mandate is your delegation of authority policy rendered in code and signed. If you have ever written a check-signing authority matrix, you have already written a mandate. The difference is that this one gets enforced automatically and produces an audit trail.
What is actually unresolved. Three things, and no vendor deck will tell you. Liability is ambiguous, because existing regimes assume human intent and direct causation, and when an agent acting under a valid mandate misdirects funds current law struggles to separate unauthorized use from user negligence. Traceability changes shape, because authorization becomes structural rather than transaction-level, so your auditor will ask how you evidence authorization for a payment nobody specifically approved and the honest answer is the mandate, the policy engine log and the agent identity, not a signature. And correlated behavior is a systemic risk nobody has priced, because if every treasury agent optimizes liquidity the same way at the same moment you get synchronized payment flows and strained settlement capacity.
Three questions before an agent touches your money. Where does authorization live? If a vendor cannot show you a deterministic rules layer sitting between the model and the rail, then the model is the control, and you should walk away. What is the mandate, and who can change it? Scope, limits, counterparties, duration and revocation, and then which humans can amend it, and whether amending it requires the same approval as amending your signing authority matrix. It should. What happens when it is wrong? Not if. Who bears the loss, what is the reversal path, and what does the audit trail look like when you have to explain the payment to someone who was not in the room.
Where to start. Do not start with payments. Start with the mandate. Take your existing delegation of authority policy, write it in the form an AP2 mandate would take, and see how much of it you can actually specify. Most finance teams discover their real approval rules are considerably fuzzier than they thought, and that exercise is worth doing whether or not you ever deploy an agent.
What this means for your stack: the scary version is not the version being built, but the liability question is genuinely open and your auditor will get there before your vendor does. Discuss with ChatGPT →
Not all MCPs are created equal (9 minute read, vendor blog)
Two vendors can both tick the MCP support box on your RFP and behave nothing alike. One dumps raw records into the model's context, so 100,000 expenses becomes roughly 20 million tokens and the model quietly answers from whatever fraction it managed to read. A better version loads your data into a temporary database and queries it, which fixes the arithmetic but means minutes of pagination per question. Brex pre-computes and pre-joins continuously and claims 2 to 3 seconds and about 800 tokens for the same 100,000-expense question. Vendor-authored, but the questions at the end are vendor-neutral. What this means for your stack: steal these three verbatim for your next evaluation. What happens when your controller asks about 100,000 expenses? What does the math, the model or a database? How do permissions hold up when the agent can query anything? Any implementation relying on the model itself to do arithmetic is the wrong one. Discuss with ChatGPT →
Documentation rather than news, but there is a segregation-of-duties model buried in it worth reading. Claude Code now lets agents running in separate sessions send each other plain-text messages, and the permission rules are explicit: a message from another agent never counts as user consent, cannot answer a pending approval prompt, and cannot change configuration. Commands embedded in a message arrive as plain text and do not execute. What this means for your stack: somebody had to sit down and decide what one agent is allowed to make another agent do, and this is a worked example of the answer. If you are designing multi-agent finance workflows, that is the control question you will be asked about first. Discuss with ChatGPT →
⚡ Quick Links
Agentic payments will need more than wallets (4 min) A wallet tells you where value sits, a card tells you the network. Neither is authorization.
The future of payments will be defined by interoperability (4 min) The winning network connects to the most others, not the most customers.
Agentic payments protocols compared: MPP, ACP, AP2 and x402 (8 min) Vendor-adjacent, but the cleanest side-by-side of what each protocol actually standardizes.
Moss reaches unicorn status in spend management (3 min) Fintech funding fell to $673M across 15 deals this week, down from $2B across 27.
🔁 ICYMI
Worth a second look
Audit's moment has arrived, and Lightspeed just led $37M into Andera Public company audit fees average $3 million a year and run north of $20 million at the Fortune 100, CFOs are demanding 200% to 300% efficiency gains from back-office functions, and the Big Four cannot build the best AI audit product without cannibalizing billable hours. That last constraint is why they will partner rather than build. Discuss with ChatGPT →
Accounting AI leaders say autonomous agents aren't there yet Resurfacing this deliberately next to this issue's agentic payments coverage, because the people building these tools stay consistently more cautious than the people selling them. Discuss with ChatGPT →
Finance teams spend 13 hours a week verifying AI outputs Thirteen hours a week is the verification tax, and it is the number that decides whether any of the productivity math in this issue actually holds. Discuss with ChatGPT →
📖 Worth The Read
Can agents use a computer yet? a16z went and got the data (15 minute read)
a16z talked to teams actually running computer-use agents in production and came back with numbers instead of projections. Benchmark scores went from 42% to 85% on OSWorld-Verified in about a year against a human baseline near 72%, but the authors argue the benchmark stopped being the story: one operator running millions of automated tasks a month could not tell them which model executes them, because his vendor swaps models underneath him the way a cloud provider swaps hardware. On economics, roughly $6 to $8 per agent hour against about $10 offshore and $30 to $45 for US back-office labor, all fully loaded. What this means for your stack: the failure mode transfers exactly. An agent extracting payment terms from contracts reads net 60 as net 30, the record looks perfectly plausible, passes every visual check, and nobody catches it until an invoice goes out wrong. Their closing frame is the one to carry into any automation business case: the scarce resource is no longer producing the work, it is vouching for it. Discuss with ChatGPT →

How agentic AI will reshape payments (45 minute read, IMF Note 2026/004, dense)
The source behind this issue's Practitioner explainer, and the thing to read instead of a vendor whitepaper. It lays out the three-layer architecture, the full protocol landscape, and a risk classification matrix that names, for each risk, who bears the cost when it materializes. It is also refreshingly willing to say what is unsolved, including that liability regimes assume human intent and that correlated agent behavior in payment initiation could strain settlement capacity in ways nobody has modeled. What this means for your stack: you do not need all 45 minutes. Skim the executive summary, read the Layer 2 section on control and authorization, then read the risk matrix. That is about 20 minutes for most of the value. Discuss with ChatGPT →
📚 Definitions
Terms that turned up in this issue, plus two readers asked for.
RAG, or retrieval-augmented generation. Instead of relying on what a model memorized in training, you first retrieve relevant documents from your own sources, then hand them to the model with the question. It is why a finance assistant can cite your actual revenue recognition policy rather than a generic one. The failure mode is the retrieval step, not the model. If the search surfaces the wrong three documents, you get a confident answer built on the wrong three documents. When a vendor says grounded in your data, ask how they measure whether retrieval found the right thing.
Reinforcement learning, or RL. Training a model by rewarding good outcomes rather than showing it correct answers. The model tries something, gets scored, adjusts, and repeats. Most of the capability jump in agents this past year came from labs spending heavily on RL environments, meaning sandboxes where a model practices real tasks and gets rewarded for finishing them. The improvement was bought deliberately through training infrastructure, which is why you should expect the curve to keep going.
MCP, or Model Context Protocol. The standard way an AI agent connects to a system like your ERP or spend platform. MCP specifies how the connection works, not what sits on the other end of it. Two vendors can both support MCP and perform completely differently, which is the entire point of the Brex piece above.
Agentic payment. A payment initiated by software operating under standing delegated authority, rather than by a person approving that specific transaction.
Mandate. A cryptographically signed object specifying what an agent may do, for whom, up to what limit, with which counterparties and for how long. This is your delegation of authority policy, written in code and signed. If you already have a signing authority matrix, you already have the raw material.
AP2 and x402. AP2, the Agent Payments Protocol, is the leading open standard for binding agent-initiated payments to verifiable mandates, introduced by Google Cloud and Coinbase. x402 is its extension for stablecoin settlement, built on the long-dormant HTTP 402 Payment Required status code.
Know Your Agent. Verifying that a piece of software making a payment request is the software it claims to be, using existing web authorization standards like OAuth 2.0 and OpenID Connect. The agent equivalent of KYC.
Deterministic. A system that produces the same output for the same input every single time, which is what rules engines and settlement rails do. The opposite of probabilistic, which is what a language model is. The whole agentic payments architecture rests on keeping the probabilistic part strictly upstream of the deterministic part. If a vendor blurs that line, that is the finding.
Computer use. An agent operating software the way a person does, through screenshots, clicks and keystrokes, rather than through an API. This is how agents reach the long tail of legacy finance systems that never got a clean integration, which is also where nobody can easily verify the output.
Prompt injection. Hiding instructions inside content an agent reads, such as a web page, a PDF or an invoice, in order to hijack what it does next. Worth naming to your AP team explicitly. A malicious instruction inside an invoice attachment is a fraud vector that did not exist three years ago.
Subagent. A separate agent instance spawned to handle one piece of a larger task, reporting back to a coordinating agent. In the Riemann work above, roughly a quarter of the 60 subagents existed only to check the others' work. That is a control design worth borrowing.