Skip to content

Finance report agent ​

This page explains how a finance report agent, under delegated user authority, extracts report data from authorized finance or billing backends, validates every figure structurally on the app side (period, currency, source, capture time), writes it to a versioned report DB, and produces a periodic report where every figure traces back to a source page. After reading it, you will understand which permission boundaries read-only financial data needs, why the delegated scope contains no transaction actions, and why the reporting basis and baselines belong in a report DB rather than Memory.

Use case ​

Analysis teams need to periodically pull data from several finance or billing backends (SaaS billing, payment platforms, bank statement pages) into a periodic report, supplemented with public filings and announcements for context. These backends have no unified export API; signing in to each one and copying figures by hand is slow, error-prone, and leaves no record of which page a number came from.

Typical triggers:

  • Before a monthly or quarterly business report is due, cost and revenue figures must be aggregated across billing backends.
  • A tracked company releases a filing or major announcement, and the analysis summary must be updated.
  • An audit or budget review requires source-page evidence for every figure in the report.

Engineering challenges ​

  • Basis consistency: the same metric can differ across backends in period, currency, and accounting basis; a report with basis drift is not comparable across periods — "up 12% QoQ" may just mean the basis changed.
  • Figure traceability: every figure in the report must survive the question "where did this number come from." One miscopied figure, or one figure taken from the wrong page version, and the whole report loses its footing in an audit.
  • Backend heterogeneity and login state: every backend differs in report structure, sign-in flow, and MFA policy; extraction must fail loudly when a page structure changes instead of silently emitting wrong numbers.

Module composition ​

ModuleRoleNotes
GenAuthCoreFinance backend login state is high-risk delegation: interactive authorization, read-only report pages, short-lived revocable credentials, and a delegated scope that excludes all transaction capability at issuance.
Web AgentCoreControlled sessions extract report figures page by page, keeping a source URL, screenshot, and capture time per figure; Profiles reuses login state, and public filings and announcements are supplemented with WebSearch.
GUMemNot usedFigures, the reporting basis, and period baselines are business data — they go to a versioned report DB for cross-period reconciliation and audit. Memory plays no part in this scenario.

Finance report agent architecture

Permission and delegation boundaries ​

The Agent holds no inherent permissions. The effective authority for each data-pull task is the intersection of three sets: what the user actually holds ∩ what was explicitly delegated ∩ what the enterprise has approved. Applied to this scenario:

  • The delegated scope covers only "read report pages in the specified finance or billing backends" — no transactions, payments, refunds, approvals, or account settings of any kind.
  • Finance backend sign-in is a high-risk operation and must use mode: 'interactive': the user approves in Qoni Console, and the credential is exchanged in a server-side callback — it never reaches the browser.
  • Delegation credentials are short-lived; minute-level validity is recommended for a single data-pull task, and periodic reports rely on scheduled re-issuance.
  • Out-of-scope attempts (for example, a payment or approval page) are rejected and recorded — the audit chain covers all attempts, not just successful actions.

Note: this sample requests doAnything product permissions. Your application and downstream services must configure and enforce business limits such as backend domains and report-page ranges. Product scopes and prompt rules do not enforce those detailed limits; this sample does not configure them. See Delegate token and attenuation.

Workflow ​

Finance report agent workflow

  1. The user selects the reporting period and the list of finance backends to pull from.

  2. Your app starts interactive delegation; the user approves in Qoni Console, and the server-side callback exchanges the credential.

  3. Your app loads the current version of the reporting basis (currency, metric definitions, period rules) from the report DB and injects it into the task.

  4. Web Agent opens each finance backend; on first access the user completes sign-in in the controlled session, and later cycles reuse login state through Profiles.

    Checkpoint: When a login wall, CAPTCHA, or risk-control page appears, the Web Agent should escalate to a human instead of silently bypassing it.

  5. Web Agent extracts report figures page by page, recording each figure's source URL, page screenshot, and capture time; public filings and announcements are added as needed.

  6. Your app validates each figure structurally: if any of period, currency, sourceUrl, or capturedAt is missing, the figure is marked pending_verification, never presented as confirmed data.

    Checkpoint: Every figure in the report should trace back to a concrete source page; untraceable figures may only appear in pending-verification state.

  7. Validated figures are written to the report DB together with the basis version; your app produces the periodic report with a figure-to-source table, the pending-verification list, and the audit id.

Example code ​

This example uses @qoniai/qoni 0.9.0, published on npm. The download includes the same SDK version, installed with npm ci. The ledger is synthetic: USD revenue 12000, cost 7000, and period demo-2026-Q3. The Mozilla Manifesto is used only to verify the organization. The SDK reads real public pages; business inputs in scenarios.ts are labeled public-demo.

This demo does not read or write GUMem; its task uses the supplied page list and explicit business inputs. DoAnything opens the supplied pages and produces the scenario output. Supply your own JSON with --input; use --interactive when the user must consent to delegation. For site sign-in and user responses, see Qoni SDK.

Download the complete runnable examples, or run from the documentation repository:

bash
cd examples/qoni
npm ci
npm run case -- finance-report-agent
# Supply your own inputs
npm run case -- finance-report-agent --input /path/to/input.json

Set server-side QONI_ACCESS_KEY and QONI_SECRET_KEY. QONI_USER_ID can identify your application's current GenAuth user; the local demo otherwise selects a user from the bound pool. Demonstration Memory writes use an isolated user rather than changing a business user's preferences.

This scenario's executable entry point:

ts
import { cliOptions } from '../runtime.js'
import { inputFile, runScenario } from '../run-case.js'

// Draft a demonstration report from the supplied synthetic reporting basis.
const report = await runScenario('finance-report-agent', cliOptions(), inputFile())
// report.items: metrics, values, periods, currencies, sources, and timestamps for reconciliation.
// Validated fields: metric, value, period, currency, sourceUrl, capturedAt.
console.log(JSON.stringify(report, null, 2))

The entry point loads the definition below by scenario ID. The code is included directly from scenarios.ts, with comments shown in the page language: the task, output fields, source field, Web Search queries (if any), whether Memory is used, and the demonstration input. The pipeline appends the input data, search sources, recalled Memory, and shared safety constraints to the task to build the final prompt; see the pipeline below for the full assembly.

ts
// Draft a demonstration report from the supplied synthetic reporting basis.
browser('finance-report-agent',
  'Use the explicit synthetic ledger and reporting basis to make a demonstration period report. Verify the organization identity from the public source page. Return [{metric,value,period,currency,sourceUrl,capturedAt}] with one item per synthetic ledger metric, copying the metric name and value verbatim from the ledger; every item uses the input period and currency, and its sourceUrl is the public page that verified the organization identity. Clearly keep ledger figures as synthetic input data; do not describe them as Mozilla financial facts. Do not transact, approve, refund or change any account.',
  ['metric','value','period','currency','sourceUrl','capturedAt'], 'sourceUrl', sample([manifesto], { period:'demo-2026-Q3',currency:'USD',ledger:[{metric:'demo_revenue',value:12000},{metric:'demo_cost',value:7000}],reportingBasis:'Synthetic cash-basis demonstration ledger, not an actual company filing.' })),

The sample implements runScenario(), browser(), research(), and sample() as application functions. The pipeline below makes the actual SDK calls: delegation and introspection → required Memory/search → browser task or monitor → validation and saving. fields and sourceField define the application's output checks. The complete application helpers are in the package's runtime.ts.

Inspect the actual SDK pipeline
ts
// browser() uses DoAnything; research() searches first; sample() labels public-demo inputs.
const firefox = 'https://www.mozilla.org/en-US/firefox/new/'
const manifesto = 'https://www.mozilla.org/en-US/about/manifesto/'
const privacy = 'https://www.mozilla.org/en-US/privacy/firefox/'
const support = 'https://support.mozilla.org/en-US/kb/get-started-firefox-overview-main-features'
// These copy rules become prompt context; server permissions and business checks remain separate.
const policy = {
  version: 'demo-2026-09',
  approvedClaims: ['Describe only features supported by the cited page.'],
  forbiddenClaims: ['guaranteed security', '100% private', 'unverified pricing or performance'],
  voice: 'concise and warm',
}
const sample = (pages: string[], business: JsonObject = {}): JsonObject => ({
  dataset: 'public-demo', pages, policy, business,
  notice: 'Business records are synthetic demonstration inputs. Referenced websites and SDK execution are real.',
})
// fields lists required output keys; sourceField identifies URL checks; memory enables GUMem calls.
const browser = (id: string, task: string, fields: string[], sourceField: string | undefined, input: JsonObject, memory = false): Scenario => ({
  id, products: ['doAnything'], task, fields, sourceField, input, memory,
})
const research = (id: string, task: string, fields: string[], sourceField: string, queries: string[], input: JsonObject, memory = false): Scenario => ({
  id, products: ['webSearch', 'doAnything'], task, fields, sourceField, queries, input, memory,
})
ts
import { QoniScopes, type JsonObject, type RunResult } from '@qoniai/qoni'
import { readFileSync } from 'node:fs'
import { getScenario, type Scenario } from './scenarios.js'
import { appendTrace, checkInputCoverage, cleanupDemoUser, createContext, delegate, handleInteraction, inputEntryCount, isolateDemoUser, object, readWithRetry, renderScreenshot,
  save, saveArtifacts, searchHits, settled, settleRun, validateItems, withCleanup, type Context, type Options } from './runtime.js'

export async function runScenario(id: string, options: Options = {}, input?: JsonObject) {
  // Load the task definition by ID; --input replaces its business inputs.
  const scenario = getScenario(id)
  const data = input ?? scenario.input
  if (scenario.memory && data.dataset !== 'public-demo' && options.mode !== 'interactive' && !options.userId && !process.env.QONI_USER_ID) {
    throw new Error('Business Memory writes require the current QONI_USER_ID; do not select an arbitrary bound user')
  }
  const context = await createContext(id, { ...options,
    skipUserResolution: scenario.memory && data.dataset === 'public-demo' })
  return withCleanup(context, async register => {
    register('isolated demonstration user', () => cleanupDemoUser(context))
    // Isolate demo preferences; business Memory belongs to the identified current user.
    if (scenario.memory && data.dataset === 'public-demo') await isolateDemoUser(context)
    return await executeScenario(context, scenario, data)
  })
}

export async function executeScenario(context: Context, scenario: Scenario, input: JsonObject) {
  const pages = input.pages
  if (!Array.isArray(pages) || !pages.length || pages.some(page => typeof page !== 'string' || !/^https:\/\//.test(page))) {
    throw new Error('Input pages must be an array of HTTPS URLs')
  }
  if (input.requiresLogin === true && context.mode !== 'interactive') {
    throw new Error('Targets that require sign-in need --interactive and user-controlled login')
  }
  const memoryScopes = scenario.memory
    ? [QoniScopes.GUMEM_MEMORY_READ, QoniScopes.GUMEM_MEMORY_WRITE, QoniScopes.GUMEM_MESSAGE_WRITE] : []
  // delegate() is an application helper around the SDK delegation methods.
  const grant = await delegate(context, scenario.id, scenario.products, memoryScopes)
  // Read the effective scopes; readWithRetry() retries only retryable read failures.
  const { data: tokenInfo } = await readWithRetry(context, 'delegation introspection',
    () => context.qoni.genauth.introspectDelegationToken({ token: grant.token }))
  const info = object(tokenInfo)
  if (info.active !== true) throw new Error('The delegation token is not active')
  const audit = { grantId: grant.grantId, auditId: grant.auditId, scopes: info.scope }
  let memory: unknown
  if (scenario.memory) {
    // A Session associates this conversation with the user; the app chooses sessionId.
    const sessionId = `${scenario.id}-${Date.now()}`
    await context.qoni.gumem.createSession({
      token: grant.token, userId: context.userId, sessionId, title: scenario.id,
    })
    const preferences = object(input.business ?? {}).confirmedPreferences
    if (Array.isArray(preferences) && preferences.length) {
      // Store confirmed preferences only; sync: true requests synchronous processing.
      await context.qoni.gumem.addMessages({ token: grant.token, userId: context.userId, sessionId, sync: true,
        messages: [{ role: 'user', content: `Confirmed demonstration preferences: ${preferences.join('; ')}` }] })
    }
    // Recall relevant preferences for the later task prompt.
    memory = (await readWithRetry(context, 'GUMem recall', () => context.qoni.gumem.recall({ token: grant.token, sessionId,
      query: 'Confirmed preferences relevant to this task', details: true }))).data
    save(context, 'memory.json', { sessionId, context: memory })
    if (Array.isArray(preferences) && preferences.length && !preferences.every(value => JSON.stringify(memory).includes(String(value)))) {
      throw new Error('Recall did not include the confirmed preferences just written by this demo')
    }
  }

  if (scenario.products.includes('track')) return runMonitor(context, scenario, input, grant.token, audit)

  let hits: ReturnType<typeof searchHits> = []
  if (scenario.queries) {
    // Web Search returns results[]; DoAnything receives these sources to inspect.
    const search = await context.qoni.webSearch.run({ token: grant.token, prompt: scenario.queries, maxResultsPerQuery: 3 })
    save(context, 'search-ref.json', { runId: search.id, audit })
    await withCleanup(context, async register => {
      register('Web Search run', () => search.cancel('Documentation demonstration cleanup'))
      const result = await settleRun(context, search)
      save(context, 'search-result.json', result)
      settled(result)
      hits = searchHits(result.output)
    })
  }

  // One-per-entry scenarios request exactly one item per input entry; others at most two.
  const requiredItems = inputEntryCount(scenario.id, input)
  // Assemble the task, inputs, search sources, and Memory as application-defined context.
  const prompt = [scenario.task, `Task inputs: ${JSON.stringify(input)}`,
    `Search sources: ${JSON.stringify(hits)}`, `Confirmed memory: ${JSON.stringify(memory ?? null)}`,
    `Actual collection time: ${new Date().toISOString()}`,
    requiredItems === undefined
      ? 'Inspect the supplied sources. Return at most two items in the requested JSON array, without prose or Markdown.'
      : `Inspect the supplied sources. Return exactly ${requiredItems} item${requiredItems === 1 ? '' : 's'} in the requested JSON array, one per input entry, without prose or Markdown.`,
    'Keep synthetic demonstration data identified as synthetic. Do not send messages, publish, pay, edit accounts or submit forms.',
    input.requiresLogin === true ? 'Request user sign-in through an interaction when required; never enter credentials yourself.' : 'Public demonstration sources only; do not sign in.',
  ].join('\n\n')
  // Start the Agent with this grant; capture receives delivered screenshots, not every step.
  const run = await context.qoni.doAnything.run({ token: grant.token, prompt, capture: { screenshots: true } })
  save(context, 'run-ref.json', { runId: run.id, session: run.sessionRef, audit })
  let result: RunResult
  const trace = (event: { type: string; data: unknown }) => {
    context.eventCounts[event.type] = (context.eventCounts[event.type] ?? 0) + 1
    if (['progress','message','done'].includes(event.type)) appendTrace(context, event)
    if (event.type === 'browserLiveUrlChanged') {
      const liveUrl = object(event.data).liveUrl
      if (typeof liveUrl === 'string') context.browserUrl = liveUrl
    }
  }
  return withCleanup(context, async register => {
    register('DoAnything run', () => run.cancel('Documentation demonstration cleanup'))
    if (context.delivery === 'events') {
      // --events streams updates; wrap interaction data in an SDK handle for user handling.
      for await (const event of run.events({ signal: AbortSignal.any([context.abort.signal, AbortSignal.timeout(context.timeoutMs)]) })) {
        trace(event)
        if (event.type === 'screenshot') renderScreenshot(context, event.image)
        if (event.type === 'interaction') await handleInteraction(context, run.interactionHandle(event.data))
      }
      result = await settleRun(context, run)
    } else {
      // Callback mode receives this run's events inside wait; helpers save images and ask the user.
      result = await settleRun(context, run, { onEvent: trace,
        onScreenshot: (image, index) => renderScreenshot(context, image, index),
        onInteraction: interaction => handleInteraction(context, interaction) })
    }
    save(context, 'result.json', result)
    settled(result)
    await saveArtifacts(context, result)
    if (scenario.id === 'landing-page-audit-agent' && context.screenshots === 0) throw new Error('The landing-page audit did not deliver the requested screenshot')
    // The app checks required fields and source URL formats; a reviewer still checks facts.
    const items = validateItems(result.output, scenario.fields, scenario.sourceField)
    // Scenarios that require one item per input entry are checked against the input.
    checkInputCoverage(scenario.id, items, input)
    const report = { scenario: scenario.id, dataset: input.dataset, passed: true, runId: run.id,
      status: result.status, items, audit, screenshots: context.screenshots, interactions: context.interactions,
      events: context.eventCounts, artifactIds: result.artifacts.map(artifact => artifact.id) }
    save(context, 'report.json', report)
    return report
  })
}

async function runMonitor(context: Context, scenario: Scenario, input: JsonObject, token: string, audit: JsonObject) {
  // Track creates a monitor with targets, extraction fields, and hourly scheduling.
  const monitor = await context.qoni.track.create({ token, prompt: scenario.task,
    targetUrls: input.pages, extractionSchema: { heading: 'string', source_url: 'string' },
    tickInstructions: `Open the target URLs and read the actual visible heading. Return a JSON object with heading and source_url. ${scenario.task}`,
    triggerDsl: { on: 'change' }, schedule: { kind: 'interval', intervalSeconds: 3600 } })
  save(context, 'monitor-ref.json', { id: monitor.id, audit })
  return withCleanup(context, async register => {
    register('Track monitor', () => monitor.delete())
    const definition = await monitor.get()
    save(context, 'monitor-definition.json', definition)
    if (object(definition.schedule).intervalSeconds !== 3600) throw new Error('Track did not persist the requested schedule interval')
    // Run one tick and inspect its extraction by runId; completed alone does not prove success.
    const tick = await monitor.runNow()
    save(context, 'tick.json', tick)
    if (tick.state !== 'completed') throw new Error(`Track execution failed: ${tick.state} / ${tick.error ?? ''}`)
    const runId = tick.runId
    if (typeof runId !== 'string') throw new Error('Track tick did not return a runId')
    const detail = await monitor.run(runId)
    save(context, 'tick-detail.json', detail)
    if (detail.state !== 'completed' || !detail.extracted || !Object.keys(object(detail.extracted)).length) {
      throw new Error('Track did not extract page data')
    }
    const extracted = object(detail.extracted)
    if (typeof extracted.heading !== 'string' || !extracted.heading.trim() ||
      typeof extracted.source_url !== 'string' || !/^https:\/\//.test(extracted.source_url)) {
      throw new Error('Track extraction is missing a heading or source URL')
    }
    const normalizeUrl = (value: string) => { const url = new URL(value); url.hash = ''; return url.href.replace(/\/$/, '') }
    if (!(input.pages as string[]).some(url => normalizeUrl(url) === normalizeUrl(String(extracted.source_url)))) {
      throw new Error('The Track source URL is not a configured target')
    }
    // Check persisted pause/resume state; withCleanup() deletes the monitor on exit.
    await monitor.pause()
    if ((await monitor.get()).status !== 'paused') throw new Error('Track did not persist the paused state')
    await monitor.resume()
    if ((await monitor.get()).status !== 'active') throw new Error('Track did not persist the active state')
    const report = { scenario: scenario.id, dataset: input.dataset, passed: true, monitorId: monitor.id,
      runId, state: detail.state, outcome: detail.outcome, extracted: detail.extracted, audit }
    save(context, 'report.json', report)
    return report
  })
}

export function inputFile(): JsonObject | undefined {
  const index = process.argv.indexOf('--input')
  return index >= 0 ? object(JSON.parse(readFileSync(process.argv[index + 1], 'utf8'))) : undefined
}

Results are written to output/finance-report-agent/report.json. report.items contains metrics, values, periods, currencies, sources, and timestamps for reconciliation, with fields metric, value, period, currency, sourceUrl, capturedAt. audit links the grant ID, audit ID, and effective scopes; http.json records redacted request statuses. The application parses DoAnything output and checks required fields and source URL formats. A business reviewer still assesses the content against the original sources.

Data and memory boundaries ​

This scenario touches four kinds of data; none of them belongs in GUMem:

  • Versioned rules: the reporting basis (currency, metric definitions, period rules) — managed by version in the report DB, referenced by version in every report; a basis change is a version bump.
  • Business state: report figures, figure-to-source mappings, page screenshots, and period baselines — kept in the report DB for cross-period reconciliation, audit, and difference explanation.
  • Audit records: the behavior chain of interactive authorization, every pull, and rejected out-of-scope attempts — maintained by GenAuth.
  • User Memory: this scenario does not use GUMem. Figures and the reporting basis are business data that must reconcile across periods, not user preferences; putting them in Memory loses version reconciliation.

Failure handling ​

SituationRecommended handling
Finance backend login state expiresSuspend the task, notify the user to sign in again, and resume from the checkpoint.
Report page structure changes break extractionTreat it as a failure and replay the session recording; mark the data point as missing instead of filling in an estimate.
Payment or approval page request outside the delegated scopeReject and record it; the attempted access remains visible in the audit chain.
A figure lacks period, currency, source, or capture timeApp-side validation marks it pending_verification and puts it on the pending list for human review before confirmation.

Production notes ​

Financial content should preserve sources and dates. Agent output should not be treated as investment advice. The Agent performs no transactions, payments, refunds, or approvals — the delegated scope excludes these capabilities at issuance. When backend figures conflict with public filings, present both sources and the difference side by side as pending verification instead of picking a side; the basis version history in the report DB is the only ground for explaining cross-period differences.

Next steps ​