Vendor monitor agent
This page explains how a vendor monitor agent uses Track to keep a long-lived watch on vendor pricing pages, SLA and terms pages, and status pages: the first run builds a baseline, later runs compare against it on a schedule, and hits notify the owner with a diff and sources. After reading it, you will understand why this scenario's backbone is a long-lived monitor rather than memory, where baselines and per-run diffs should live, and why watching public pages only needs a silently issued runtime credential.
Use case
Procurement, legal, or security teams need to monitor vendor pricing, terms of service, security announcements, and product changes. Vendors rarely announce these updates: pricing pages shift quietly, SLA terms change during a redesign, and status pages record incidents that never reach email. Manual patrols consume time and cannot prove "what the terms looked like at the last check."
Typical triggers:
- Before a contract renewal, the team must confirm how the vendor's SLA and terms changed since signing.
- An incident appears on a vendor status page, and its impact on the business must be assessed.
- An annual procurement review needs a change timeline of each vendor's pricing page as negotiation input.
Engineering challenges
- Signal-to-noise in change detection: vendor pages get redesigned often, and visual or structural churn far outnumbers substantive change. Misreporting "the page was redesigned" as "the terms changed" trains owners to ignore alerts within a few rounds.
- Baseline management: deciding "did it change" requires knowing "what it looked like last time." Baselines must be versioned, replayable, and rebuildable after a redesign with human confirmation — snapshots scattered on a script's local disk cannot support evidence at renewal negotiations.
- Alert credibility: an alert like "SLA dropped from 99.9% to 99.5%" must point back to the exact page, the section-level diff, and the capture time; a change verdict without a source has no actionable value.
Module composition
| Module | Role | Notes |
|---|---|---|
| GenAuth | Core | Silent delegation issues the short-lived runtime credential every product call requires; interactive consent is only needed when login state or write actions enter the picture — see "When interactive consent is needed" below. |
| Web Agent | Core | Track maintains the long-lived monitor: the first run builds a baseline, scheduled runs re-fetch and compare, and hits notify via callback; a one-off check can fall back to a single-round patrol. |
| GUMem | Not used | Page baselines and per-run diffs are business state, not user memory — they live in Track's run snapshots and your baseline store, managed by version. Memory plays no part in this scenario. |
When interactive consent is needed
Every product call requires a GenAuth delegate token; public read-only scenarios are covered by silent delegation. This scenario watches public pages by default (pricing pages, public terms pages, status pages) and needs no interactive confirmation in the Console:
- Public reads: issue a runtime credential silently with
delegateToken. The credential is short-lived and revocable at any time; after revocation the monitor's next fetch fails immediately, and every issuance and fetch enters the audit chain. - Signed-in targets or write actions: only when the monitoring target is a vendor's signed-in backend (for example, contract pricing or billing pages), or any write action on a site is involved, switch to
mode: 'interactive'so the owner approves in Qoni Console. The example on this page includes no such target.
Note: this sample requests track product permissions. Your application and downstream services must configure and enforce business limits such as vendor domain lists and page ranges. Product scopes and prompt rules do not enforce those detailed limits; this sample does not configure them. See Delegate token and attenuation.
Workflow
The team picks the vendor pages to watch, the sections that matter, and the check frequency.
Your app silently issues a runtime credential and loads the current version of the monitor policy from your policy store.
Your app creates the Track monitor; the first run fetches the target pages and builds the baseline (
baseline_extracted).Every scheduled run afterwards fetches the current pages and compares against the baseline; whether something changed is decided by rules, not by model wording.
Checkpoint: When a full-page redesign breaks the comparison, treat it as a failure — replay that run and rebuild the baseline after human confirmation, instead of misreporting the redesign as a terms change.
On a hit, Track calls back to your app with the section-level diff, source URLs, and snapshots.
Your app validates the change entries (no source, no alert), writes the new baseline and diff to the baseline store, and notifies the owner.
Checkpoint: Every alert should trace back to a concrete section diff and source page; risk grading and follow-up stay with the owner — the Agent never issues high-risk verdicts on its own.
Example code
Legacy Track sample compatibility
This case retains the published npm @qoniai/qoni@0.9.0 sample code. Its legacy /track/monitors route returns 404 against the current /track/tracks backend, so this case will not pass there. Callbacks, snapshots, and profile_id describe the old monitor contract and are not current Track capabilities. For the current HTTP contract and unpublished Unreleased 0.10.0 SDK candidate, see Track.
Check extraction results
Record success only when the monitor state and extracted data match the expected output. If the state is completed but extracted is empty, the sample fails. Check the target page, extraction fields, and execution state.
This example uses @qoniai/qoni 0.9.0, published on npm. The download includes the same SDK version, installed with npm ci. The demo monitors the Firefox public product-page heading under a synthetic demo-v1 policy. The SDK reads real public pages; business inputs in scenarios.ts are labeled public-demo.
The sample creates a Track monitor, runs one extraction, checks the result and pause/resume state, then deletes the monitor. This sample does not cover long-running scheduling, change notifications, or private-portal sign-in persistence. Supply your own JSON with --input; use --interactive when the user must consent to delegation. For site sign-in and user responses, see Qoni SDK.
Download the complete runnable examples, or run from the documentation repository:
cd examples/qoni
npm ci
npm run case -- vendor-monitor-agent
# Supply your own inputs
npm run case -- vendor-monitor-agent --input /path/to/input.jsonSet server-side QONI_ACCESS_KEY and QONI_SECRET_KEY. QONI_USER_ID can identify your application's current GenAuth user; the local demo otherwise selects a user from the bound pool. Demonstration Memory writes use an isolated user rather than changing a business user's preferences.
This scenario's executable entry point:
import { cliOptions } from '../runtime.js'
import { inputFile, runScenario } from '../run-case.js'
// Create a public-page monitor and check extraction, pause, and resume.
const report = await runScenario('vendor-monitor-agent', cliOptions(), inputFile())
// report.extracted: monitor and run IDs, headings, and sources; the demo deletes its monitor on exit.
// Validated fields: heading, source_url.
console.log(JSON.stringify(report, null, 2))The entry point loads the definition below by scenario ID. The code is included directly from scenarios.ts, with comments shown in the page language: the task, output fields, source field, Web Search queries (if any), whether Memory is used, and the demonstration input. The pipeline appends the input data, search sources, recalled Memory, and shared safety constraints to the task to build the final prompt; see the pipeline below for the full assembly.
// Create a public-page monitor and check extraction, pause, and resume.
{ id:'vendor-monitor-agent', products:['track'], task:'Monitor the supplied public vendor page for heading changes. Extract heading and source_url. Report evidence and changes without placing orders, submitting forms or changing any vendor account.', fields:[], input:sample([firefox], { policyVersion:'demo-v1',watchedSections:['heading'] }) },The sample implements runScenario(), browser(), research(), and sample() as application functions. The pipeline below makes the actual SDK calls: delegation and introspection → required Memory/search → browser task or monitor → validation and saving. fields and sourceField define the application's output checks. The complete application helpers are in the package's runtime.ts.
Inspect the actual SDK pipeline
// browser() uses DoAnything; research() searches first; sample() labels public-demo inputs.
const firefox = 'https://www.mozilla.org/en-US/firefox/new/'
const manifesto = 'https://www.mozilla.org/en-US/about/manifesto/'
const privacy = 'https://www.mozilla.org/en-US/privacy/firefox/'
const support = 'https://support.mozilla.org/en-US/kb/get-started-firefox-overview-main-features'
// These copy rules become prompt context; server permissions and business checks remain separate.
const policy = {
version: 'demo-2026-09',
approvedClaims: ['Describe only features supported by the cited page.'],
forbiddenClaims: ['guaranteed security', '100% private', 'unverified pricing or performance'],
voice: 'concise and warm',
}
const sample = (pages: string[], business: JsonObject = {}): JsonObject => ({
dataset: 'public-demo', pages, policy, business,
notice: 'Business records are synthetic demonstration inputs. Referenced websites and SDK execution are real.',
})
// fields lists required output keys; sourceField identifies URL checks; memory enables GUMem calls.
const browser = (id: string, task: string, fields: string[], sourceField: string | undefined, input: JsonObject, memory = false): Scenario => ({
id, products: ['doAnything'], task, fields, sourceField, input, memory,
})
const research = (id: string, task: string, fields: string[], sourceField: string, queries: string[], input: JsonObject, memory = false): Scenario => ({
id, products: ['webSearch', 'doAnything'], task, fields, sourceField, queries, input, memory,
})import { QoniScopes, type JsonObject, type RunResult } from '@qoniai/qoni'
import { readFileSync } from 'node:fs'
import { getScenario, type Scenario } from './scenarios.js'
import { appendTrace, checkInputCoverage, cleanupDemoUser, createContext, delegate, handleInteraction, inputEntryCount, isolateDemoUser, object, readWithRetry, renderScreenshot,
save, saveArtifacts, searchHits, settled, settleRun, validateItems, withCleanup, type Context, type Options } from './runtime.js'
export async function runScenario(id: string, options: Options = {}, input?: JsonObject) {
// Load the task definition by ID; --input replaces its business inputs.
const scenario = getScenario(id)
const data = input ?? scenario.input
if (scenario.memory && data.dataset !== 'public-demo' && options.mode !== 'interactive' && !options.userId && !process.env.QONI_USER_ID) {
throw new Error('Business Memory writes require the current QONI_USER_ID; do not select an arbitrary bound user')
}
const context = await createContext(id, { ...options,
skipUserResolution: scenario.memory && data.dataset === 'public-demo' })
return withCleanup(context, async register => {
register('isolated demonstration user', () => cleanupDemoUser(context))
// Isolate demo preferences; business Memory belongs to the identified current user.
if (scenario.memory && data.dataset === 'public-demo') await isolateDemoUser(context)
return await executeScenario(context, scenario, data)
})
}
export async function executeScenario(context: Context, scenario: Scenario, input: JsonObject) {
const pages = input.pages
if (!Array.isArray(pages) || !pages.length || pages.some(page => typeof page !== 'string' || !/^https:\/\//.test(page))) {
throw new Error('Input pages must be an array of HTTPS URLs')
}
if (input.requiresLogin === true && context.mode !== 'interactive') {
throw new Error('Targets that require sign-in need --interactive and user-controlled login')
}
const memoryScopes = scenario.memory
? [QoniScopes.GUMEM_MEMORY_READ, QoniScopes.GUMEM_MEMORY_WRITE, QoniScopes.GUMEM_MESSAGE_WRITE] : []
// delegate() is an application helper around the SDK delegation methods.
const grant = await delegate(context, scenario.id, scenario.products, memoryScopes)
// Read the effective scopes; readWithRetry() retries only retryable read failures.
const { data: tokenInfo } = await readWithRetry(context, 'delegation introspection',
() => context.qoni.genauth.introspectDelegationToken({ token: grant.token }))
const info = object(tokenInfo)
if (info.active !== true) throw new Error('The delegation token is not active')
const audit = { grantId: grant.grantId, auditId: grant.auditId, scopes: info.scope }
let memory: unknown
if (scenario.memory) {
// A Session associates this conversation with the user; the app chooses sessionId.
const sessionId = `${scenario.id}-${Date.now()}`
await context.qoni.gumem.createSession({
token: grant.token, userId: context.userId, sessionId, title: scenario.id,
})
const preferences = object(input.business ?? {}).confirmedPreferences
if (Array.isArray(preferences) && preferences.length) {
// Store confirmed preferences only; sync: true requests synchronous processing.
await context.qoni.gumem.addMessages({ token: grant.token, userId: context.userId, sessionId, sync: true,
messages: [{ role: 'user', content: `Confirmed demonstration preferences: ${preferences.join('; ')}` }] })
}
// Recall relevant preferences for the later task prompt.
memory = (await readWithRetry(context, 'GUMem recall', () => context.qoni.gumem.recall({ token: grant.token, sessionId,
query: 'Confirmed preferences relevant to this task', details: true }))).data
save(context, 'memory.json', { sessionId, context: memory })
if (Array.isArray(preferences) && preferences.length && !preferences.every(value => JSON.stringify(memory).includes(String(value)))) {
throw new Error('Recall did not include the confirmed preferences just written by this demo')
}
}
if (scenario.products.includes('track')) return runMonitor(context, scenario, input, grant.token, audit)
let hits: ReturnType<typeof searchHits> = []
if (scenario.queries) {
// Web Search returns results[]; DoAnything receives these sources to inspect.
const search = await context.qoni.webSearch.run({ token: grant.token, prompt: scenario.queries, maxResultsPerQuery: 3 })
save(context, 'search-ref.json', { runId: search.id, audit })
await withCleanup(context, async register => {
register('Web Search run', () => search.cancel('Documentation demonstration cleanup'))
const result = await settleRun(context, search)
save(context, 'search-result.json', result)
settled(result)
hits = searchHits(result.output)
})
}
// One-per-entry scenarios request exactly one item per input entry; others at most two.
const requiredItems = inputEntryCount(scenario.id, input)
// Assemble the task, inputs, search sources, and Memory as application-defined context.
const prompt = [scenario.task, `Task inputs: ${JSON.stringify(input)}`,
`Search sources: ${JSON.stringify(hits)}`, `Confirmed memory: ${JSON.stringify(memory ?? null)}`,
`Actual collection time: ${new Date().toISOString()}`,
requiredItems === undefined
? 'Inspect the supplied sources. Return at most two items in the requested JSON array, without prose or Markdown.'
: `Inspect the supplied sources. Return exactly ${requiredItems} item${requiredItems === 1 ? '' : 's'} in the requested JSON array, one per input entry, without prose or Markdown.`,
'Keep synthetic demonstration data identified as synthetic. Do not send messages, publish, pay, edit accounts or submit forms.',
input.requiresLogin === true ? 'Request user sign-in through an interaction when required; never enter credentials yourself.' : 'Public demonstration sources only; do not sign in.',
].join('\n\n')
// Start the Agent with this grant; capture receives delivered screenshots, not every step.
const run = await context.qoni.doAnything.run({ token: grant.token, prompt, capture: { screenshots: true } })
save(context, 'run-ref.json', { runId: run.id, session: run.sessionRef, audit })
let result: RunResult
const trace = (event: { type: string; data: unknown }) => {
context.eventCounts[event.type] = (context.eventCounts[event.type] ?? 0) + 1
if (['progress','message','done'].includes(event.type)) appendTrace(context, event)
if (event.type === 'browserLiveUrlChanged') {
const liveUrl = object(event.data).liveUrl
if (typeof liveUrl === 'string') context.browserUrl = liveUrl
}
}
return withCleanup(context, async register => {
register('DoAnything run', () => run.cancel('Documentation demonstration cleanup'))
if (context.delivery === 'events') {
// --events streams updates; wrap interaction data in an SDK handle for user handling.
for await (const event of run.events({ signal: AbortSignal.any([context.abort.signal, AbortSignal.timeout(context.timeoutMs)]) })) {
trace(event)
if (event.type === 'screenshot') renderScreenshot(context, event.image)
if (event.type === 'interaction') await handleInteraction(context, run.interactionHandle(event.data))
}
result = await settleRun(context, run)
} else {
// Callback mode receives this run's events inside wait; helpers save images and ask the user.
result = await settleRun(context, run, { onEvent: trace,
onScreenshot: (image, index) => renderScreenshot(context, image, index),
onInteraction: interaction => handleInteraction(context, interaction) })
}
save(context, 'result.json', result)
settled(result)
await saveArtifacts(context, result)
if (scenario.id === 'landing-page-audit-agent' && context.screenshots === 0) throw new Error('The landing-page audit did not deliver the requested screenshot')
// The app checks required fields and source URL formats; a reviewer still checks facts.
const items = validateItems(result.output, scenario.fields, scenario.sourceField)
// Scenarios that require one item per input entry are checked against the input.
checkInputCoverage(scenario.id, items, input)
const report = { scenario: scenario.id, dataset: input.dataset, passed: true, runId: run.id,
status: result.status, items, audit, screenshots: context.screenshots, interactions: context.interactions,
events: context.eventCounts, artifactIds: result.artifacts.map(artifact => artifact.id) }
save(context, 'report.json', report)
return report
})
}
async function runMonitor(context: Context, scenario: Scenario, input: JsonObject, token: string, audit: JsonObject) {
// Track creates a monitor with targets, extraction fields, and hourly scheduling.
const monitor = await context.qoni.track.create({ token, prompt: scenario.task,
targetUrls: input.pages, extractionSchema: { heading: 'string', source_url: 'string' },
tickInstructions: `Open the target URLs and read the actual visible heading. Return a JSON object with heading and source_url. ${scenario.task}`,
triggerDsl: { on: 'change' }, schedule: { kind: 'interval', intervalSeconds: 3600 } })
save(context, 'monitor-ref.json', { id: monitor.id, audit })
return withCleanup(context, async register => {
register('Track monitor', () => monitor.delete())
const definition = await monitor.get()
save(context, 'monitor-definition.json', definition)
if (object(definition.schedule).intervalSeconds !== 3600) throw new Error('Track did not persist the requested schedule interval')
// Run one tick and inspect its extraction by runId; completed alone does not prove success.
const tick = await monitor.runNow()
save(context, 'tick.json', tick)
if (tick.state !== 'completed') throw new Error(`Track execution failed: ${tick.state} / ${tick.error ?? ''}`)
const runId = tick.runId
if (typeof runId !== 'string') throw new Error('Track tick did not return a runId')
const detail = await monitor.run(runId)
save(context, 'tick-detail.json', detail)
if (detail.state !== 'completed' || !detail.extracted || !Object.keys(object(detail.extracted)).length) {
throw new Error('Track did not extract page data')
}
const extracted = object(detail.extracted)
if (typeof extracted.heading !== 'string' || !extracted.heading.trim() ||
typeof extracted.source_url !== 'string' || !/^https:\/\//.test(extracted.source_url)) {
throw new Error('Track extraction is missing a heading or source URL')
}
const normalizeUrl = (value: string) => { const url = new URL(value); url.hash = ''; return url.href.replace(/\/$/, '') }
if (!(input.pages as string[]).some(url => normalizeUrl(url) === normalizeUrl(String(extracted.source_url)))) {
throw new Error('The Track source URL is not a configured target')
}
// Check persisted pause/resume state; withCleanup() deletes the monitor on exit.
await monitor.pause()
if ((await monitor.get()).status !== 'paused') throw new Error('Track did not persist the paused state')
await monitor.resume()
if ((await monitor.get()).status !== 'active') throw new Error('Track did not persist the active state')
const report = { scenario: scenario.id, dataset: input.dataset, passed: true, monitorId: monitor.id,
runId, state: detail.state, outcome: detail.outcome, extracted: detail.extracted, audit }
save(context, 'report.json', report)
return report
})
}
export function inputFile(): JsonObject | undefined {
const index = process.argv.indexOf('--input')
return index >= 0 ? object(JSON.parse(readFileSync(process.argv[index + 1], 'utf8'))) : undefined
}After a successful tick, output/vendor-monitor-agent/report.json contains the monitor ID, tick run ID, actual extraction, and effective scopes. http.json records redacted request statuses. A completed state with empty extraction fails validation. If your deployment returns empty data, check the Track scraper, extractor, and worker configuration before retrying; this example cannot certify that deployment as working.
Data and memory boundaries
This scenario touches four kinds of data; none of them belongs in GUMem:
- Versioned rules: monitored-section lists, trigger thresholds, alert escalation rules — managed by version in your policy store, with every alert referencing the policy version.
- Business state: page baselines, per-run diffs, snapshots — carried by Track's run records and your baseline store, versioned and replayable.
- Audit records: the behavior chain of credential issuance and every fetch — maintained by GenAuth.
- User Memory: this scenario does not use GUMem. Baselines and diffs are business state, not user preferences; writing them into Memory only creates a second, unreconcilable source of truth.
Failure handling
| Situation | Recommended handling |
|---|---|
| A full-page redesign breaks the comparison | Treat it as a failure and replay that run; rebuild the baseline after human confirmation, never misreport the redesign as a terms change. |
The monitor keeps failing (consecutive_failures grows) | The monitor enters an abnormal state; check target reachability and rule configuration, then verify once with run_now after the fix. |
| A target page shows a login wall or CAPTCHA | Track signals that a human is needed; the owner responds (intervene) — no silent bypass. |
| A callback entry lacks a diff or source | App-side validation drops the entry and records how many were dropped; no source, no alert. |
Production notes
A page change is not yet a real risk: Track decides "it changed," while risk grading and follow-up always stay with the owner. Monitoring is read-only and notify-only — it never executes contract or procurement actions on a human's behalf. Set a floor on the schedule interval based on how often the pages actually update, to avoid load on vendor sites; monitors are long-lived, so transfer ownership and re-review notify channels when team members change.
Next steps
- Read Track for the monitor's full lifecycle and event stream.
- Read the Quickstart to run the shortest path for Agent identity and delegation.
- Continue with the Partner portal monitoring agent for an adjacent signed-in monitoring scenario.