Skip to content

Personal research agent ​

This page explains how a personal research agent runs multi-source search and cross-checking on the public web, and persists user-confirmed preferences and conclusions as cross-session memory. After reading it, you will understand why Web Agent and GUMem are the core modules here, when interactive delegation is actually needed, and what belongs in long-term Memory.

Use case ​

Users repeatedly research similar topics — hardware choices, competitor moves, papers, investment information — and personal deep dives like purchase decisions, school selection, or gathering medical information. A single search rarely makes this kind of research solid: sources need cross-checking, conclusions need citations, and the next round should continue from previously confirmed conclusions instead of starting over.

Typical triggers:

  • Before a major purchase, specs and real-world feedback must be cross-checked across review sites, forums, and official pages.
  • Before choosing a school or program, information from many parties must be consolidated, separating official statements from third-party opinions.
  • Around a medical visit, public medical information must be organized into a sourced reading list for the user and their doctor.

Engineering challenges ​

  • Uneven source reliability: review sites may have commercial bias, forum takes mix truth with noise, and official pages only state the upside. Without multi-source cross-checking, a single-source claim easily ships as a conclusion.
  • Freshness is hard to judge: prices, models, and policies change constantly, and last year's conclusion may already be stale. Without a source URL and collection time on every finding, the user cannot decide what to trust.
  • Unclear long-term retention scope: budget, decision criteria, and confirmed conclusions deserve cross-session memory; full page dumps and unconfirmed claims belong only to the current task. A blurry boundary turns Memory into a store of stale web pages.
  • Interactive consent only appears in a minority of cases: every product call needs a GenAuth delegate token, but public-web research runs on a silent grant; only when the research must open subscription databases, private forums, or other signed-in sources does the user need to interactively confirm "who is accessing on my behalf".

Module composition ​

ModuleRoleNotes
GenAuthCoreThe runtime issues short-lived credentials via silent delegation (required for every product call); interactive consent is not needed by default here — upgrade to interactive only for signed-in sources or write actions (posting, ordering).
Web AgentCoreRuns multi-source search and cross-checking via WebSearch, keeping a source URL and collection time per finding and flagging single-source facts.
GUMemCoreStores personal preferences (budget, region, key metrics), user-confirmed conclusions, and research topic history — the memory that genuinely spans sessions.

Personal research agent architecture

Every Web Agent and GUMem product call requires a GenAuth delegate token; for public, read-only research a silently issued grant is enough — no enterprise-style consent ceremony. The silent runtime credential already provides the constraints this scenario needs: it is short-lived, the user can revoke it at any time, and every call carries a grantId and auditId for traceability.

Only two situations require upgrading to mode: 'interactive' delegation, with explicit user confirmation in the Qoni Console:

  • Signed-in sources: the research needs pages visible only after the user signs in — subscription databases, paid reviews, private forums.
  • Write actions: the task would post, comment, or order — actions that change external state, which this scenario excludes by default.

Note: this sample requests webSearch, doAnything product permissions and GUMem read/write scopes. Your application and downstream services must configure and enforce business limits such as domain lists. Product scopes and prompt rules do not enforce those detailed limits; this sample does not configure them. See Delegate token and attenuation.

Workflow ​

Personal research agent workflow

  1. The user submits a research question, such as "camera options within a $3,000 budget".

  2. Your app obtains a silent runtime credential scoped to web search and GUMem memory read/write.

  3. GUMem recalls relevant preferences and confirmed conclusions on the same topic — budget, region, key metrics, and the outcome of the last round.

  4. Web Agent runs multiple queries across official pages, reviews, and community threads, keeping a source URL and collection time per finding.

    Checkpoint: Key facts backed by a single source are flagged "not cross-checked"; conflicting sources are presented as-is, never guessed away.

  5. The Agent assembles the report, attaching a citation and confidence label to every conclusion; for medical or legal topics, the output states it is an information digest, not professional advice.

  6. The user reviews the report and confirms which conclusions and preferences are worth keeping long-term.

    Checkpoint: Only user-confirmed preferences and conclusions are written back to long-term Memory; unconfirmed web content stays within this task.

Example code ​

This example uses @qoniai/qoni 0.9.0, published on npm. The download includes the same SDK version, installed with npm ci. The demo researches which public Firefox features support browser privacy using official privacy and getting-started sources. The SDK reads real public pages; business inputs in scenarios.ts are labeled public-demo.

The sample writes and recalls confirmed preferences under an isolated demo user, then includes that context in the browser task. Web Search first supplies results[], which DoAnything then reads and analyzes. Supply your own JSON with --input; use --interactive when the user must consent to delegation. For site sign-in and user responses, see Qoni SDK.

Download the complete runnable examples, or run from the documentation repository:

bash
cd examples/qoni
npm ci
npm run case -- personal-research-agent
# Supply your own inputs
npm run case -- personal-research-agent --input /path/to/input.json

Set server-side QONI_ACCESS_KEY and QONI_SECRET_KEY. QONI_USER_ID can identify your application's current GenAuth user; the local demo otherwise selects a user from the bound pool. Demonstration Memory writes use an isolated user rather than changing a business user's preferences.

This scenario's executable entry point:

ts
import { cliOptions } from '../runtime.js'
import { inputFile, runScenario } from '../run-case.js'

// Research cited findings using the confirmed concise-explanation preference.
const report = await runScenario('personal-research-agent', cliOptions(), inputFile())
// report.items: findings, source lists, timestamps, confidence, and caveats.
// Validated fields: finding, sourceUrls, capturedAt, confidence, caveats.
console.log(JSON.stringify(report, null, 2))

The entry point loads the definition below by scenario ID. The code is included directly from scenarios.ts, with comments shown in the page language: the task, output fields, source field, Web Search queries (if any), whether Memory is used, and the demonstration input. The pipeline appends the input data, search sources, recalled Memory, and shared safety constraints to the task to build the final prompt; see the pipeline below for the full assembly.

ts
// Research cited findings using the confirmed concise-explanation preference.
research('personal-research-agent',
  'Research the supplied question using actual sources. Return [{finding,sourceUrls,capturedAt,confidence,caveats}]. Cite every finding and flag single-source or conflicting evidence. Include only verified facts or explicitly qualified findings. Do not sign in or modify any website.',
  ['finding','sourceUrls','capturedAt','confidence','caveats'], 'sourceUrls', ['Firefox enhanced tracking protection Mozilla official'], sample([privacy,support], { question:'What public Firefox features support browser privacy?',confirmedPreferences:['Use concise explanations.'] }), true),

The sample implements runScenario(), browser(), research(), and sample() as application functions. The pipeline below makes the actual SDK calls: delegation and introspection → required Memory/search → browser task or monitor → validation and saving. fields and sourceField define the application's output checks. The complete application helpers are in the package's runtime.ts.

Inspect the actual SDK pipeline
ts
// browser() uses DoAnything; research() searches first; sample() labels public-demo inputs.
const firefox = 'https://www.mozilla.org/en-US/firefox/new/'
const manifesto = 'https://www.mozilla.org/en-US/about/manifesto/'
const privacy = 'https://www.mozilla.org/en-US/privacy/firefox/'
const support = 'https://support.mozilla.org/en-US/kb/get-started-firefox-overview-main-features'
// These copy rules become prompt context; server permissions and business checks remain separate.
const policy = {
  version: 'demo-2026-09',
  approvedClaims: ['Describe only features supported by the cited page.'],
  forbiddenClaims: ['guaranteed security', '100% private', 'unverified pricing or performance'],
  voice: 'concise and warm',
}
const sample = (pages: string[], business: JsonObject = {}): JsonObject => ({
  dataset: 'public-demo', pages, policy, business,
  notice: 'Business records are synthetic demonstration inputs. Referenced websites and SDK execution are real.',
})
// fields lists required output keys; sourceField identifies URL checks; memory enables GUMem calls.
const browser = (id: string, task: string, fields: string[], sourceField: string | undefined, input: JsonObject, memory = false): Scenario => ({
  id, products: ['doAnything'], task, fields, sourceField, input, memory,
})
const research = (id: string, task: string, fields: string[], sourceField: string, queries: string[], input: JsonObject, memory = false): Scenario => ({
  id, products: ['webSearch', 'doAnything'], task, fields, sourceField, queries, input, memory,
})
ts
import { QoniScopes, type JsonObject, type RunResult } from '@qoniai/qoni'
import { readFileSync } from 'node:fs'
import { getScenario, type Scenario } from './scenarios.js'
import { appendTrace, checkInputCoverage, cleanupDemoUser, createContext, delegate, handleInteraction, inputEntryCount, isolateDemoUser, object, readWithRetry, renderScreenshot,
  save, saveArtifacts, searchHits, settled, settleRun, validateItems, withCleanup, type Context, type Options } from './runtime.js'

export async function runScenario(id: string, options: Options = {}, input?: JsonObject) {
  // Load the task definition by ID; --input replaces its business inputs.
  const scenario = getScenario(id)
  const data = input ?? scenario.input
  if (scenario.memory && data.dataset !== 'public-demo' && options.mode !== 'interactive' && !options.userId && !process.env.QONI_USER_ID) {
    throw new Error('Business Memory writes require the current QONI_USER_ID; do not select an arbitrary bound user')
  }
  const context = await createContext(id, { ...options,
    skipUserResolution: scenario.memory && data.dataset === 'public-demo' })
  return withCleanup(context, async register => {
    register('isolated demonstration user', () => cleanupDemoUser(context))
    // Isolate demo preferences; business Memory belongs to the identified current user.
    if (scenario.memory && data.dataset === 'public-demo') await isolateDemoUser(context)
    return await executeScenario(context, scenario, data)
  })
}

export async function executeScenario(context: Context, scenario: Scenario, input: JsonObject) {
  const pages = input.pages
  if (!Array.isArray(pages) || !pages.length || pages.some(page => typeof page !== 'string' || !/^https:\/\//.test(page))) {
    throw new Error('Input pages must be an array of HTTPS URLs')
  }
  if (input.requiresLogin === true && context.mode !== 'interactive') {
    throw new Error('Targets that require sign-in need --interactive and user-controlled login')
  }
  const memoryScopes = scenario.memory
    ? [QoniScopes.GUMEM_MEMORY_READ, QoniScopes.GUMEM_MEMORY_WRITE, QoniScopes.GUMEM_MESSAGE_WRITE] : []
  // delegate() is an application helper around the SDK delegation methods.
  const grant = await delegate(context, scenario.id, scenario.products, memoryScopes)
  // Read the effective scopes; readWithRetry() retries only retryable read failures.
  const { data: tokenInfo } = await readWithRetry(context, 'delegation introspection',
    () => context.qoni.genauth.introspectDelegationToken({ token: grant.token }))
  const info = object(tokenInfo)
  if (info.active !== true) throw new Error('The delegation token is not active')
  const audit = { grantId: grant.grantId, auditId: grant.auditId, scopes: info.scope }
  let memory: unknown
  if (scenario.memory) {
    // A Session associates this conversation with the user; the app chooses sessionId.
    const sessionId = `${scenario.id}-${Date.now()}`
    await context.qoni.gumem.createSession({
      token: grant.token, userId: context.userId, sessionId, title: scenario.id,
    })
    const preferences = object(input.business ?? {}).confirmedPreferences
    if (Array.isArray(preferences) && preferences.length) {
      // Store confirmed preferences only; sync: true requests synchronous processing.
      await context.qoni.gumem.addMessages({ token: grant.token, userId: context.userId, sessionId, sync: true,
        messages: [{ role: 'user', content: `Confirmed demonstration preferences: ${preferences.join('; ')}` }] })
    }
    // Recall relevant preferences for the later task prompt.
    memory = (await readWithRetry(context, 'GUMem recall', () => context.qoni.gumem.recall({ token: grant.token, sessionId,
      query: 'Confirmed preferences relevant to this task', details: true }))).data
    save(context, 'memory.json', { sessionId, context: memory })
    if (Array.isArray(preferences) && preferences.length && !preferences.every(value => JSON.stringify(memory).includes(String(value)))) {
      throw new Error('Recall did not include the confirmed preferences just written by this demo')
    }
  }

  if (scenario.products.includes('track')) return runMonitor(context, scenario, input, grant.token, audit)

  let hits: ReturnType<typeof searchHits> = []
  if (scenario.queries) {
    // Web Search returns results[]; DoAnything receives these sources to inspect.
    const search = await context.qoni.webSearch.run({ token: grant.token, prompt: scenario.queries, maxResultsPerQuery: 3 })
    save(context, 'search-ref.json', { runId: search.id, audit })
    await withCleanup(context, async register => {
      register('Web Search run', () => search.cancel('Documentation demonstration cleanup'))
      const result = await settleRun(context, search)
      save(context, 'search-result.json', result)
      settled(result)
      hits = searchHits(result.output)
    })
  }

  // One-per-entry scenarios request exactly one item per input entry; others at most two.
  const requiredItems = inputEntryCount(scenario.id, input)
  // Assemble the task, inputs, search sources, and Memory as application-defined context.
  const prompt = [scenario.task, `Task inputs: ${JSON.stringify(input)}`,
    `Search sources: ${JSON.stringify(hits)}`, `Confirmed memory: ${JSON.stringify(memory ?? null)}`,
    `Actual collection time: ${new Date().toISOString()}`,
    requiredItems === undefined
      ? 'Inspect the supplied sources. Return at most two items in the requested JSON array, without prose or Markdown.'
      : `Inspect the supplied sources. Return exactly ${requiredItems} item${requiredItems === 1 ? '' : 's'} in the requested JSON array, one per input entry, without prose or Markdown.`,
    'Keep synthetic demonstration data identified as synthetic. Do not send messages, publish, pay, edit accounts or submit forms.',
    input.requiresLogin === true ? 'Request user sign-in through an interaction when required; never enter credentials yourself.' : 'Public demonstration sources only; do not sign in.',
  ].join('\n\n')
  // Start the Agent with this grant; capture receives delivered screenshots, not every step.
  const run = await context.qoni.doAnything.run({ token: grant.token, prompt, capture: { screenshots: true } })
  save(context, 'run-ref.json', { runId: run.id, session: run.sessionRef, audit })
  let result: RunResult
  const trace = (event: { type: string; data: unknown }) => {
    context.eventCounts[event.type] = (context.eventCounts[event.type] ?? 0) + 1
    if (['progress','message','done'].includes(event.type)) appendTrace(context, event)
    if (event.type === 'browserLiveUrlChanged') {
      const liveUrl = object(event.data).liveUrl
      if (typeof liveUrl === 'string') context.browserUrl = liveUrl
    }
  }
  return withCleanup(context, async register => {
    register('DoAnything run', () => run.cancel('Documentation demonstration cleanup'))
    if (context.delivery === 'events') {
      // --events streams updates; wrap interaction data in an SDK handle for user handling.
      for await (const event of run.events({ signal: AbortSignal.any([context.abort.signal, AbortSignal.timeout(context.timeoutMs)]) })) {
        trace(event)
        if (event.type === 'screenshot') renderScreenshot(context, event.image)
        if (event.type === 'interaction') await handleInteraction(context, run.interactionHandle(event.data))
      }
      result = await settleRun(context, run)
    } else {
      // Callback mode receives this run's events inside wait; helpers save images and ask the user.
      result = await settleRun(context, run, { onEvent: trace,
        onScreenshot: (image, index) => renderScreenshot(context, image, index),
        onInteraction: interaction => handleInteraction(context, interaction) })
    }
    save(context, 'result.json', result)
    settled(result)
    await saveArtifacts(context, result)
    if (scenario.id === 'landing-page-audit-agent' && context.screenshots === 0) throw new Error('The landing-page audit did not deliver the requested screenshot')
    // The app checks required fields and source URL formats; a reviewer still checks facts.
    const items = validateItems(result.output, scenario.fields, scenario.sourceField)
    // Scenarios that require one item per input entry are checked against the input.
    checkInputCoverage(scenario.id, items, input)
    const report = { scenario: scenario.id, dataset: input.dataset, passed: true, runId: run.id,
      status: result.status, items, audit, screenshots: context.screenshots, interactions: context.interactions,
      events: context.eventCounts, artifactIds: result.artifacts.map(artifact => artifact.id) }
    save(context, 'report.json', report)
    return report
  })
}

async function runMonitor(context: Context, scenario: Scenario, input: JsonObject, token: string, audit: JsonObject) {
  // Track creates a monitor with targets, extraction fields, and hourly scheduling.
  const monitor = await context.qoni.track.create({ token, prompt: scenario.task,
    targetUrls: input.pages, extractionSchema: { heading: 'string', source_url: 'string' },
    tickInstructions: `Open the target URLs and read the actual visible heading. Return a JSON object with heading and source_url. ${scenario.task}`,
    triggerDsl: { on: 'change' }, schedule: { kind: 'interval', intervalSeconds: 3600 } })
  save(context, 'monitor-ref.json', { id: monitor.id, audit })
  return withCleanup(context, async register => {
    register('Track monitor', () => monitor.delete())
    const definition = await monitor.get()
    save(context, 'monitor-definition.json', definition)
    if (object(definition.schedule).intervalSeconds !== 3600) throw new Error('Track did not persist the requested schedule interval')
    // Run one tick and inspect its extraction by runId; completed alone does not prove success.
    const tick = await monitor.runNow()
    save(context, 'tick.json', tick)
    if (tick.state !== 'completed') throw new Error(`Track execution failed: ${tick.state} / ${tick.error ?? ''}`)
    const runId = tick.runId
    if (typeof runId !== 'string') throw new Error('Track tick did not return a runId')
    const detail = await monitor.run(runId)
    save(context, 'tick-detail.json', detail)
    if (detail.state !== 'completed' || !detail.extracted || !Object.keys(object(detail.extracted)).length) {
      throw new Error('Track did not extract page data')
    }
    const extracted = object(detail.extracted)
    if (typeof extracted.heading !== 'string' || !extracted.heading.trim() ||
      typeof extracted.source_url !== 'string' || !/^https:\/\//.test(extracted.source_url)) {
      throw new Error('Track extraction is missing a heading or source URL')
    }
    const normalizeUrl = (value: string) => { const url = new URL(value); url.hash = ''; return url.href.replace(/\/$/, '') }
    if (!(input.pages as string[]).some(url => normalizeUrl(url) === normalizeUrl(String(extracted.source_url)))) {
      throw new Error('The Track source URL is not a configured target')
    }
    // Check persisted pause/resume state; withCleanup() deletes the monitor on exit.
    await monitor.pause()
    if ((await monitor.get()).status !== 'paused') throw new Error('Track did not persist the paused state')
    await monitor.resume()
    if ((await monitor.get()).status !== 'active') throw new Error('Track did not persist the active state')
    const report = { scenario: scenario.id, dataset: input.dataset, passed: true, monitorId: monitor.id,
      runId, state: detail.state, outcome: detail.outcome, extracted: detail.extracted, audit }
    save(context, 'report.json', report)
    return report
  })
}

export function inputFile(): JsonObject | undefined {
  const index = process.argv.indexOf('--input')
  return index >= 0 ? object(JSON.parse(readFileSync(process.argv[index + 1], 'utf8'))) : undefined
}

Results are written to output/personal-research-agent/report.json. report.items contains findings, source lists, timestamps, confidence, and caveats, with fields finding, sourceUrls, capturedAt, confidence, caveats. audit links the grant ID, audit ID, and effective scopes; http.json records redacted request statuses. The application parses DoAnything output and checks required fields and source URL formats. A business reviewer still assesses the content against the original sources.

Memory strategy ​

  • Into Memory: user-confirmed preferences (budget, region, key metrics), confirmed conclusions (with source pointers, confidence, and time), and research topic history; preference memories decay quarterly so old tastes do not dominate new decisions.
  • Not into Memory: full page dumps, unconfirmed web claims, and one-off comparison data — discarded when the task ends, re-collected when needed.
  • Corrections: when new evidence overturns a conclusion (for example, a model hit by a quality scandal), mark the old conclusion invalidated and point it to the new memory instead of physically deleting it.

Failure handling ​

SituationRecommended handling
A source page requires sign-in or shows a CAPTCHASkip the source and record it; if signed-in access is genuinely needed, start a separate interactive delegation — never silently bypass.
A key fact rests on a single sourceFlag it "not cross-checked"; never treat an uncorroborated claim as verified.
A result entry lacks a source or collection timeApp-side validation drops the entry and the report notes how many were dropped.
A recalled conclusion conflicts with new evidenceNew evidence wins; mark the old conclusion invalidated and write the correction back to GUMem.

Production notes ​

Never write full web pages into long-term Memory: only user-confirmed preferences, conclusions, or long-term constraints are written back; everything else is discarded when the task ends. Research output on medical or legal topics must state it is an information digest, not professional advice — final decisions belong to the user and licensed professionals. Time-sensitive data in the report (prices, policies) should keep its collection time and be re-collected rather than reused once past a reasonable window.

Next steps ​