Skip to content

Recruiting sourcing agent ​

This page explains how a recruiting sourcing agent searches public candidate profiles against approved role criteria and produces a sourcing list where every match reason carries a source. After reading it, you will understand why Web Agent is the core module here, why role rules live in the ATS or a policy store rather than Memory, and how bias risk in candidate evaluation is handled as an engineering concern.

Use case ​

Recruiting teams need to find public candidate profiles for a role and generate candidate summaries. Public material is scattered across personal sites, tech communities, and public résumé pages — searching manually is slow and inconsistent across batches, and handing a recruiting-system account to a script leaves candidate-data access and retention completely unbounded. The Agent only processes publicly visible information, and its output is a sourcing list for the recruiter to review.

Typical triggers:

  • A new role opens, and a first sourcing list is needed within days.
  • Role requirements change, and the existing candidate pool must be re-screened against the new criteria.
  • A hard-to-fill role stays open, and newly appearing public candidate leads must be added periodically.

Engineering challenges ​

  • The public-data collection boundary: candidate material sits on personal sites, communities, and portfolios; what is truly public versus sign-in-only must be enforced as a hard boundary — collecting candidate data past a login wall is a compliance incident, not a technical optimization.
  • Criteria drift and bias entrenchment: if each script interprets the JD its own way, batches cannot be compared; and calibrating criteria from "who historically advanced to interviews" hardens past bias into a ranking signal. Criteria must come from approved, versioned role requirements — never from historical interview outcomes.
  • Evidence and reviewability: every match reason must map to a concrete public source and capture time so the recruiter can verify it; an unsourced "feels like a fit" is neither auditable nor defensible.

Module composition ​

ModuleRoleNotes
GenAuthCoreThe runtime issues short-lived credentials through silent delegation (required for every product call): short-lived, revocable, audited; no interactive confirmation is needed by default in this scenario — upgrade to interactive delegation only when signing in to an ATS or recruiting system.
Web AgentCoreSearches public candidate profiles via WebSearch, keeping a source URL and capture time per item; sign-in-only pages are skipped and recorded.
GUMemNot usedRole requirements are approved, versioned recruiting rules, injected per version from the ATS or a policy store; past interview outcomes must not serve as a candidate ranking signal, so there is no "screening preference memory" worth persisting across sessions.

Recruiting sourcing agent architecture

Every product call must carry a GenAuth-issued delegation token — what is optional is not delegation itself but interactive confirmation. Public candidate searching does not need per-task user consent: a silently issued runtime credential already provides the constraints this scenario needs — it is short-lived, revocable at any time, and every search carries a grantId and auditId attributable to a specific role and task.

Situations that require upgrading to mode: 'interactive' delegation, with explicit recruiter confirmation in the Qoni Console:

  • Signed-in access: the task needs candidate data behind the ATS, a recruiting system, or any other sign-in wall — not covered by this page's example.
  • Write actions: messaging candidates or modifying the candidate pool — excluded by default here; outreach is always sent by a human after the recruiter approves the draft.

Note: this sample requests webSearch, doAnything product permissions. Your application and downstream services must configure and enforce business limits such as searchable site lists. Product scopes and prompt rules do not enforce those detailed limits; this sample does not configure them. See Delegate token and attenuation.

Workflow ​

Recruiting sourcing agent workflow

  1. The recruiter picks the role and confirms the sourcing scope.

  2. Your app loads the approved role criteria (with a version number) from the ATS or policy store and obtains a silent runtime credential.

  3. Web Agent searches public candidate profiles, extracting experience, skills, and public work per item, with a source URL and capture time.

    Checkpoint: Only publicly visible information is collected; sign-in-only candidate pages are skipped and recorded, never bypassed.

  4. The Agent generates candidate summaries and match reasons against the role criteria, each reason tied to a concrete source and citing the criteria version.

    Checkpoint: Summaries must not contain sensitive attributes (age, family status, ethnicity, and so on) or signals unrelated to the role criteria; anything found is removed and recorded. Past interview outcomes play no part in ranking.

  5. The recruiter reviews the list, marking advance or exclude; review results are recorded in the ATS and feed the next human revision review of the role criteria.

  6. When outreach is wanted, the Agent drafts the message for the recruiter to approve, and a human sends it — candidates are never contacted automatically.

Example code ​

This example uses @qoniai/qoni 0.9.0, published on npm. The download includes the same SDK version, installed with npm ci. The demo researches Mozilla public leadership information for a synthetic role and screening criteria. The SDK reads real public pages; business inputs in scenarios.ts are labeled public-demo.

This demo does not read or write GUMem; its task uses the supplied page list and explicit business inputs. Web Search first supplies results[], which DoAnything then reads and analyzes. Supply your own JSON with --input; use --interactive when the user must consent to delegation. For site sign-in and user responses, see Qoni SDK.

Download the complete runnable examples, or run from the documentation repository:

bash
cd examples/qoni
npm ci
npm run case -- recruiting-sourcing-agent
# Supply your own inputs
npm run case -- recruiting-sourcing-agent --input /path/to/input.json

Set server-side QONI_ACCESS_KEY and QONI_SECRET_KEY. QONI_USER_ID can identify your application's current GenAuth user; the local demo otherwise selects a user from the bound pool. Demonstration Memory writes use an isolated user rather than changing a business user's preferences.

This scenario's executable entry point:

ts
import { cliOptions } from '../runtime.js'
import { inputFile, runScenario } from '../run-case.js'

// Draft candidate research from public professional evidence.
const report = await runScenario('recruiting-sourcing-agent', cliOptions(), inputFile())
// report.items: names, evidence, and sources for recruiting review.
// Validated fields: name, evidence, sourceUrl.
console.log(JSON.stringify(report, null, 2))

The entry point loads the definition below by scenario ID. The code is included directly from scenarios.ts, with comments shown in the page language: the task, output fields, source field, Web Search queries (if any), whether Memory is used, and the demonstration input. The pipeline appends the input data, search sources, recalled Memory, and shared safety constraints to the task to build the final prompt; see the pipeline below for the full assembly.

ts
// Draft candidate research from public professional evidence.
research('recruiting-sourcing-agent',
  'Use public professional information to prepare a candidate-research draft for the synthetic role. Return [{name,evidence,sourceUrl}]. Evidence must be visible in the public sources. Do not infer protected traits, collect private contacts, contact anyone or make a hiring decision.',
  ['name','evidence','sourceUrl'], 'sourceUrl', ['Mozilla public leadership professional biography'], sample(['https://www.mozilla.org/en-US/about/leadership/'], { roleId:'demo-open-web-role',approvedCriteria:['Public browser-industry experience'] })),

The sample implements runScenario(), browser(), research(), and sample() as application functions. The pipeline below makes the actual SDK calls: delegation and introspection → required Memory/search → browser task or monitor → validation and saving. fields and sourceField define the application's output checks. The complete application helpers are in the package's runtime.ts.

Inspect the actual SDK pipeline
ts
// browser() uses DoAnything; research() searches first; sample() labels public-demo inputs.
const firefox = 'https://www.mozilla.org/en-US/firefox/new/'
const manifesto = 'https://www.mozilla.org/en-US/about/manifesto/'
const privacy = 'https://www.mozilla.org/en-US/privacy/firefox/'
const support = 'https://support.mozilla.org/en-US/kb/get-started-firefox-overview-main-features'
// These copy rules become prompt context; server permissions and business checks remain separate.
const policy = {
  version: 'demo-2026-09',
  approvedClaims: ['Describe only features supported by the cited page.'],
  forbiddenClaims: ['guaranteed security', '100% private', 'unverified pricing or performance'],
  voice: 'concise and warm',
}
const sample = (pages: string[], business: JsonObject = {}): JsonObject => ({
  dataset: 'public-demo', pages, policy, business,
  notice: 'Business records are synthetic demonstration inputs. Referenced websites and SDK execution are real.',
})
// fields lists required output keys; sourceField identifies URL checks; memory enables GUMem calls.
const browser = (id: string, task: string, fields: string[], sourceField: string | undefined, input: JsonObject, memory = false): Scenario => ({
  id, products: ['doAnything'], task, fields, sourceField, input, memory,
})
const research = (id: string, task: string, fields: string[], sourceField: string, queries: string[], input: JsonObject, memory = false): Scenario => ({
  id, products: ['webSearch', 'doAnything'], task, fields, sourceField, queries, input, memory,
})
ts
import { QoniScopes, type JsonObject, type RunResult } from '@qoniai/qoni'
import { readFileSync } from 'node:fs'
import { getScenario, type Scenario } from './scenarios.js'
import { appendTrace, checkInputCoverage, cleanupDemoUser, createContext, delegate, handleInteraction, inputEntryCount, isolateDemoUser, object, readWithRetry, renderScreenshot,
  save, saveArtifacts, searchHits, settled, settleRun, validateItems, withCleanup, type Context, type Options } from './runtime.js'

export async function runScenario(id: string, options: Options = {}, input?: JsonObject) {
  // Load the task definition by ID; --input replaces its business inputs.
  const scenario = getScenario(id)
  const data = input ?? scenario.input
  if (scenario.memory && data.dataset !== 'public-demo' && options.mode !== 'interactive' && !options.userId && !process.env.QONI_USER_ID) {
    throw new Error('Business Memory writes require the current QONI_USER_ID; do not select an arbitrary bound user')
  }
  const context = await createContext(id, { ...options,
    skipUserResolution: scenario.memory && data.dataset === 'public-demo' })
  return withCleanup(context, async register => {
    register('isolated demonstration user', () => cleanupDemoUser(context))
    // Isolate demo preferences; business Memory belongs to the identified current user.
    if (scenario.memory && data.dataset === 'public-demo') await isolateDemoUser(context)
    return await executeScenario(context, scenario, data)
  })
}

export async function executeScenario(context: Context, scenario: Scenario, input: JsonObject) {
  const pages = input.pages
  if (!Array.isArray(pages) || !pages.length || pages.some(page => typeof page !== 'string' || !/^https:\/\//.test(page))) {
    throw new Error('Input pages must be an array of HTTPS URLs')
  }
  if (input.requiresLogin === true && context.mode !== 'interactive') {
    throw new Error('Targets that require sign-in need --interactive and user-controlled login')
  }
  const memoryScopes = scenario.memory
    ? [QoniScopes.GUMEM_MEMORY_READ, QoniScopes.GUMEM_MEMORY_WRITE, QoniScopes.GUMEM_MESSAGE_WRITE] : []
  // delegate() is an application helper around the SDK delegation methods.
  const grant = await delegate(context, scenario.id, scenario.products, memoryScopes)
  // Read the effective scopes; readWithRetry() retries only retryable read failures.
  const { data: tokenInfo } = await readWithRetry(context, 'delegation introspection',
    () => context.qoni.genauth.introspectDelegationToken({ token: grant.token }))
  const info = object(tokenInfo)
  if (info.active !== true) throw new Error('The delegation token is not active')
  const audit = { grantId: grant.grantId, auditId: grant.auditId, scopes: info.scope }
  let memory: unknown
  if (scenario.memory) {
    // A Session associates this conversation with the user; the app chooses sessionId.
    const sessionId = `${scenario.id}-${Date.now()}`
    await context.qoni.gumem.createSession({
      token: grant.token, userId: context.userId, sessionId, title: scenario.id,
    })
    const preferences = object(input.business ?? {}).confirmedPreferences
    if (Array.isArray(preferences) && preferences.length) {
      // Store confirmed preferences only; sync: true requests synchronous processing.
      await context.qoni.gumem.addMessages({ token: grant.token, userId: context.userId, sessionId, sync: true,
        messages: [{ role: 'user', content: `Confirmed demonstration preferences: ${preferences.join('; ')}` }] })
    }
    // Recall relevant preferences for the later task prompt.
    memory = (await readWithRetry(context, 'GUMem recall', () => context.qoni.gumem.recall({ token: grant.token, sessionId,
      query: 'Confirmed preferences relevant to this task', details: true }))).data
    save(context, 'memory.json', { sessionId, context: memory })
    if (Array.isArray(preferences) && preferences.length && !preferences.every(value => JSON.stringify(memory).includes(String(value)))) {
      throw new Error('Recall did not include the confirmed preferences just written by this demo')
    }
  }

  if (scenario.products.includes('track')) return runMonitor(context, scenario, input, grant.token, audit)

  let hits: ReturnType<typeof searchHits> = []
  if (scenario.queries) {
    // Web Search returns results[]; DoAnything receives these sources to inspect.
    const search = await context.qoni.webSearch.run({ token: grant.token, prompt: scenario.queries, maxResultsPerQuery: 3 })
    save(context, 'search-ref.json', { runId: search.id, audit })
    await withCleanup(context, async register => {
      register('Web Search run', () => search.cancel('Documentation demonstration cleanup'))
      const result = await settleRun(context, search)
      save(context, 'search-result.json', result)
      settled(result)
      hits = searchHits(result.output)
    })
  }

  // One-per-entry scenarios request exactly one item per input entry; others at most two.
  const requiredItems = inputEntryCount(scenario.id, input)
  // Assemble the task, inputs, search sources, and Memory as application-defined context.
  const prompt = [scenario.task, `Task inputs: ${JSON.stringify(input)}`,
    `Search sources: ${JSON.stringify(hits)}`, `Confirmed memory: ${JSON.stringify(memory ?? null)}`,
    `Actual collection time: ${new Date().toISOString()}`,
    requiredItems === undefined
      ? 'Inspect the supplied sources. Return at most two items in the requested JSON array, without prose or Markdown.'
      : `Inspect the supplied sources. Return exactly ${requiredItems} item${requiredItems === 1 ? '' : 's'} in the requested JSON array, one per input entry, without prose or Markdown.`,
    'Keep synthetic demonstration data identified as synthetic. Do not send messages, publish, pay, edit accounts or submit forms.',
    input.requiresLogin === true ? 'Request user sign-in through an interaction when required; never enter credentials yourself.' : 'Public demonstration sources only; do not sign in.',
  ].join('\n\n')
  // Start the Agent with this grant; capture receives delivered screenshots, not every step.
  const run = await context.qoni.doAnything.run({ token: grant.token, prompt, capture: { screenshots: true } })
  save(context, 'run-ref.json', { runId: run.id, session: run.sessionRef, audit })
  let result: RunResult
  const trace = (event: { type: string; data: unknown }) => {
    context.eventCounts[event.type] = (context.eventCounts[event.type] ?? 0) + 1
    if (['progress','message','done'].includes(event.type)) appendTrace(context, event)
    if (event.type === 'browserLiveUrlChanged') {
      const liveUrl = object(event.data).liveUrl
      if (typeof liveUrl === 'string') context.browserUrl = liveUrl
    }
  }
  return withCleanup(context, async register => {
    register('DoAnything run', () => run.cancel('Documentation demonstration cleanup'))
    if (context.delivery === 'events') {
      // --events streams updates; wrap interaction data in an SDK handle for user handling.
      for await (const event of run.events({ signal: AbortSignal.any([context.abort.signal, AbortSignal.timeout(context.timeoutMs)]) })) {
        trace(event)
        if (event.type === 'screenshot') renderScreenshot(context, event.image)
        if (event.type === 'interaction') await handleInteraction(context, run.interactionHandle(event.data))
      }
      result = await settleRun(context, run)
    } else {
      // Callback mode receives this run's events inside wait; helpers save images and ask the user.
      result = await settleRun(context, run, { onEvent: trace,
        onScreenshot: (image, index) => renderScreenshot(context, image, index),
        onInteraction: interaction => handleInteraction(context, interaction) })
    }
    save(context, 'result.json', result)
    settled(result)
    await saveArtifacts(context, result)
    if (scenario.id === 'landing-page-audit-agent' && context.screenshots === 0) throw new Error('The landing-page audit did not deliver the requested screenshot')
    // The app checks required fields and source URL formats; a reviewer still checks facts.
    const items = validateItems(result.output, scenario.fields, scenario.sourceField)
    // Scenarios that require one item per input entry are checked against the input.
    checkInputCoverage(scenario.id, items, input)
    const report = { scenario: scenario.id, dataset: input.dataset, passed: true, runId: run.id,
      status: result.status, items, audit, screenshots: context.screenshots, interactions: context.interactions,
      events: context.eventCounts, artifactIds: result.artifacts.map(artifact => artifact.id) }
    save(context, 'report.json', report)
    return report
  })
}

async function runMonitor(context: Context, scenario: Scenario, input: JsonObject, token: string, audit: JsonObject) {
  // Track creates a monitor with targets, extraction fields, and hourly scheduling.
  const monitor = await context.qoni.track.create({ token, prompt: scenario.task,
    targetUrls: input.pages, extractionSchema: { heading: 'string', source_url: 'string' },
    tickInstructions: `Open the target URLs and read the actual visible heading. Return a JSON object with heading and source_url. ${scenario.task}`,
    triggerDsl: { on: 'change' }, schedule: { kind: 'interval', intervalSeconds: 3600 } })
  save(context, 'monitor-ref.json', { id: monitor.id, audit })
  return withCleanup(context, async register => {
    register('Track monitor', () => monitor.delete())
    const definition = await monitor.get()
    save(context, 'monitor-definition.json', definition)
    if (object(definition.schedule).intervalSeconds !== 3600) throw new Error('Track did not persist the requested schedule interval')
    // Run one tick and inspect its extraction by runId; completed alone does not prove success.
    const tick = await monitor.runNow()
    save(context, 'tick.json', tick)
    if (tick.state !== 'completed') throw new Error(`Track execution failed: ${tick.state} / ${tick.error ?? ''}`)
    const runId = tick.runId
    if (typeof runId !== 'string') throw new Error('Track tick did not return a runId')
    const detail = await monitor.run(runId)
    save(context, 'tick-detail.json', detail)
    if (detail.state !== 'completed' || !detail.extracted || !Object.keys(object(detail.extracted)).length) {
      throw new Error('Track did not extract page data')
    }
    const extracted = object(detail.extracted)
    if (typeof extracted.heading !== 'string' || !extracted.heading.trim() ||
      typeof extracted.source_url !== 'string' || !/^https:\/\//.test(extracted.source_url)) {
      throw new Error('Track extraction is missing a heading or source URL')
    }
    const normalizeUrl = (value: string) => { const url = new URL(value); url.hash = ''; return url.href.replace(/\/$/, '') }
    if (!(input.pages as string[]).some(url => normalizeUrl(url) === normalizeUrl(String(extracted.source_url)))) {
      throw new Error('The Track source URL is not a configured target')
    }
    // Check persisted pause/resume state; withCleanup() deletes the monitor on exit.
    await monitor.pause()
    if ((await monitor.get()).status !== 'paused') throw new Error('Track did not persist the paused state')
    await monitor.resume()
    if ((await monitor.get()).status !== 'active') throw new Error('Track did not persist the active state')
    const report = { scenario: scenario.id, dataset: input.dataset, passed: true, monitorId: monitor.id,
      runId, state: detail.state, outcome: detail.outcome, extracted: detail.extracted, audit }
    save(context, 'report.json', report)
    return report
  })
}

export function inputFile(): JsonObject | undefined {
  const index = process.argv.indexOf('--input')
  return index >= 0 ? object(JSON.parse(readFileSync(process.argv[index + 1], 'utf8'))) : undefined
}

Results are written to output/recruiting-sourcing-agent/report.json. report.items contains names, evidence, and sources for recruiting review, with fields name, evidence, sourceUrl. audit links the grant ID, audit ID, and effective scopes; http.json records redacted request statuses. The application parses DoAnything output and checks required fields and source URL formats. A business reviewer still assesses the content against the original sources.

Data and memory boundaries ​

This scenario touches four kinds of data; none of them belongs in GUMem:

  • Recruiting rules: role requirements — approved versions only, managed in the ATS or a policy store, with every match reason citing the criteria version. Past interview outcomes must not serve as a candidate ranking signal, nor be written back as "screening preferences".
  • Business state: sourcing lists, candidate summaries, and review annotations — archived in the ATS, retained and used in line with local recruiting and personal-data regulations.
  • Audit records: the delegation and behavior chain formed by grantId and auditId — maintained by GenAuth, attributing every search to a specific role and task.
  • User Memory: not used in this scenario. Candidate evaluation must not depend on personalized screening memory accumulated across sessions — that is exactly where bias hardens; criteria changes go only through the policy store's human revision review.

Failure handling ​

SituationRecommended handling
A candidate page shows a login wall or CAPTCHASkip the source and record it; escalate to the recruiter to decide on a manual look.
Public material cannot support a match judgmentState the insufficient evidence honestly; never guess a candidate's background.
An output entry lacks a source or criteria versionApp-side validation drops the entry and the list notes how many were dropped.
A sensitive attribute or unrelated signal appears in a summaryRemove the field and record the event for bias-detection review.

Production notes ​

Candidate ranking uses approved role criteria only; past interview outcomes must not serve as a ranking signal. Sensitive attributes (age, family status, ethnicity, and so on) never enter automated scoring, and proxy variables (school, region, name style — signals that can indirectly stand in for protected attributes) plus group-level screening skew must be checked periodically, with findings feeding the role-criteria revision review. Candidate data is used only for sourcing this role, retained and used in line with local recruiting and personal-data regulations. The Agent never messages candidates automatically — outreach is always sent by a human after the recruiter approves the draft.

Next steps ​

  • Read the Quickstart to run the shortest path for Agent identity and delegation.
  • Read WebSearch for how public profile search works.
  • Continue with the Sales lead agent for an adjacent scenario.