Skip to content

Quickstart: let an Agent place an order ​

This page follows one scenario — finding the lowest price on Amazon and placing the order for the user — to show how a consumer app calls the Qoni SDK in order: create an Agent with its own email address for a user who is already signed in, bind the two, delegate limited access, then let the Agent shop through Web Agent. Whenever it needs the user, the Agent pauses and raises an interaction, handing the decision back to the user.

Toward the agentic web ​

Qoni's goal is to make agents first-class citizens of the web. Models think, the web is where agents act, and Qoni is the layer in between. The diagram shows the layers: a person delegates scoped, time-bound authority to the agent, and every step is audited; the agent works on web pages for the person in hosted cloud browsers, reusing sign-ins across tasks; when it is done, it hands the results back to the person.

Toward the agentic webA layered model: a person delegates scoped, audited authority to the agent, which works on the web in hosted cloud browsers and hands the results back.HTMLWebAgentHumanDelegate to agentEvery step auditedResults to humanHosted cloud browsersNo API neededSign-ins carry overRuns tasks in parallel

Qoni is managed infrastructure for personal agents. It gives an agent the three things it needs — identity, action and memory: GenAuth gives the agent its own identity, bound to a person, with grants that are scoped, time-bound and audited at every step; Web Agent lets the agent work on real websites in hosted cloud browsers, no API needed; GUMem remembers what the user said and did, so nothing has to be taught twice. All three come through one Qoni SDK.

Scenario ​

In your app, the user tells their shopping assistant:

Open Amazon, search for the MacBook Neo 512GB, find the lowest price, and place the order for me.

Turning that sentence into a real order takes four things:

What the Agent needsWhyStep
Its own Agent IdentityIts own identity and email address, so it can receive codes and send notifications without borrowing the user's accounts1
A binding between themProof that this Agent belongs to this user and nobody else2
A delegation tokenLimits what the Agent may read and do, and for how long3
Execution with the user in the loopBrowse, compare, check out — and ask the user to sign in, answer, take over, fill in details or confirm when needed4

The user's own Identity is not in this table: the user has already signed up and signed in to your app through GenAuth. See Sign users up and in.

How it fits together ​

The diagram shows the overall logic: the top half is onboarding, which runs once per user; the bottom half is the task flow that runs for every order.

OnboardingOnce per user · redo step 3 when the token expires
  1. GenAuthUser signed inSigned up or in on GenAuth’s hosted page
  2. 1GenAuthCreate the Agent IdentityWith an inbox the user names
  3. 2GenAuthBind the user and the AgentProved by the user’s session
  4. 3GenAuthDelegate scoped accessObtained silently on your server
grant.tokencarries the user’s consent
Every orderOnce per task

Onboarding does the minimum: for a signed-in user, your app creates their Agent, binds it, and obtains a delegation token. When the token expires, repeat step 3 only. The shipping address and card are not collected up front: at the first checkout, the Agent asks the user to fill them in and saves them to the Profile for later orders.

Before you start ​

  • Node.js 20.11 or later.
  • A Qoni AccessKey ID and Secret. The AccessKey's permission policy must allow every scope requested on this page.
  • The user has signed up and signed in to your app through GenAuth, and your server keeps their user ID and session (Access Token). See Sign users up and in.
  • A way to put a request in front of the user, such as an in-app dialog or a push notification. Step 4 uses it to handle interactions.

Install the SDK and create the client on your app server:

bash
npm install @qoniai/qoni
ts
import { Qoni, type InteractionHandle } from '@qoniai/qoni'

// Create the client on your app server only; never ship the AccessKey to a browser or phone
const qoni = new Qoni({
  accessKey: process.env.QONI_ACCESS_KEY!,
  secretKey: process.env.QONI_SECRET_KEY!,
})

// The user has signed in through GenAuth: read their user ID and Access Token from the session you stored
const userId = '<genauth-user-id>'
const userAccessToken = '<user-access-token>'

Integration guide ​

1. Create the Agent Identity ​

This step creates the user's own shopping assistant Agent. Qoni assigns it a dedicated email address whose name the user chooses: when a site sends a verification code there, the Agent reads it itself, and after the order it uses the same inbox to send the user an order summary.

ts
const agent = await qoni.genauth.agents.create({
  displayName: "Alex's shopping assistant",
  description: 'Compares prices, places orders and sends order updates for Alex',
  channels: {
    email: '<mailbox-name-chosen-by-the-user>', // for example alex-shopper; Qoni adds the domain and creates an inbox that can send and receive
  },
})
console.log(agent.id, agent.email)
  • agent.email belongs to the Agent, not to the user. When the Agent signs up on a site or receives a code, it uses its own identity, so the user's personal email stays private.
  • The user enters the mailbox name in your app. If the name is taken, creation fails and your app should ask the user for another one.
  • Whether the Agent may read from or send with the inbox is decided by the scopes in step 3. Creating an Agent grants no permissions by itself.

Checkpoint: you have agent.id, and agent.email combines the user's chosen mailbox name with the domain Qoni assigns.

2. Bind the user and the Agent ​

This step declares that the Agent belongs to this user. Binding uses the Access Token the user obtained when signing in to GenAuth, and GenAuth identifies the user from it.

ts
const binding = await qoni.genauth.agents.bindUser({
  agentId: agent.id,
  userAccessToken, // proves the user is present
  relation: 'owner',
})
console.log(binding.bindingId, binding.userId === userId)
  • Binding requires the user's session, not a bare user ID. Your app therefore cannot bind someone else's Identity to an Agent while that person is away.
  • An Agent has exactly one owner. Only a bound Agent can request this user's delegation in step 3.

Checkpoint: binding.userId equals the user's GenAuth user ID.

3. Delegate scoped access ​

This step gets the Agent a delegation token for this user. The Agent was bound to the user with the user's own session in step 2, so your app gets the token silently on the server, with no redirect to a consent page.

ts
const grant = await qoni.genauth.delegateAgent({
  mode: 'silent',
  userId, // the user's GenAuth user ID
  agentId: agent.id,
  scopes: ['*'], // request every scope the AccessKey policy allows
  expiresIn: '30m',
})
console.log(grant.grantId, grant.auditId)
  • The example uses * to request every scope the current AccessKey policy allows. This scenario needs reading and completing the user's Profile, filling the card, reading and sending from the Agent's inbox, running DoAnything and site sign-in. In production, request only the scopes the task needs. For every scope you can request and what each one allows, see Qoni SDK: requestable scopes.
  • Silent delegation skips the consent page and returns grant.token directly. It requires the Agent to be bound to this user in step 2: only a bound Agent can get this user's delegation, and userId must be the user's own GenAuth user ID, never someone else's.
  • * expands to every scope the current AccessKey policy allows. When you need to record the scopes actually granted in your audit log, call qoni.genauth.introspectDelegationToken({ token: grant.token }); running the task does not need it.
  • If the AccessKey policy does not allow user.payment:use, the Agent never fills the card: at the payment page it asks the user to take over the browser and pay themselves.

Checkpoint:

  • You have grant.token, valid for 30 minutes.
  • Your server recorded grant.grantId and grant.auditId for later audit.

4. Run DoAnything and handle interactions ​

This step hands the user's sentence to Web Agent. DoAnything opens Amazon in a controlled browser, searches, compares and checks out. Whenever it needs the user, it pauses and raises an interaction, and waits for your app to submit the user's decision.

Interactions come in five types based on what the user has to do, independent of any particular business:

TypeMeaningSDK methods
site_loginSign-in: the user signs in to a site themselves in the controlled browseropenLogin(), confirmSignedIn()
take_controlHuman takeover: anything other than sign-in that needs the user's own hands on the browser, such as a CAPTCHAconnectControl(), refreshControl(), releaseControl()
ask_userAsk the user: one question, with a structured answer described by answerType and optionsanswer(), skip()
confirmationConfirmation: the only outcomes are allow or denyconfirm(), reject()
fill_formFill in details: a generic form rendered from fields that collects structured datasubmit(), skip()

In this scenario, all five types happen in order; filling in details only happens at the first order:

  1. AgentOpens amazon.com in a controlled browser
  2. site_loginSign-inWaits for the user
    Sign in to Amazon
    The user signs in to the site themselves in the live browser. Neither the Agent nor your app sees the password.
    SDK methodshandle.openLogin()handle.confirmSignedIn()
  3. AgentSearches “MacBook Neo 512GB” and finds it comes in two colors
  4. ask_userAsk the userWaits for the user
    Pick a color
    One question with a structured answer. Here answerType is single_choice and options lists both colors.
    SDK methodshandle.answer()handle.skip()
  5. AgentKeeps the chosen color, compares item + shipping, adds the cheapest to the cart and hits a CAPTCHA at checkout
  6. take_controlHuman takeoverWaits for the user
    Solve a CAPTCHA
    Anything other than sign-in that needs the user’s own hands on the browser. Codes sent by email go to the Agent’s inbox and need no takeover.
    SDK methodshandle.connectControl()handle.releaseControl()
  7. AgentReaches checkout and finds no saved shipping address or card in the Profile
  8. fill_formFill in detailsWaits for the userFirst order only
    Fill in the address and card
    A generic form rendered from fields. Fields with a profileField are saved to the Profile, so later orders do not ask again.
    SDK methodshandle.submit()handle.skip()
  9. AgentSaves the details to the user’s Profile, fills in the address and prepares the order summary
  10. confirmationConfirmationOrder gate
    Confirm the payment and order
    Allow or deny, nothing else. Only after the user allows it does the Agent fill the card and click “Place your order”.
    SDK methodshandle.confirm()handle.reject()
  11. AgentFetches the card from GenAuth and fills it in, places the order and emails the summary from the Agent’s own inbox

Start the task with the delegation token:

ts
const run = await qoni.doAnything.run({
  token: grant.token,
  prompt: `
    Open https://www.amazon.com and search for "MacBook Neo 512GB".
    If the model comes in several colors, ask me which one I want first.
    Compare only listings with exactly that model, storage and color, and pick the lowest item + shipping total.
    Check out with the shipping address and card in my Profile; ask me to fill them in if they are missing.
    Before placing the order, ask me to confirm the item, total and payment method.
    After the order is placed, email the order number, total and delivery estimate to me from your inbox.
    Return only JSON: { orderId, title, color, seller, totalPrice, currency, deliveryEstimate }.
  `,
  capture: { screenshots: true },
})

Then handle interactions by type in run.wait(). The handler's skeleton works for any business: five branches for the five types, each listing the SDK methods it uses. The comments use this scenario as the example: In this scenario shows what this run's payload looks like and how often it occurs, and → marks what your app does, such as drawing UI or waiting for the user. For your own business, keep the skeleton and replace what those two kinds of comments describe.

ts
const handled = new Set<string>() // IDs of interactions already handled

async function handleInteraction(handle: InteractionHandle) {
  const request = handle.interaction
  // One interaction triggers the callback on creation, on status updates and on event replay:
  // handle only pending ones, once per ID
  if (request.status !== 'pending' || handled.has(request.id)) return
  handled.add(request.id)

  // Branches follow the order they occur in this scenario
  switch (request.type) {
    // Sign-in: the user signs in to the sites in request.payload.sites themselves
    // In this scenario: once per site that needs sign-in, here only Amazon; once the sign-in is saved,
    //   later tasks skip it. sites is
    //   [{ siteId: 'amazon', displayName: 'Amazon', loginUrl: 'https://www.amazon.com/ap/signin' }]
    case 'site_login': {
      const login = await handle.openLogin() // open the controlled sign-in browser; returns the live view URL login.liveUrl
      // → open login.liveUrl in your app; the user signs in to Amazon there, then taps "I'm signed in"
      await handle.confirmSignedIn() // call after "I'm signed in": DoAnything saves the sign-in, re-checks and goes on
      break
    }

    // Ask the user: request.payload.question is the question; answerType and options describe the answer
    // In this scenario: once per choice the Agent is unsure about. Here it is the color:
    //   question is "Which color do you want?", answerType is 'single_choice', options is
    //   [{ value: 'space-gray', label: 'Space Gray' }, { value: 'silver', label: 'Silver' }]
    case 'ask_user': {
      // → your app draws a single-choice control; the user picks Space Gray and you submit its value
      await handle.answer('space-gray')
      // await handle.skip() // the user declines to choose; the Agent continues on its own judgment
      break
    }

    // Human takeover: the user's own hands are needed on the browser; request.payload.reason says why
    // In this scenario: once per check that only a person can pass. Here it is an image CAPTCHA
    //   at checkout, and reason is 'captcha_required'
    case 'take_control': {
      await handle.connectControl() // take over the browser
      // → open request.payload.liveUrl in your app; the user solves the CAPTCHA by hand
      // await handle.refreshControl() // refresh when the live view expires
      // call after the user taps "Done": hand the browser back to the Agent, which goes on checking out
      await handle.releaseControl()
      break
    }

    // Fill in details: render a form from request.payload.fields and submit values keyed by field.name
    // In this scenario: only when the Profile lacks checkout details, usually on the first order. fields is
    //   [{ name: 'shippingAddress', label: 'Shipping address', type: 'address',
    //      required: true, profileField: 'shippingAddress' },
    //    { name: 'paymentCard', label: 'Card', type: 'payment_card',
    //      required: true, sensitive: true, profileField: 'paymentCard' }]
    case 'fill_form': {
      // → your app renders an address input and a card input from fields, and submits once they are filled in
      // fields with a profileField are saved to the user's Profile, so later orders do not ask again
      await handle.submit({
        shippingAddress: {
          line1: '500 Howard St', city: 'San Francisco', state: 'CA', postalCode: '94105', country: 'US',
        },
        paymentCard: { number: '<card-number>', expMonth: 12, expYear: 2029, cvc: '<cvc>', holderName: 'Alex Chen' },
      })
      // await handle.skip() // the user gives up: the Agent does not continue to checkout
      break
    }

    // Confirmation: request.payload.summary describes what the user is asked to allow; allow or deny only
    // In this scenario: once before each order; if the price changes after that, the Agent asks again.
    //   summary is "Pay $1,299.00 at Amazon.com with Visa •••• 4242 for Apple MacBook Neo 512GB in Space Gray"
    case 'confirmation': {
      // → your app shows the summary with "Allow" and "Deny" buttons
      await handle.confirm() // the user allows it: the Agent fills the card from GenAuth, then orders
      // await handle.reject() // the user denies it: the Agent does not order and the task ends
      break
    }
  }
}

const order = await run.wait({ onInteraction: handleInteraction })
console.log(order.status, order.output)

The same type can occur more than once:

  • Every interaction has its own request.id and payload, and one type can occur several times in a task: one site_login per site that needs sign-in, one ask_user per choice the Agent is unsure about, one take_control per image CAPTCHA. The color question and the CAPTCHA here are just one example; with another product, the Agent may also ask about the configuration or the seller, or hit a CAPTCHA while signing in.
  • Interactions come one at a time: the task pauses at each one, and the Agent continues — and the next interaction appears — only after the user has handled it.
  • One interaction also triggers the callback more than once: its creation, its status moving from pending to active and resolved, and event replay after a reconnect all call onInteraction again. That is why the handler only handles the pending status and remembers each handled request.id, so nothing pops up or gets submitted twice. Your UI should likewise update the same card by request.id instead of creating a new one per callback.

A few notes:

  • Passwords and card numbers never go through the Agent: sign-in and CAPTCHAs happen in the live browser, by the user's own hand. In fill_form the card is a sensitive field that GenAuth stores encrypted, and every later read returns only the brand and last four digits; the card number passes through your server only in this one submit() request, so keep it out of your database, logs and error reports. Once the user allows the confirmation, the Web Agent runtime uses user.payment:use to fetch the card from GenAuth and fills it straight into the payment form in the controlled browser; the card number and CVC never enter the model context, events, screenshots or logs.
  • Saving to the Profile needs consent: writing profileField fields back to the Profile requires user.profile:write in the delegation. When the same user orders again, no fill_form appears.
  • Not every code needs a takeover: when a site sends a code to the Agent's inbox, the Agent reads it itself with agent.mail:read and raises no take_control; only challenges that need a person, such as image CAPTCHAs, trigger a takeover.
  • Waiting and expiry: interactions wait on the user and may take minutes. The handler should call each method only after the user has acted. When an interaction expires, the SDK method throws, and the Agent decides whether to retry based on the task description.

Checkpoint ​

When the task ends, confirm that:

  • order.status is succeeded, and order.output is the order for the lowest-priced listing in the color the user chose:

    json
    {
      "orderId": "112-4839201-5528261",
      "title": "Apple MacBook Neo 512GB",
      "color": "Space Gray",
      "seller": "Amazon.com",
      "totalPrice": 1299,
      "currency": "USD",
      "deliveryEstimate": "2026-10-09"
    }
  • The user's inbox received an order summary sent from agent.email.

  • In step 4, the user handled sign-in, the question, the takeover, filling in details and the confirmation in that order; denying any of them stops the task at that point.

  • When the same user orders again, no fill-in interaction appears: the Agent uses the address and card already in the Profile.

  • Your server stored grant.grantId, grant.auditId and order.runId, so the Qoni Console can trace this order back to the user's delegation.

  • The full card number appears nowhere in your app logs, task events, screenshots or artifacts.

Next steps ​