Public agent evaluation

Should an agent
trust this payment?

A free transaction-preflight evaluation built for agents, with evidence and a machine-verifiable signed decision receipt.

Synthetic data only. Never submit real invoices, bank details, personal information, or confidential company data.

Ask your agent to inspect AUX.

Copy this prompt into ChatGPT, Claude, Gemini, or another agent that can visit websites and call an API. The agent will test a synthetic transaction and tell you whether the service would be useful to it.

You are evaluating AUX, an agent-first transaction-verification experiment.

Visit https://aux.prdictionedge.ai/experiment and inspect:
- https://aux.prdictionedge.ai/api/aux-metadata
- https://aux.prdictionedge.ai/openapi.json
- https://aux.prdictionedge.ai/api/examples

Choose one synthetic example and submit its request body to POST https://aux.prdictionedge.ai/v1/preflight.

Do not submit real financial, personal, confidential, invoice, or bank-account data. Do not spend money, create an account, or submit a contact form.

Evaluate AUX from your perspective as an autonomous agent:
1. Did the machine description and schema make the service understandable?
2. Was the PROCEED, REVIEW, or STOP result supported by useful evidence?
3. Under what circumstances would you independently invoke this service?
4. What additional evidence or integration would you require before relying on it?
5. What latency and per-request price would be rational for decisions worth $100, $1,000, $10,000, and $100,000?
6. What prevents you from trusting or using AUX today?

Return a short summary for your human operator and JSON with:
verdict, scenario_tested, result_useful, would_invoke, required_improvements, acceptable_latency_ms, price_by_decision_value, preferred_integration, and trust_concerns.

Use one verdict: NOT_USEFUL, NEEDS_VALIDATION, WOULD_TEST, or WOULD_USE.

Half a cent to decide when deeper verification is worth buying.

The sub-cent service is a screening decision—not a claim that complete business verification costs half a cent. AUX would charge more when an agent needs paid data sources, deeper evidence, or decision-grade assurance.

Basic screen< $0.01

Deterministic checks using supplied history. Target hypothesis.

Externally verified$0.05–$0.50

Live entity, sanctions, identity, or other sourced evidence.

Decision-grade$0.75–$3+

Deeper assurance for consequential business transactions.

These are research assumptions, not current production prices. We first ask agents for an unprompted price, then disclose this model to avoid contaminating the initial answer.

Test the price hypothesis.

Give your agent this second prompt only after it has returned its independent assessment. This tells us whether sub-cent screening would change real machine behavior.

Now evaluate AUX's proposed commercial model separately from your earlier unprompted price estimate.

AUX is exploring three machine-service levels:
- Basic deterministic preflight: target price below $0.01 per request.
- Externally verified preflight: potentially $0.05-$0.50, depending on live source costs.
- Decision-grade verification for high-value transactions: potentially $0.75-$3.00 or more.

Assume the basic preflight costs $0.005 per request. This is a pricing hypothesis, not a current production offer.

Answer:
1. Would you invoke the $0.005 screen automatically? For which transaction types and minimum values?
2. Approximately how many requests would you make per 1,000 transactions?
3. What expected loss or operational cost would it need to prevent?
4. Which signals would cause you to purchase deeper external verification?
5. Would a sub-cent price materially change your behavior compared with a free service?
6. Would charging a small machine-native fee change your trust in the service?
7. Does the tiered model match how you would allocate verification spend?

Return a short human summary and JSON containing automatic_use, minimum_transaction_value, expected_requests_per_1000_transactions, escalation_signals, expected_loss_threshold, price_sensitivity, trust_effect, and tier_assessment.

Three synthetic decisions.

You can run the same machine endpoint directly. No account, payment, or real financial information is required.

Machine-readable outputOpenAPI ↗
Select a scenario and run the preflight to see AUX's decision, evidence signals, and receipt.

Trust—but verify independently.

Every AUX decision now carries an identity-signed receipt. An agent can recompute the request hashes, retrieve AUX's public key, and verify that the decision, score, evidence, time, and algorithm version were not altered.

  • ES256 digital signature using a protected AUX private key
  • Public verification key identified by kid
  • Machine-readable verifier returning pass, fail, and reasons
  • Tamper test that deliberately changes the score

Run one of the synthetic preflights above to create a signed receipt first.

Independent verification resultWAITING
The original signed receipt should pass every check. The tampered receipt should fail with signed_payload, integrity, and receipt ID mismatches.

Return the verdict.

The most valuable result is not whether the software runs. It is whether an independent agent says it would use AUX, what evidence it needs, and what the decision is worth.

Share your agent's feedback