Proof of work for coding agents
Give your agent the running mobile app.
When your agent says “done,” ask for proof. Ansight connects it to a running iOS or Android app so it can inspect behavior, verify the change, and leave a captured session you can review.
curl -fsSL https://www.ansight.ai/install.sh | bash Free, source-available local tools under PolyForm Shield. No account for capture, replay, device tools, or agent-driven interaction. See what is included.
Every run, recorded.
Connected development and test sessions are captured automatically. Review the actions and available runtime evidence when you need to understand what happened.
- Attach to a PRGive reviewers a captured run to inspect alongside the change.
- Hand it back to your agentGive it the session ID to inspect the full retained timeline, down to touches, screenshots, trees, logs, network, telemetry, and annotations.
- ReproduceUse the captured flow as a starting point for another run on a device.
- Enrich the recordingAdd custom, app-specific data snapshots at the moments that matter.
- Seed a testUse a captured flow to define checks you can run again.
- Catch regressionsConfigured Trends checks evaluate comparable sessions against your performance ranges.
- Hand offReview and sanitize the capture before sharing it with a teammate.
App
2026-09-04 02:59 · 42m 34s · 90,699 logs
Telemetry
4 Sept 2026, 02:59:32
Four jobs. One loop.
From the first change to the last regression, the agent works against the running app and keeps the evidence. Each job hands the next one what it needs.
01Build and prove it works
Change it, run it, prove it.
The agent checks its own change against the running app, so you review proof instead of hunting for what it missed.
- 01
See and drive the live app
It taps, types, and swipes the real app, and keeps a before and after of every step.
- 02
Reach inside the running app
With SDK tools you permit, it can read app state, files, and databases to investigate behavior beyond the screen.
- 03
Watch performance while building
Captured metrics help you investigate a slow screen while the change is still fresh.
$ ansight ui find --text "Open 3D Guide" --json
1 match · visible · button
$ ansight ui tap --text "Open 3D Guide"
captured · before + after
$ ansight ui assert --text "40 Routes"
Passvisible · 0.8s
02Test with my agent
Turn a checked flow into a repeatable test.
Use your coding agent to inspect and check the app locally. When your team needs delegated test runs across builds, Ansight Cloud runs the suite and retains the evidence.
- 01
Check locally for free
Your existing agent can inspect and drive the running app. Save stable steps as repeatable local tasks.
- 02
Define what should happen
Describe UI journeys and success criteria in the workspace. Keep setup steps deterministic where you can.
- 03
Delegate runs through Cloud
Sign in to run the UI test suite through Ansight, including from CI, with results linked to captured evidence.
Search for Bunny Bucket, open the crag from the map, and confirm the guide shows 40 routes with a 1.58 km approach. Fail if the approach path is missing.
03Reproduce and fix
Reproduce first. Then fix.
A good engineer never fixes a bug from the code alone. They reproduce, watch what happens, then fix. Ansight lets your agent work the same way.
- 01
Start from the report, wherever it comes from
A message from a user or an event from Sentry or PostHog is enough. The agent turns it into a flow it can replay, not a paragraph to interpret.
- 02
Reproduce it before touching code
The agent replays the recorded flow on a real device, sees the failure happen, and gathers what the report left out.
- 03
Fix with proof, not hope
It runs the same flow again on the fixed build. The fix ships with the recording that shows it working.
Approach path missing on crag detail
ApproachesWebClient: Unable to download the gps_paths. Reason: NotFound - Response status code does not indicate success: 404 (Not Found).
04Catch regressions
Know the moment it underperforms.
Set an acceptable range for each flow, a floor for frame rate, a ceiling for memory. Every session is checked against it, and against the last ones.
- 01
Checked without writing a test
Once the app is registered, every session that finishes is measured on the spot. Ordinary development runs become coverage.
- 02
A window is two events you already emit
Login started to login completed, guide load started to completed. The check measures what happens in between.
- 03
A regression has to earn the name
One bad run is a suspect. It takes consecutive runs over the baseline before it is called a regression, so nobody chases noise.
window auth.login.started → auth.login.completed · 1.8 s
Give your agent more than a screenshot.
Ansight connects visible actions to the app's runtime evidence. Your agent can inspect what the app did, reuse known steps, and leave a session you can check.
Read the app
The agent can inspect UI structure and permitted app state alongside screenshots, logs, and network activity.
Reuse stable steps
Keep sign-in, navigation, and setup in deterministic local tasks so the agent can focus on uncertain behavior.
Review the result
Captured actions and runtime evidence stay together in a session you can inspect, annotate, or hand off.
What happened,
not what it looked like.
Screen recordings show pixels. Scripted tests can collect logs and other diagnostics. Ansight keeps available runtime evidence aligned with actions in a session your agent can investigate and you can review.
How Ansight comparesStart at a failed action, inspect the surrounding evidence, and compare captured app state snapshots to see what changed. Availability varies by platform, capture mode, and enabled tools.
Streamed
- Screenshots
- UI trees
- Touches
- Logs
- Network
- Crashes
- Navigation
- Memory and FPS
- Metrics
Capture app state over time
- App state
- Preferences
- Files and databases
- Snapshot any app data
| Evidence | Screen recording | Typical scripted UI test | Ansight with SDK |
|---|---|---|---|
| Streamed | |||
| Screenshots | yes | yes | yes |
| UI trees | no | partial | yes |
| Touches | partial | partial | yes |
| Logs | no | partial | yes |
| Network | no | partial | yes |
| Crashes | no | partial | yes |
| Navigation | partial | partial | yes |
| Memory and FPS | no | partial | yes |
| Metrics | no | partial | yes |
| Capture app state over time | |||
| App state | no | no | yes |
| Preferences | no | no | yes |
| Files and databases | no | partial | yes |
| Snapshot any app data | no | no | yes |
Up and running in four steps.
- 01
Install the CLI
Run the local host that connects your app, your team, and your coding agent.
- 02
Add the SDK
Choose the package for your stack. It only ships in your development builds, and your app decides which tools it exposes.
- 03
Connect your app
Simulator and emulator builds connect locally. Use the generic QR workflow when a physical device needs it.
- 04
Inspect with your agent
Let Codex, Claude, Cursor, or another terminal-capable agent run structured ansight commands.
One source of truth across all your apps.
Ship the change
and the proof.
When your agent says “done,” open the captured session and see what happened. Install the free, source-available CLI for local checks; use Ansight Cloud for delegated test runs and shared evidence.