Adept
PaidAgentic AI for your enterprise tech stack.
Adept is an agentic AI company building enterprise agents that operate across SaaS applications and browsers using a multimodal action model. After Amazon hired the original co-founders into its AGI team in June 2024, the remaining team continued under CEO Zach Brock with a focus on agentic AI solutions for enterprise tech stacks.

What is it
Adept builds AI agents that take actions across enterprise software and the web by understanding UIs visually rather than relying on APIs. Its full-stack approach combines proprietary agent training data (trillions of tokens of real software usage), a suite of multimodal models, a custom domain-specific actuation layer, and tooling for ongoing feedback-driven model improvement.
What it can do
Adept's models β including ACT-1 (Action Transformer, 2022) and the open-sourced Fuyu-8B multimodal model (2023) β can locate elements on a webpage or application, reason about documents, PDFs, charts, and tables, and plan end-to-end enterprise workflows. On Adept's internal evaluations the system reports scores of Locate 93, Web VQA 88.2, and Planning 88 (vs. GPT-4 at 59). Demonstrated tasks include condensing 10+ Salesforce clicks into a single sentence, working inside spreadsheets, and chaining actions across multiple programs.
Who is it for
Adept is sold to enterprises that want to automate cross-application, manual workflows in tools like Salesforce, SAP, Workday, and custom web apps. The product is in private beta β there is no self-serve signup β and the team explicitly targets supply-chain, financial services, and healthcare operations use cases.
Key Features
ACT-1 Action Transformer β Browser-Native Task Execution
ACT-1 is Adept's foundational Action Transformer, a large-scale model trained to use digital tools through a Chrome extension that observes the browser viewport and emits actions (click, type, scroll). It takes high-level natural-language commands and executes them across UI elements, demonstrated on workflows that would otherwise take 10+ clicks inside tools like Salesforce.
Fuyu-8B Multimodal Model β Decoder-Only Vision for UIs
Fuyu-8B is Adept's open-sourced (CC-BY-NC) multimodal model that powers UI understanding. It is a vanilla decoder-only transformer with no separate image encoder, supports arbitrary image resolutions, and returns answers on large images in under 100 milliseconds β designed from the ground up for screen, chart, diagram, and document understanding rather than natural-photo benchmarks.
Locate β Pixel-Level Element Targeting
Adept's agents accurately locate items on a webpage or application β buttons, links, text fields, table cells β at the pixel level rather than relying on DOM selectors or APIs. On Adept's internal eval, Locate scores 93, enabling reliable action execution across software with incomplete or missing programmatic access.
Web VQA β Reasoning over Documents, PDFs, and Charts
Agents reason about and answer questions over websites, documents, PDFs, charts, graphs, and tables, scoring 88.2 on Adept's internal Web VQA eval. This powers downstream tasks like contract review, license-application processing, and pulling structured data out of unstructured source materials.
Use Cases
Supply-chain shipping availability checks across hundreds of sites
Run cross-site queries to assemble accurate delivery plans, replacing a manual sweep across vendor portals with a single agent invocation.
Financial services PDF & contract data extraction with system updates
Pull key information out of PDFs and contracts, update internal systems with the extracted fields, and send a notification email when the work is complete β all without integration code.
Healthcare license application processing with human-in-the-loop
Process license applications against defined business logic and route the final submission to a human for sign-off, combining throughput with regulatory accountability.
Salesforce data updates from natural-language commands
Replace 10+ click sequences inside Salesforce with a single natural-language sentence, demonstrated on the ACT-1 capability preview.
Pricing plans
Frequently Asked Questions
Discussion
No comments yet. Be the first to start the thread.