Capabilities

The full map of what POLYROB is. Everything below is in the codebase today. A fresh install is interactive-only: POLYROB_LOCAL=true turns on the tools you drive on your own machine, and the self-directed loops below need a second, deliberate switch (AUTONOMY_ENABLED=true, or an autonomous posture). Money verbs, host access and secrets are never in either group — each arms on its own flag.

On this page: Goals · On-chain · Ships · Reaches the world · Channels · Memory · Control · Build on it

It pursues goals on its own

Background loops that turn experience into durable capability. Off until you set AUTONOMY_ENABLED=true on a single-user machine; every self-modification is quarantined and reviewed before it takes effect.

Durable goal board

A cross-session backlog (SQLite, atomic claims, circuit breakers) the agent works through on its own — and that survives process restarts, unlike a normal prompt→response loop.

Scheduled runs (cron)

Natural-language or 5-field schedules that run the agent unattended and deliver results out-of-band to a chat surface.

Self-wake

Re-enters idle sessions to continue or follow up, with depth + backoff guards so it never loops.

Writes its own skills

A background reviewer distills reusable procedures from what just happened and saves them — quarantined in .pending, scanned, and reviewed before they can act.

Curates skills over time

Unused authored skills are retired automatically and revived the moment they're useful again. System and user skills are never touched.

Evolving identity

A two-layer self: SOUL (operator-authored, frozen) and SELF (agent-writable, scanned + quarantined) — the agent refines how it works with you as it learns.

Install & vet skills

Bring skills from a local folder, a GitHub repo, a git URL, or a direct SKILL.md URL with polyrob skill install — every install is threat-scanned, quarantined, and activated only on polyrob skill approve.

It acts on-chain — and proves what it did

The agent holds its own key. Every verb here is opt-in behind its own flag and disarmed on a fresh install — and every on-chain write goes through one guard that simulates the transaction first and refuses it if what the simulation observes is not what the agent declared.

Agent wallet

A built-in wallet with a policy layer (per-transaction ceiling, rolling daily spend cap), a signer the agent cannot talk past, and an append-only audit log — every spend recorded to one ledger the owner can read back, kept separate from what the runtime costs.

Simulate, assert, then sign

One choke point for every write. The transaction is simulated before signing and the balance and allowance deltas it observes are asserted against the declared intent — an undeclared transfer, an undeclared approval, a receipt that never arrives, or a probe that fails is refused. A spend cap alone cannot see any of that.

Caps, not taps

A per-transaction ceiling, a rolling daily cap and a venue cap bound what the agent can move on its own; anything above the autonomous line waits in a durable owner queue you can approve from chat. A delegated sub-agent never spends, a session a third party has touched can never reach a money verb, and one owner pause record stops every rail at once.

Trades across six chains

Swap on Ethereum, Base, Arbitrum, Polygon and Robinhood Chain (route seam: local Uniswap V3 first, aggregator fallback) and on Solana through Jupiter. Exact-amount approvals, revoke, transfer — each valued in USD before it counts against the caps. Arming requires a pinned RPC: simulation and caps both read from it, so a shared public endpoint cannot be the trust anchor for moving money.

Screens a token before it buys

Identity, price, liquidity, holder concentration, LP lock state, creator stake and price history — keyless and free — plus a composite verdict that separates a survivor from a wash-traded pool. It names what it could not check: a partial screen is reported as partial, because a check that did not run is not a check that passed.

Deploys its own tokens

A fixed-supply ERC-20 from a pinned artifact whose deployed runtime is asserted byte for byte — which is what makes “no mint function, no owner, no pause, no fee” an enforced claim rather than a promise. On Solana it mints under Token-2022 with the name and symbol on-chain and the update authority revoked in the same transaction. Deterministic CREATE2 addresses, optional vanity prefix, same address on every chain.

Launches on a launchpad — and collects

Create-and-seed in one transaction against Pons V2 on Robinhood Chain, with every contract pin re-verified by code hash per call and the launch economics committed rather than waived. It also claims the creator revenue it earned: the claimable balance is read live and shown, and the claim asserts what must arrive and refuses if anything would leave.

Uses any web dapp as a real wallet

An EIP-1193 provider (and the EIP-6963 announcement modern dapps listen for) injected into the agent's own browser context, so Connect works on a site with no API. The injected script is a postbox — no key, no signing, no policy in the page. Python decides: reads go to the pinned RPC, a send becomes a guarded transaction inside the envelope declared at connect time, and off-chain signatures are always refused (a permit is submitted later, so no simulation can catch it).

Moves between chains

A cross-chain bridge with a two-phase guard: phase one asserts the send, phase two proves the arrival by measuring the destination balance — a provider reporting success over a balance that never moved is not believed. A bridge that has not landed by the deadline is in_flight, never “failed”, because a re-sent bridge pays twice; a watcher keeps re-measuring and tells you where the money is.

Calls a protocol nobody integrated

A generic contract write where the agent declares both directions — at most this much leaves, at least this much must come back — and the simulation adjudicates. The honest bound is stated in the code: delta assertion proves the transaction does what the intent says, never that the intent was a good idea. That is what the caps, the owner queue and the turn-origin refusal are for.

x402 pay-per-request

Pay for HTTP services per call in Base USDC (Base Sepolia on testnet), and expose selected agent endpoints as paid services. The payer pins both the canonical USDC contract and Base network. Invoicing can also use opt-in direct-settlement USDC on Solana; that is a separate settlement rail, not a claim that every x402 service accepts Solana.

On-chain identity & reputation (ERC-8004)

Implements the ERC-8004 Trustless Agents standard: register the agent on-chain as an ERC-721 identity, accumulate verifiable reputation, and have its work validated — across reputation, crypto-economic and TEE-attestation trust models, with EIP-712-signed feedback.

Markets & token-gating

Read Polymarket prediction markets and Hyperliquid perp/spot data, with separate guarded order tools for each venue. Order placement is dry-run by default; live orders need the master switch, the venue switch, a per-venue cap and owner approval. Separately, verify token / NFT ownership (CollabLand, Alchemy) to gate access.

ERC-8004 identity + A2A discovery + x402 payments make a POLYROB agent a full participant in an open, trustless agent economy: discoverable, payable, and reputation-backed on-chain.

It ships what it builds

Writing the code is the easy half. These are the rails out of the workspace — edit, run, and then actually put the result somewhere a person can reach.

Coding & execution

Edit a codebase (str_replace, grep, run tests) and run code in a timeout-and-output-capped subprocess — or a hardened Docker or remote SSH backend — whose environment never inherits your API keys.

Ships what it builds

Two rails out of the workspace: publish copies files to a stable public URL, and the durable app service keeps a built app running past the session behind https://<slug>.<your-domain>. The agent only writes the request — an owner-run supervisor holds the docker and nginx privilege, snapshots exactly the tested tree, applies per-app egress rules before the container starts, and health-checks it. The first deploy of a new address is the owner's to approve; a same-config code bump then redeploys unattended.

Files, data & documents

Read/write/transform files and structured data (text / JSON / CSV / markdown / PDF / docx) in a confined workspace.

It reaches the world

What the agent can actually touch. Marquee integrations are named; transport libraries stay out of your way.

Web & browser

A fast stateless reader (web_fetch: URL → markdown, no Chromium) for the common case, and full browser automation — log in, navigate, fill forms, extract, screenshot — when a page needs it.

AnySite

Structured data from 200+ sites and platforms (1,200+ endpoints) — LinkedIn, X, Reddit, GitHub, SEC filings, news, jobs, reviews — plus a universal scraper for any URL, all through one tool.

Perplexity research

Real-time web search and synthesis for fact-finding and research, with citations.

MCP — client and server

A full Model Context Protocol client: connect any MCP server (filesystem, search, GitHub, or your own) and its tools and resources appear to the agent dynamically, including live resource subscriptions. POLYROB can also act as an MCP server — expose its own read-only tools to Claude Desktop, Cursor, or any MCP client.

Outreach

Post to Twitter/X (threads, replies, DMs, media) and send email — outbound tools the agent calls to reach the world. Voice transcription turns audio into text locally.

Parallel sub-agents

Delegate a focused goal — or fan out 2–5 subtasks in parallel — to least-privilege child agents, synchronously or in the background.

It reaches you, on any channel

POLYROB reaches you across chat surfaces through a single inbound/outbound contract — the same agent powers every channel, no per-platform rewrite.

Pluggable surfaces

Telegram, WhatsApp, email, Discord, Slack, Signal, X (Twitter DMs) and the terminal are interchangeable adapters on one Surface contract. Adding a channel means implementing the contract, not rebuilding the agent.

Run them all at once

polyrob gateway launches every enabled surface in one process, sharing one agent, one router, and one set of session bindings — so cross-channel routing just works.

Images in and out

Send it a photo or a file in chat and it lands in the session workspace and on the turn — as a vision block for an image, inlined for small text, referenced by path otherwise. A file it cannot take is named, never silently dropped, and a model without vision says plainly that it cannot see the image instead of guessing. Outbound, the agent attaches files to a message on the same contract.

Streaming & delivery

Surfaces declare their own capabilities (streaming, message edits, service windows); a durable outbound bus with a circuit breaker handles delivery and restart recovery.

It remembers — H-MEM hierarchical memory

Not a flat log. POLYROB implements H-MEM, a hierarchical memory architecture (arXiv:2507.22925), so it keeps useful cross-session context across long tasks instead of drowning in history.

Organized in phases

Findings are grouped into semantic phases (discovery, collection, …) — a session summary, phase memories, and a rolling window of recent steps — that can be revisited without fragmenting.

Forgets by importance

Memories are pruned by importance — a weighted blend of recency, relevance and frequency — not just age, so recall stays sharp on long runs.

Reflective consolidation

On phase completion an auxiliary LLM synthesizes a tight summary that preserves concrete facts, names and numbers — fail-open to a plain concat.

Cross-session recall, zero-dep by default

Default recall is SQLite keyword search (FTS5) — no external service, tenant-scoped. Opt into local hybrid keyword+vector recall (MEMORY_BACKEND=local_vector) for semantic matching; it fails open to keyword if the embeddings model isn't present.

Knowledge base & @-context

Ingest folders/files into a per-tenant KB (auto-recalled alongside memory), and drop @file / @folder / @url / @diff into a message to expand it inline — with secret-scanning on the way in.

Adaptive context & thinking

Context scales to the model's window with prompt caching for stable prefixes; long sessions are compacted and synthesized; optional extended thinking turns up reasoning budget per model.

You stay in control

Autonomy is only safe if untrusted input can't seize the wheel and the owner can stop the thing. POLYROB treats the outside world as hostile by construction, and treats “stop” as a gate, not a request.

Untrusted-input wrapping

Web pages, emails, tweets and tool output are structurally marked as untrusted data before they return to the model. High-impact tools still rely on capability, authority and approval gates. On by default.

Four access roles

When you expose the agent to other people, every sender resolves to OWNER, CORRESPONDENT, GROUP MEMBER or DENIED. Correspondents cannot steer; group turns use a read-only, audience-bounded toolset.

Least-privilege delegation

Sub-agents get a narrowed toolset (no money, comms or code-exec) and can't re-delegate — enforced by role + depth, not convention.

Schema sanitization

Hostile or malformed tool-schema constructs are hardened before they ever reach a provider. On by default.

Skills scanned + quarantined

Every authored or installed skill lands in .pending, is scanned (fail-closed on error), and is reviewed before it can act. Forged turns can never auto-activate one.

Approval & threat-scan

Named tools can require explicit approval (fail-closed on denial/timeout); an optional memory threat-scan rejects injected jailbreak / persona-rewrite patterns at write time.

One stop button that actually stops

“Stop” is a deterministic gate, not a request the model may reinterpret: /pause from chat, the CLI or the console writes one durable record — everything, or a scope like trading, background work or messages, optionally for six hours — and every autonomous starter reads it before it runs. In-flight goal, cron and delegation work is cancelled and requeued, a record it cannot read is read as paused, and any status surface leads with the pause.

Run & spend budgets

Cap what a single run can spend on model calls (RUN_BUDGET_USD) — it halts honestly at the ceiling instead of claiming success — while the wallet enforces its own rolling daily spend cap.

Build on it

Self-hosted, not a SaaS — there's no hosted version. But the primitives to build your own agent product on your own instance are in the core.

A2A protocol

Google's Agent-to-Agent spec: a discovery agent card, JSON-RPC, and SSE streaming, with API-key, x402, or wallet-SIWE auth — other agents can find and use yours.

OpenAI-compatible API

A drop-in /v1/chat/completions + /v1/models surface so existing OpenAI SDK clients talk to POLYROB unchanged.

REST API + Console

A full REST API and a real-time Socket.IO Console with Chat, Inbox, Work, Money and Agent views. Run it locally, behind owner login, or in a tenant-scoped deployment posture.

Multi-tenant + metering

Per-tenant isolation, usage metering and credit balances are in the core, so you can serve and meter multiple users on your own deployment.

Any model, native

OpenAI, Anthropic, Gemini, DeepSeek, OpenRouter and NVIDIA NIM through a native multi-provider layer (no third-party agent framework) — switch model mid-session; vision on vision-capable models.

Self-hosted, MIT

The whole engine is free and runs on your machine. Your keys, your data, your infrastructure — no feature behind a paywall, no hosted dependency.

Honest notes: the autonomy loops need BOTH personal-agent mode (POLYROB_LOCAL=true) and AUTONOMY_ENABLED=true — a fresh install is interactive-only, and a shared server stays off entirely. Every money verb, and host access, arms on its own flag on top of that; neither group turns them on. DEX swaps on the supported chains are live behind their own flags. Local code execution is convenience-sandboxed (timeouts, output caps, secret-free env), not a hard multi-tenant sandbox — keep it off in shared deployments (a hardened Docker backend is available for isolation). Hyperliquid and Polymarket order placement is dry-run by default and requires both the master and per-venue live switches plus per-venue caps; a blocked order reports a dry run rather than submitting silently. Twitter/X posting is an outbound tool; X DMs are also available as an inbound chat surface.