The primitives
Functions for short work, sandboxes for untrusted code, agents for the conversation.
AI for partner networks
Each partner gets their own agent — memory, tools, and a spend cap. Their data does not mix. You keep the brand.
Product
An agent per partner, with its own tools and memory. Their customers never see another partner’s context.
Gemini, Gemma, or a compatible endpoint. No rewrite of the partner’s stack.
Chat, batch, or a spike at month-end. The endpoint follows the load.
We run the runtime. The partner sees your product, not ours.
Fine-tune when you have the data. The job runs, you get a versioned endpoint back.
SFT or LoRA. You send the dataset. We return a version you can promote.
Jobs go up on Cloud Run or GKE and come down when they finish.
Ask for the size the job needs. Don’t pay for a rack you are not using.
A short-lived machine for an agent that needs to run code or open a file. Snapshot it. Throw it away.
Spin up a lot of sandboxes when a rollout needs them. They do not share disks.
Start from a default image or the one the channel already uses.
CPU and memory follow the job. Idle time does not sit on the bill.
A function becomes an HTTP endpoint. It grows with traffic and goes quiet when nobody calls.
Fan out a batch without writing your own queue.
Leave a run going. Checkpoint lands in Cloud Storage.
Auth, a cap per partner, and a log you can read later.
Platform
Functions for short work, sandboxes for untrusted code, agents for the conversation.
Weights, datasets, and images sit close to where the run happens.
A runtime for large images and a queue when a job needs an accelerator.
Where it runs
Default region is São Paulo. We can stand up another Google Cloud region if a customer asks.
Idle tenants should not keep machines warm.
Training jobs can take a GPU or TPU on Vertex AI when you need them.
Logs and alerts per tenant, so a partner issue stays a partner issue.
In production
Access control, logs, and a bill you can show a partner.
IAM, audit logs, and data that stays where you put it. Built with LGPD in mind.
What ran, who called it, what it cost. Export if you already have a sink.
Pay for what the code uses. Set a ceiling per partner so a runaway job does not surprise you.
How we run it
OrbySoft runs on Google Cloud. São Paulo is the home region. Agents use Gemini through Vertex AI.
What the agents call today, through Vertex AI. We can point at another model if the job needs it.
Where functions and sandboxes run. They scale down when the channel is quiet.
Telemetry, evals, and versioned weights. One partner’s data does not train another’s model.
Secrets stay out of the code. Each partner is a separate scope.
São Paulo by default. Another region only if you ask for it.
A platform fee, then tokens and compute. The invoice can split by partner.
Notes
Memory, tools, and policy stay inside that tenant. Even the logs.
ProductFreeze a short-lived machine and bring it back without a cold start speech.
PlatformNew environments land in southamerica-east1 unless you say otherwise.
DataTell us about the network. We’ll stand up an environment with a spend cap.