Go Local.ai

Product — Hosting

Hosting and maintenance

AI without the meter running and the planet dying.

A flat monthly rate for all the AI your business needs, no surprise API bills, no John in accounting drafting emails with Fable, just right-sized models tailored to your needs managed by us.

The market will keep changing. 
Your operating costs should not have to change with it.

Inference, on your terms, no silent model adjustments and tweaks, no random usage limits, no sudden model withdrawals. Stability for every workflow with consistent capacity and isolation.

What hosting with us would look like for you

Managed sustainable inference for your business. We run the models your workflows need on rented compute, then provide the API endpoint your systems connect to. If you already use an API, moving can be as simple as changing the endpoint rather than rebuilding your workflows.

Model and API hosting. We do not require you to replace your existing business tools with a single Go Local AI chatbot.

It is managed. We monitor, maintain the agreed configuration, and provide the support set out in your agreement.

It can take over an existing setup. If you already run a local or open-weight AI system, we can assess whether we can host the existing configuration or improve it.

Hosting can follow one of our pilots, or we can discuss it directly if you already know what you want to deploy.

Our monthly rate covers

All the capacity your AI workflows need, and no live meter charging per token. It covers:

Intelligent Delegation & Routing. Automatically sending your requests to the appropriate model. So you don’t have to constantly decide which request would be most appropriate for what model. Efficiency that’s not just good for your wallet, but good for the planet as well.

The compute. Inference capability sized to your needs, workflows and usage, without any over-tiering.

Model serving. Running the models and configurations that provide the required quality at an affordable price through your API endpoint.

Monitoring. Constantly measuring inference, performance, capacity, and logging AI usage through our endpoint. Allowing detailed insight into how much AI your business really uses.

Support. Deploying a new tool or workflow that requires AI? We’re here to help integrate it smoothly into your existing plan.

A configuration designed around your needs. The named models, versions, settings and isolation tier that you have approved to serve your business best.

If you would like more information on the exact support hours, monitoring level and response arrangements send us an email.

Models fit to task

There is no fixed model shortlist, we stay up to date with the latest breakthroughs and releases in the open weight space, so you always have the option of utilising cutting edge AI.

The right model depends on the task, the quality your team needs, the amount of work, the response time and the information each workflow handles.

Context dependent deployments. Routine drafting, summarising, extracting and classifying may not need the same model or intelligence level as complex analysis, coding or long-context synthesis.

Configuration fit to hardware. The same model can need very different hardware at different levels of quantisation. We choose the model and configuration together, based on the work and the capacity which minimises any waste of computing resources, time, money or energy.

No silent model changes. We do not replace a model, alter its configuration or change its weights without your request or consent.

Hybrid is possible. If a small part of your work genuinely needs a cloud model, the router can send those requests to an agreed cloud API. Your systems still use one endpoint.

Hybrid is not automatically UK-only. Any request sent to a cloud provider leaves the Go Local hosted environment. We will name the provider, jurisdiction, retention terms and training position for that route before it is approved.

An isolation promise you can trust

We offer three levels of keeping your data separate and secure, depending on your needs, comfort and regulatory requirements.

Logical

For the standard tier our clients share the same hardware and the same inference models.

The model is a fixed, read-only set of weights. Every request reads from the same weights, but no request ever writes to them. Your data passes through the model to produce a response. It does not become part of the model, and it does not affect what any other client receives.

Each request is allocated its own isolated memory. No other client's request can access it. When the request completes, that memory is released. Nothing carries over between clients. The model retains no state between requests.

Your identity and access controls are enforced at the gateway before any request reaches the model. Each client has their own API credentials, enforced before the model sees anything; their own system prompt, assembled per request; their own logs, separated at the gateway; no shared context, no shared retrieval, no cross-tenant access.

The hardware is shared. Everything above it is separated. Shared inference divides the compute between clients who would otherwise each pay for hardware they use a fraction of the time. It raises utilisation, which is better for the environment, and it lowers your bill. It does not turn shared hardware into dedicated hardware.

Physical

Dedicated hardware for one client. The same gateway level separation as in logical separation, on a machine no other client's requests ever touch. We provide detailed documentation for every step of your data's processing. For clients who need physical separation or whose workload makes dedicated capacity the right choice.

Regulated (Coming Soon)

Single-tenant with the extra compliance documentation and controls that regulated sectors require. Availability, sector coverage and documentation are still being finalised.

We hold the same standard for every term we use. Local, on-premise, sovereign, isolated. We tell you exactly where the service runs, who operates the facility, and how your data is kept separate.

A rate you can plan around

Cloud providers can change prices, models, access terms, usage rules and model behaviour. They can also retire or replace models while a business is still building around them. That makes accurate planning difficult.

Our hosting agreement gives you a stable documented configuration and rate that you can truly depend on.

No per-token billing. You do not pay more because a request happened to contain more tokens, more rewrites, more reasoning, or discarded outputs.

The configuration is agreed in advance. You know which models and settings the service is built around.

There are no silent model swaps. If a new model becomes available, we can assess it with you. We do not move you onto it without consent.

Growth can be forecast. We size the service around your expected work, concurrency and planned growth, with headroom for peaks. Every deployment is designed with the ability to scale with you as your needs change.

It might not always be the cheapest option, but what does stability cost for you? And if a direct cloud API is the better option, we will tell you, see our AI cost saving audit for more details.

Usage and growth

Our hosting plan is designed around your real workloads and the capacity your business requires. We build in ample headroom for peaks rather than sizing the system to the average request. Your AI stays usable whether just early birds or the entire company are actively working with it.

No usage allowance. No unexpected limits. Capacity is designed around your peaks and then some, not your average day.

Normal peaks are part of the design. We size around expected concurrency and workload patterns.

Sustained growth is reviewed openly. If your workload grows enough to require additional compute, we discuss the adjustment with you and agree to any revised service and fee before making changes.

No surprise overage billing. We do not let an unnoticed usage spike turn into an unexplained invoice.

Where your data goes

Before any client data moves, we sign an NDA and a data processing agreement. We name every sub-processor, including the partner data centre.

Your data stays in the UK for the hosted service. It is not used to train a model for another client.

We provide inference rather than a general-purpose data store. We do not create a separate Go Local AI document repository for your business data as part of standard hosting.

While a request is being processed, information may be present temporarily in the serving process and the model's KV cache. That is working memory for inference, not a permanent client data store.

We keep operational logs needed to run, monitor and troubleshoot the service. The fields collected, whether request content is included and the retention period will be documented in the agreement and data processing documentation.

RAG for your system. Your documents and retrieval index can stay with you, or we can host them as a separate agreed service. We will document where they are held, who can access them and what maintenance is included.

Hybrid has a clear boundary

Hybrid routing is useful when routine work can stay on open-weight models but a small number of difficult requests need a cloud model. It is not the same as fully local hosting.

For a hybrid setup:

some requests leave the Go Local AI hosted environment

the cloud provider and jurisdiction are named

retention and training terms are checked

the route is agreed before use

Your chosen cloud API service gets billed directly to you

UK-only residency is not promised unless every route stays in the UK

If your data cannot leave the UK or your agreed environment, we will not present hybrid as a solution unless the routing architecture can genuinely meet that requirement.

The service is part of a wider responsibility

Hosting is not only a cost decision. It is also a decision about where your data goes, which infrastructure you use and what values & commitments your business is willing to stand behind.

Your data is not training material for someone else. We do not create or train large language or visual language models. With our full hosting we cannot supply any of your data to people who create or train them.

We do not have defence or surveillance ties. Our aim is to provide useful AI without contributing client data to uses and agencies that could harm people.

The model choice is part of the environmental decision. Right-sizing and routing can reduce the compute needed for routine work, and is part of taking the important step of reducing your business's environmental impact.

Pooling can improve utilisation. Where it is appropriate, sharing capacity can avoid dedicated hardware sitting idle. This is described as logical isolation.

We measure what we can. We aim to report energy and carbon information at model and workload level where the measurements and facility information support it.

We label uncertainty. Measured figures, estimates and unknowns are kept separate.

We do not claim that a hosting facility is renewable powered or carbon neutral without verified information. Hosting is delivered through a partner data centre. The facility, energy basis and supporting information are listed on our sustainability page.

A portion of Go Local AI revenue and profit will support named charitable causes.

What hosting does not include

No arbitrary limits. You do not have a weekly or hourly token limit, or a credit pool that runs out. The service is appropriately sized to the workload and capacity set for your business’s needs. If you start nearing capacity we can work together on expanding it.

Silent charges. If sustained growth needs more compute, we discuss it and agree it first, instead of quietly tacking it onto your bill.

A physical server that you own. This product uses rented compute. Deployment on your own hardware is a separate option.

Automatic model upgrades. New releases are not applied without your consent, and often not necessary. Upgrading without a need is wasteful.

A new business application or chatbot. We host models and APIs. New interfaces and business applications are separate work.

Implementation beyond the agreed setup. Wider migration, new workflows requiring more capacity, and company-wide rollout are quoted separately.

Rebuilding your data or processes. If your data, retrieval or processes need work, that is a separate service.

Email us

Send us the details of what you're working with and what you're hoping to change. We reply to every enquiry.

contact@go-local.ai

Book a meeting

Pick a time that suits you for a free 30-minute consultation. No commitment, no sales pressure, just a conversation about where you are.

Book a 30-minute call