01
Local model serving
Exploring model selection, quantisation, memory and context limits using local runtimes such as Ollama.
AI & research / Local & distributed AI
Local language models, retrieval and distributed execution across CPU and GPU resources.
Research & prototyping
The question
Where a model runs changes its cost, latency, data boundary and operational requirements. Our research explores how local and distributed AI can become a practical part of a business platform.
The approach
01
Exploring model selection, quantisation, memory and context limits using local runtimes such as Ollama.
02
Working with embeddings, vector retrieval and metadata to provide a model with relevant information from the surrounding system.
03
Investigating queue-based coordination, CPU/GPU scheduling and execution across different machines and gateways.
04
Assessing response quality, throughput, latency, memory use and operational complexity against a defined task.
Work with Gala Arch
Interested in a practical AI use case, a research collaboration or an evaluation? Tell us what you want to learn.