AI / DevOps Engineer β five years across HFT, fintech and banking, building the full stack behind production LLM systems.
Currently at an international HFT hedge fund: self-hosted inference on vLLM and llama.cpp for confidential data, agentic and RAG pipelines, and the Kubernetes, observability and CI/CD platforms they depend on. I measure everything and optimise for cost per token, latency and reliability.
AI & Agentic Systems
- Self-hosted LLM inference on vLLM and llama.cpp with open-weight models (Gemma 4 E4B / 12B, FP8 and 8-bit), plus whisper.cpp for speech-to-text
- Inference tuning: quantization, speculative decoding, prefix caching, scheduler sizing β on power-capped GPUs
- Agentic and RAG pipelines, harness engineering, and a memory engine for AI agents
- Hybrid retrieval over PostgreSQL: BM25, pgvector, graph search
- AI application development with LangGraph (Python), rig (Rust), and eino (Go)
- LLM evaluation, tracing, and observability with LangChain, Langfuse, Prometheus, Grafana
Infrastructure & Platform
- Cloud and bare-metal automation with AWS, Terraform, Pulumi
- Kubernetes (bare-metal clusters via Kubespray, Talos Linux, GitLab CI + Argo CD), including GPU node pools for inference
- Monitoring: Zabbix, Prometheus, Grafana LGTM Stack, VictoriaMetrics, ELK
- Tracing & profiling: Grafana Alloy, Tempo, Pyroscope
Languages
- π Python β daily driver for automation, data pipelines, and AI applications
- 𦦠Go β when simplicity matters
- π¦ Rust β when correctness and performance demands it
- π Bash β glue that holds the universe together
- π§ C β when feeling adventurous
- The $250 server that replaces up to $7900 in API tokens β one month of self-hosted vLLM inference: 1 billion tokens for $250 versus $95β$7,900 on hosted APIs
- 612 output tokens a second on 70 watts: how to tune an AI server β prefill vs decode, speculative decoding, prefix caching and scheduler sizing, explained for non-specialists
π Resume
| Period | Role | Company |
|---|---|---|
| 2023βPresent | AI Infrastructure / DevOps Engineer | International HFT Hedge Fund |
| 2022β2023 | DevOps Engineer | CentrInform |
| 2022 | Release Engineer | Banks Soft Systems (BSS) |
| 2021β2022 | Python & Blockchain Developer | Genesis Block |







