# Gilad Maoz — AI/LLM & DevOps Engineering Leader (Full) > Gilad Maoz is an Engineering Manager specializing in AI/LLM product strategy and DevOps with 15+ years experience building scalable cloud infrastructure, leading technical teams, and delivering 100+ projects across 8 companies. He is available for full-time roles and consulting engagements. Based in Budapest, Hungary, working remote-first with global availability. --- ## About Gilad Gilad has spent over 15 years as "the tech guy behind the screen" — ensuring systems work flawlessly, scale effortlessly, and never fail at 3 AM. From early days in QA and automation to architecting cloud infrastructure at scale. His journey includes roles at Verimatrix, Yadati, Finqle, vidaXL, and many others, where he built DevOps pipelines, cloud architectures, and automation systems that serve millions of users across diverse industries. His most recent engagement was leading engineering at nxty.ai (2025–2026), where he built production LLM infrastructure on Cloudflare Workers and achieved a 40% reduction in AI infra costs through multi-LLM routing via OpenRouter, LangGraph orchestration, and LangFuse observability. He is now available for new engagements — both full-time remote roles and consulting projects — focused on AI/LLM infrastructure, platform engineering, and hands-on technical leadership. ## Career Timeline ### Mar 2025 – Jan 2026: Head of Engineering & Platform Lead — nxty.ai Led engineering operations for early-stage AI startup, managing team of 6 developers. Architected production LLM infrastructure on Cloudflare (AI Gateway, Workers AI, Vectorize) using LangGraph, LangFuse, and OpenRouter. Reduced infrastructure costs by 40% migrating from Vercel+Supabase+OpenAI to Cloudflare Workers. Handled 10,000+ daily requests at sub-200ms latency. ### 2021 – 2024: Senior DevOps Engineer — Verimatrix Built cloud tools and automation, spearheaded AWS CDK migration, managed artifact repositories and supported CI/CD pipelines. ### 2020 – 2021: Cloud & Software Architect — Yadati Functioned as software architect, moved entire operation from Azure to Firebase, designed Cloudflare Worker reverse proxy, managed development team. ### 2018 – 2020: Backend Developer & DevOps — vidaXL & Finqle Implemented microservices using Firebase Functions and Cloudflare Workers, automated CI/CD pipelines, migrated legacy Python codebase. ### 2014 – 2018: Automation & DevOps Engineer — Early Career Journey Started in QA and automation, established test environments, built automation infrastructure, led AWS to GCP migrations. ## Core Expertise ### AI Infrastructure & LLM Integration Design and implement production-ready LLM systems using Cloudflare Workers AI, multi-model routing, and cost optimization strategies. Includes Cloudflare Workers AI integration architecture, multi-provider LLM routing via OpenRouter (OpenAI, Claude, open-source), LangGraph orchestration and workflow implementation, LangFuse observability and cost monitoring, and inference cost optimization (40%+ savings achieved). Real example: Implemented multi-LLM routing at nxty.ai, reducing costs by 40% while improving quality through intelligent model selection. Best for: Startups building AI-first products, teams adding LLM features, companies with high inference costs. ### Cloudflare-Native Architecture & Migration Build cost-effective, globally distributed infrastructure using Cloudflare Workers, AI Gateway, Vectorize, and edge computing patterns. Includes Cloudflare Workers serverless architecture, AI Gateway integration and optimization, Vectorize (vector database) implementation, AutoRAG system design, and AWS/GCP to Cloudflare migrations. Real example: Migrated infrastructure to Cloudflare Workers ecosystem, reducing hosting costs by 60% and eliminating regional latency issues. Best for: Companies with high cloud bills, teams seeking global distribution, organizations outgrowing traditional cloud providers. ### VP R&D / Fractional Technical Leadership Part-time or full-time technical leadership for startups and growing teams. Strategy, team building, and hands-on development combined. Includes engineering team leadership (10-30 hrs/week), technical strategy and product roadmap, cross-functional collaboration (Product/Eng bridge), DevOps culture and best practices, and technical hiring and team scaling. Real example: Served as Head of Engineering at nxty.ai (2025–2026), managing a team of 6 developers building AI-powered products on Cloudflare infrastructure. Best for: Startups scaling engineering teams, companies needing interim CTO, organizations bridging product/engineering gaps. ### DevOps & Infrastructure Modernization Transform legacy infrastructure into modern, cost-effective systems using Infrastructure as Code, serverless patterns, and intelligent automation. Includes cloud migration strategy (AWS/GCP/Azure → Cloudflare), Terraform Infrastructure as Code implementation, CI/CD pipeline optimization (GitHub Actions, AWS CDK), cloud cost reduction audits and implementation, and legacy system modernization roadmaps. Real example: Built AWS CDK pipelines at Verimatrix, reducing deployment time 70% while improving reliability through automation. Best for: Companies with technical debt, organizations with high cloud costs, teams modernizing infrastructure. ## Key Projects ### thriv.es — AI Product Studio Two-founder studio (Budapest & Tel Aviv) in AI Product Studio / Build, Advise, Lab. 2024 – Present. Co-founder & CTO. Two-founder AI product studio that turns validated ideas into production-ready AI products in fixed-scope, fixed-price sprints of six to eight weeks. Founders work directly with every client — no account managers, no hand-offs — and all code lands in the client's own repository from day one. The problem: Early-stage teams and enterprises with a clear AI idea struggle to ship production systems quickly: prototype-to-production migrations stall, multi-provider AI is expensive and unreliable, and most agencies hand off to junior staff after the sale. The approach: Architected and built the studio's flagship platform end-to-end as technical co-founder: edge-native AI orchestration on Cloudflare Workers, multi-provider routing across Cloudflare AI, Claude, and OpenAI, and a unified monorepo consolidating three repos into a single type-safe workspace. Key results: - 40% cost reduction through intelligent multi-provider routing - <200ms API response times via Cloudflare Workers edge computing - 6–8 week sprints delivering production-ready AI products - 3 AI providers (Cloudflare AI, Claude, OpenAI) with smart routing - 2,500+ lines of duplicate code eliminated through monorepo migration - 6+ open-source libraries released (domain-knowledge, movers, prompt-builder, slides-maker, transcriber, kid-runtime) Tech stack: React 18, TypeScript 5.9, Vite 7, Tailwind CSS 4, Hono, Cloudflare Workers, Cloudflare AI, Anthropic Claude, OpenAI, Drizzle ORM, Cloudflare D1, Cloudflare R2, Whisper, ElevenLabs, Azure TTS, Langfuse, JWT, OAuth, WebSockets, Turborepo, pnpm Workspaces, Playwright Links: https://thriv.es/ | https://github.com/thriv-es ### HansCard — AI-Powered Language Learning Open Source Project in EdTech. 10 months (Mar 2024 – Present). Solo Developer. Full-stack language learning SaaS combining spaced repetition with AI tutoring, mobile apps, and real-time sentence generation. The problem: Traditional language apps teach isolated vocabulary. Real fluency requires learning words in context with proper sentence structures. The approach: Built modern web application with React/TypeScript frontend, Supabase backend, and AI tutoring using GPT-4 and Claude. Key results: 3 languages supported, 2 AI models (GPT-4 and Claude), mobile app builds with Capacitor for Android, real-time sentence generation from Tatoeba API. Tech stack: React, TypeScript, Vite, Chakra UI, Supabase, PostgreSQL, Edge Functions, OpenAI API, Anthropic Claude, Capacitor, Vercel, Sentry, Vitest Links: https://card.evilurge.com/ | https://github.com/evilUrge/HansCard | https://www.evilurge.com/posts/hanscard/ ### ShellGems — Technical Blog & Experiments Personal Project in Developer Education. Ongoing since 2018. Solo. Technical blog showcasing hands-on experiments with AI, serverless architecture, and DevOps automation. 12+ in-depth articles. The problem: Needed platform to document learning, share insights, and establish thought leadership in AI/DevOps space. The approach: Built modern static blog using Hugo and Firebase Hosting with automated CI/CD via CircleCI. Key results: 12+ articles published, <2s build time, <100ms global latency, active since 2018. Tech stack: Hugo, Firebase Hosting, CircleCI, Markdown, Git Links: https://www.evilurge.com/ | https://github.com/evilUrge/ShellGems ## Key Achievements - 40% reduction in AI infrastructure costs at nxty.ai through multi-LLM routing - 10,000+ daily requests handled at sub-200ms latency on Cloudflare Workers - 60% hosting cost reduction migrating from AWS/GCP to Cloudflare Workers ecosystem - 70% deployment time reduction through AWS CDK pipelines at Verimatrix - 100+ projects delivered across 8 companies and diverse industries - 99.7% project success rate across 17+ engagements - Built 3+ production SaaS products (thriv.es, HansCard, ShellGems) ## Frequently Asked Questions ### How do we start working together? We begin with a complimentary discovery call to understand your specific challenges and goals. Gilad assesses your technical environment, business objectives, and team dynamics. Following this conversation, he proposes a tailored roadmap with clear milestones and success metrics. ### Do you work remotely or on-site? Gilad primarily works remotely for maximum efficiency and flexibility. However, for strategic workshops, architecture sessions, or team alignment meetings, he is available to travel globally. Based in Europe, he regularly works with teams across multiple time zones. ### What industries do you specialize in? Expertise spans across FinTech, Healthcare, Enterprise SaaS, and AI startups. Focus is on technical implementation and scalable architecture rather than industry-specific solutions. This approach ensures proven methodologies translate effectively across sectors. ### How long do typical engagements last? Strategy projects typically run 4-8 weeks, while transformation initiatives can extend 6-12 months. Flexible project-based and retainer arrangements are available, with the ability to scale team involvement based on needs and budget. ### Do you sign NDAs and handle sensitive data? Absolutely. All engagements include comprehensive NDAs, and enterprise-grade security practices are followed for data handling and communication. ### What's your typical ROI and success rate? Clients typically see 3-10x ROI within 6 months through improved efficiency, reduced costs, and accelerated time-to-market. With a 99.7% project success rate across 17+ engagements, the focus is on delivering measurable business impact. ## Core Technical Skills - AI/LLM: Cloudflare Workers AI, OpenRouter, LangGraph, LangFuse, OpenAI API, Anthropic Claude, Prompt Engineering, RAG systems, Vectorize - Cloud Infrastructure: Cloudflare Workers/Pages/D1/R2/KV/AI Gateway, AWS (CDK, Lambda, ECS, S3), GCP, Firebase, Vercel, Supabase - DevOps: Terraform, CI/CD (GitHub Actions, CircleCI), Docker, Infrastructure as Code - Languages: TypeScript, Python, JavaScript, SQL, Bash - Frameworks: React, Next.js, Hono, Hugo, RedwoodJS, Vite, Tailwind CSS - Leadership: Engineering team management, technical strategy, product roadmap, cross-functional collaboration, technical hiring ## Contact - Email: gilad@maoz.dev - Website: https://maoz.dev - GitHub: https://github.com/evilurge - LinkedIn: https://linkedin.com/in/gilad-maoz - Location: Budapest, Hungary (UTC+2) - Availability: Remote-first, global travel for strategic engagements