Founding Engineer, General Infrastructure
Description
About Goaly At Goaly, our mission is to make custom AI affordable for every business. Our founding team comes from the front lines of top AI labs and tech giants (Meta MSL, TikTok AI, Google DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering trillion-parameter models and scaled GenAI models to a global user base. Now, we are building something we wish we had before: a platform that makes training and adapting custom AI affordable for all modern companies, not just Big Tech. Our north star is ambitious: for a domain-specific task, reach 90% of SOTA performance at less than 10% of the cost. To get a taste of what we are doing, see our first tech blog. The role We are looking for a Founding Infrastructure Engineer & Strong "Fixer" who has scaled large distributed systems and thrives on building foundational infrastructure. You will design the data processing systems that power our LLM training platform; implement scalable core primitives (tokenization, deduplication, chunking), and build comprehensive monitoring to ensure reliability at scale. You are a builder who loves reliable code, automated infrastructure, and solving the most complex technical challenges in AI systems. Key responsibilities System Architecture: Design, implement, and maintain core infrastructure platforms, including large-scale distributed training systems, model serving architecture, data processing pipelines etc. DevOps & Containerization: Orchestrate containerized environments (Docker/Kubernetes) for massive scale. Manage deployment pipelines and ensure high availability for model serving. Data Infrastructure: Manage complex data pipelines involving real-time streaming (Kafka/Redpanda) and vector databases. High-Performance Engineering: Optimize system performance to support high concurrency and massive data processing. Build reusable base service components to eliminate redundancy. Agentic Platforms: Design the backend systems powering our agentic workflows, including flow engines, knowledge base services, plugin markets, and user management systems. Who you are Education: Bachelor's or Master's in CS or related field. Senior Engineering Experience: 5+ years of distributed systems experience. Proficiency in at least one: Python, Go, C++, or Java. Distributed Systems Expert: Deep understanding of CS fundamentals (data structures, algorithms). Proven experience with high availability (HA), high concurrency, and service governance. Modern Data Stack: Hands-on experience with streaming (Kafka, Redpanda); transactional DBs (PostgreSQL, TimescaleDB); vector DBs (Pinecone, Milvus, Weaviate); storage (S3, Azure Blob, data lakes). Cloud Native: Deep expertise in Kubernetes (K8s), Docker, and IaC (Terraform/Pulumi). AI Curiosity: A strong interest in LLMs, RAG, and agent architectures. Bonus points Ex-FAANG Background: Experience working on large-scale distributed systems at major tech companies. ML Infrastructure: Experience with ML DevOps and model serving (e.g., vLLM / SGLang, Ray Serve, Triton). Data Processing: Experience building embedding pipelines or large-scale ETL workflows (Airflow/Dagster). AI Integration: Prior experience deploying open-source LLMs (e.g., Llama 3, DeepSeek, Qwen) or AI training/inference experiences. Why join us? Architect for Unprecedented Scale & Efficiency: Apply your experience in building robust, scalable systems to the unique, demanding world of AI model training. You'll tackle challenges that far exceed traditional web services, optimizing for GPU clusters and exabyte-scale data pipelines. Build Foundational Technology, Not Just Features: This is an opportunity to design core infrastructure that will become the standard for an entire industry. Your work will empower hundreds of companies to leverage advanced AI, fundamentally changing their business operations. Direct Line to Impact: In our small, agile team, your architectural decisions and engineering prowess directly translate into our product's performance and our customers' success. Mentorship from AI-Native Scale Leaders: Want to dive deeper into AI infrastructure? Work alongside our founding members who trained trillion-parameter models at scale. Learn and apply best practices from the highest echelons of AI infrastructure, while contributing your own deep systems knowledge. Meaningful Equity: Receive an early-stage package with massive upside—you directly capture the value of the infrastructure breakthroughs you create. How to apply Send your resume and relevant project portfolio to recruiting@goaly.ai. Show more Show less