- Home
- Jobs
- Platform / Infrastructure
- AI Infrastructure Engineer
AI Infrastructure Engineer
On this page
TechJobs.ie Job Insights
At a glance
- Employment type
- Full-time
- Workplace
- Remote
- Category
- Platform / Infrastructure
Technologies & skills
What you'll be doing
Fin is seeking Senior+ AI Infrastructure Engineers to build and scale systems for training and serving next-gen AI products. You'll work on model training, inference, and GPU-level performance, collaborating with ML scientists and mentoring team members. The role is remote within Ireland.
- Implement and scale training pipelines for large transformer and LLM models
- Build and optimize inference services for low-latency, high-reliability customer experiences
- Work on GPU-level performance: tuning kernels, improving utilization, identifying bottlenecks
- Collaborate with ML scientists to implement and productionize training and inference methods
- Mentor and develop other engineers on the team
- Raise technical standards, reliability, and operational excellence across the AI platform
Key requirements
Must-have
- 5+ years of experience in software engineering
- Degree in Computer Science, Computer Engineering, or related field (or equivalent experience)
- Hands-on experience with model training (especially transformers and LLMs), model inference at scale, or low-level GPU work (e.g. CUDA, Triton)
- Deep knowledge of at least one programming language (e.g. Python, Ruby, Java, Go, etc.)
- Comfortable working in production environments at meaningful scale
- Strong technical fundamentals and clear communication skills
Nice-to-have
- Experience at AI-native companies training or running inference for their own models
- Experience running training or inference workloads on Kubernetes
- Experience with AWS or other major cloud providers
- Production experience with Python in ML or infrastructure contexts
- Passion for technology demonstrated through personal projects, open source, or publishing content
Experience: Senior+ (5+ years)
Role signals
- Technical focus
- AI infrastructure, model training, inference, GPU performance
- Leadership
- Mentoring
- Architecture / system design
- Indicated in the listing
- Hands-on vs management
- Hands-on
Similar jobs
Check your CV against this job
Get a certified technical review aimed at this role.
Full job description
Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences.
Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams.
Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers.
##** What's the opportunity?**
We’re looking for Senior+ AI Infrastructure Engineers to build the systems that train and serve Fin's next generation of AI products.
Fin is an AI company that builds from the GPU all the way up to a user agent that resolves millions of customer service queries a month.
You’ll join a small, highly technical team working at the cutting edge of modern AI infrastructure. The AI Infra team built the training pipelines and runs the inference for custom models like Fin Apex, which outperforms frontier models in customer service tasks, and is the foundation of the AI Group's full stack approach to AI.
** We’re particularly interested in engineers who have:**
- A track record of working on ** model training ** or ** model inference at scale **, or on ** low‑level GPU coding **(e.g. CUDA, Triton). Experience with one is great, multiple is even better.
##** What will I be doing?**
As a Senior AI Infrastructure Engineer focused on model training and inference, you will:
-
Implement and scale ** training pipelines ** for large transformer and LLM models, from data ingestion and preprocessing through distributed training and evaluation.
-
Build and optimize ** inference services ** that deliver low‑latency, high‑reliability experiences for our customers, including autoscaling, routing, and fallbacks.
-
Work on ** GPU‑level performance **: tuning kernels, improving utilization, and identifying bottlenecks across our training and inference stack.
-
Collaborate closely with ML scientists to implement cutting edge training and inference methods and bring them to production.
-
Play an active role in ** hiring, mentoring, and developing ** other engineers on the team.
-
Raise the bar for ** technical standards, reliability, and operational excellence ** across Fin's AI platform.
##** Profile we’re looking for:**
These are indicative, not hard requirements
We’re looking to hire ** Senior+ AI Infrastructure Engineers **. You’re likely a great fit if:
-
You have ** 5+ years of experience in software engineering **, with a strong track record of shipping high‑quality products or platforms.
-
You hold a ** degree in Computer Science, Computer Engineering, or a related field **(or you have equivalent experience with very strong fundamentals).
-
You have hands‑on experience with ** one or more ** of the following:
-** Model training **(especially transformers and LLMs).
-** Model inference ** at scale (again, especially transformers and LLMs).
-** Low‑level GPU work **, such as writing ** CUDA ** or ** Triton ** kernels.
Comfortable working in production environments at ** meaningful scale **(traffic, data, or organizational).
You communicate clearly, can explain complex technical topics to different audiences, and enjoy close collaboration with both engineers and non‑engineers.
You take pride in strong ** technical fundamentals **, love learning, and are willing to invest in your own development.
Have deep knowledge of at least one programming language (for example Python, Ruby, Java, Go, etc.). Specific language experience is less important than your ability to write clean, reliable code and learn new stacks quickly.
##** Bonus skills & attributes **
None of these are required, but they’re nice to have:
-
Experience at ** AI native companies ** that ** train and/or run inference for their own models **(e.g. modern AI labs or AI‑native product companies).
-
Experience running training or inference workloads on ** Kubernetes **.
-
Experience with ** AWS ** or other major cloud providers.
-
Production experience with ** Python ** in ML or infrastructure contexts.
-
Demonstrated ** passion for technology ** through personal projects, open source, meetups, or publishing content about your work and learnings
Benefits
We are a well treated bunch, with awesome benefits! If there’s something important to you that’s not on this list, talk to us!
-
Competitive salary and equity in a fast-growing start-up
-
We serve lunch every weekday, plus a variety of snack foods and a fully stocked kitchen
-
Regular compensation reviews - we reward great work!
-
Unlimited access to Claude Code and best-in-class AI tools; experimentation & building is encouraged & celebrated.
-
Pension scheme & match up to 4%
-
Peace of mind with life assurance, as well as comprehensive health and dental insurance for you and your dependents
-
Flexible paid time off policy
-
Paid maternity leave, as well as 6 weeks paternity leave for fathers, to let you spend valuable time with your loved ones
-
If you’re cycling, we’ve got you covered on the Cycle-to-Work Scheme. With secure bike storage too
-
MacBooks are our standard, but we also offer Windows for certain roles when needed.
#LI-Hybrid
** Policies **
Fin has a hybrid working policy. We believe that working in person helps us stay connected, collaborate easier and create a great culture while still providing flexibility to work from home. We expect employees to be in the office at least three days per week.
We have a radically open and accepting culture at Fin. We avoid spending time on divisive subjects to foster a safe and cohesive work environment for everyone. As an organization, our policy is to not advocate on behalf of the company or our employees on any social or political topics out of our internal or external communications. We respect personal opinion and expression on these topics on personal social platforms on personal time, and do not challenge or confront anyone for their views on non-work related topics. Our goal is to focus on doing incredible work to achieve our goals and unite the company through our core values.
Fin values diversity and is committed to a policy of Equal Employment Opportunity. Fin will not discriminate against an applicant or employee on the basis of race, color, religion, creed, national origin, ancestry, sex, gender, age, physical or mental disability, veteran or military status, genetic information, sexual orientation, gender identity, gender expression, marital status, or any other legally recognized protected basis under federal, state, or local law.
Interview prep pack
Grounded in this listing. Use it to prepare examples before you apply.
Your interview focus
You will be assessed on your ability to design, implement, and scale AI infrastructure for training and inference of large language models, with a focus on GPU-level optimization, distributed systems, and collaboration with ML scientists. Strong technical fundamentals and experience with low-level GPU work are key.
- Model training pipelines·High
- Model inference at scale·High
- GPU-level optimization·High
- Distributed systems·High
Only have 30 minutes?
Follow a focused preparation plan based on this job.
Start 30-minute prep
Your 30-minute plan
Review Distributed Training and Inference Projects
0–8 minSummarize your hands-on experience with large model pipelines, focusing on transformers and LLMs.
Refresh GPU Optimization Techniques
8–15 minRevisit your work with CUDA or Triton, and note specific performance improvements you've achieved.
Prepare Mentorship and Collaboration Examples
15–23 minIdentify stories where you mentored others or worked closely with ML scientists to deliver results.
Research Intercom's AI Platform
23–30 minUnderstand Fin's AI agent, its integration with help desk products, and recent technical milestones.
Talking points
, 6 itemsDesigning and Scaling Model Training Pipelines
You will build and scale pipelines for large transformer and LLM models, covering data ingestion, distributed training, and evaluation.
Optimizing Model Inference Services
Delivering low-latency, reliable inference is central to the customer experience and a core responsibility of this role.
Low-level GPU Programming (CUDA, Triton)
Hands-on experience with GPU kernel tuning and bottleneck identification is explicitly required for both training and inference.
Distributed Systems and Autoscaling
The role involves building systems that operate at meaningful scale, requiring knowledge of autoscaling, routing, and reliability.
Collaboration with ML Scientists
You will work closely with ML scientists to implement and productionize new training and inference methods.
Technical Mentorship and Standards
Raising the bar for technical standards and mentoring others is a stated expectation for this senior role.
What to research
, 3 itemsModel Training and Inference at Scale
This role emphasizes building and scaling pipelines for large transformer and LLM models, including distributed training and high-reliability inference services.
Low-level GPU Programming and Optimization
Hands-on experience with CUDA or Triton for kernel tuning and performance bottleneck identification is a key requirement.
Technical Mentorship and Raising Standards
You are expected to mentor others and set high technical standards, so be ready to discuss your approach and impact in these areas.
Questions to ask
, 6 itemsWhat are the biggest technical challenges the AI Infra team is currently facing?
Why ask this? Reveals the team's priorities and where your skills could have the most impact.
How do engineers and ML scientists typically collaborate on new model deployments?
Why ask this? Clarifies the workflow and expectations for cross-functional work.
What tools and frameworks are used for GPU-level optimization and monitoring?
Why ask this? Helps you understand the technical stack and where your experience fits.
How does the team approach technical mentorship and onboarding for new members?
Why ask this? Shows the support structure and opportunities for growth.
What does success look like for this role in the first six months?
Why ask this? Clarifies expectations and how your performance will be measured.
How is operational excellence and reliability measured and improved within the AI platform?
Why ask this? Gives insight into quality standards and processes.
Register now to upload your CV
Create a free account, save a PDF or Word CV, and quick apply on roles that take applications here.
