Abstrabit is a Private AI engineering company building AI-powered products and solutions for businesses. We work across the AI stack — from deploying and evaluating private and open-weight models to building production applications, agents, workflows, and integrations on top of them.
Our engineering work is organized across two closely connected layers: the AI layer, focused on models, inference, evaluation, and AI infrastructure; and the application layer, focused on building reliable, scalable software products that use those AI capabilities.
Job Description
The Context
Using an AI model through an API is very different from deploying and operating AI systems in production.
Private AI systems require decisions around model selection, inference, evaluation, latency, infrastructure, data pipelines, retrieval, cost, and reliability.
Our AI Engineering team works on this layer.
We are looking for interns who want to understand how modern AI systems work beyond prompt engineering and hosted APIs, and who are interested in learning how models are deployed, evaluated, and integrated into real products.
About the Role
As an AI Engineering Intern, you will work primarily on the AI layer of our products and solutions.
You will help deploy and evaluate open-weight models, build inference and retrieval pipelines, experiment with model configurations, and expose AI capabilities for application teams to use.
You will work primarily with Python and gain hands-on experience with technologies around model serving, LLM inference, embeddings, vector databases, RAG, evaluation frameworks, cloud infrastructure, and open-weight models.
You are not expected to already be an expert in model deployment or fine-tuning. We are looking for strong technical fundamentals, curiosity, and the ability to learn unfamiliar systems quickly.
What You'll Work On
Depending on the project, you may work on:
Deploying open-weight language and embedding models.
Building and testing model inference APIs and services.
Comparing models based on quality, latency, throughput, memory usage, and cost.
Building RAG pipelines using embeddings, vector databases, and retrieval systems.
Creating evaluation datasets and running model and application-level evaluations.
Experimenting with prompts, model parameters, retrieval strategies, and inference configurations.
Working with GPUs, containers, cloud infrastructure, and model-serving frameworks.
Investigating model failures and understanding why different models behave differently.
Documenting experiments so results can be reproduced and compared.
Working with Software Engineers to expose AI capabilities to production applications.
Qualifications
We care more about strong fundamentals and your ability to learn than whether you already know every AI framework.
You should have:
Strong Python fundamentals.
Good programming and problem-solving ability.
Basic understanding of APIs, data structures, Git, and software development.
Interest in understanding how machine learning and large language models work beyond using hosted APIs.
Ability to read technical documentation and implement unfamiliar concepts.
Comfort experimenting, measuring results, and investigating unexpected behaviour.
At least one meaningful technical project through coursework, a personal project, internship, research, hackathon, or similar work.
Ability to clearly explain what you built, what worked, what failed, and what you learned.
Prior experience with model deployment, RAG, PyTorch, Hugging Face, vector databases, GPUs, or cloud infrastructure is useful, but not required.
We use AI tools as part of our engineering workflow, and you are welcome to use them. We are not interested in testing whether you can memorize APIs or write every line of code without assistance.
What matters is whether you can understand what the system is doing, verify results, question incorrect assumptions, debug failures, and explain your reasoning.
Using AI effectively is useful. Blindly accepting its output is not.