Blog/RAG and LLM Skills for AI Jobs: What to Learn

RAG and LLM Skills for AI Jobs: What to Learn

Calling a chatbot API is no longer enough. Here are the RAG and LLM skills, from embeddings to evaluation, that AI job postings in India expect.

Last updated: 8 October 2026 · By the Asuraa Team

Quick answer: For LLM and generative AI jobs, learn to build retrieval-augmented generation (RAG) systems end to end: Python, prompting, embeddings, chunking, vector databases, hybrid search and reranking, evaluation and deployment. AWS describes RAG as grounding a model in an authoritative knowledge base without retraining it. Employers want proof through a deployed project with a test set and measured results, not just a PDF chatbot demo.

Key takeaways

  • RAG retrieves relevant information from a trusted source and passes it to a language model so answers are grounded in company data without retraining the model.
  • AWS breaks RAG into four steps: create embedded external data, retrieve relevant chunks, augment the prompt, and keep the data updated.
  • pgvector runs exact nearest-neighbour search by default, while HNSW or IVFFlat indexes switch to faster approximate search that can miss some results.
  • In Anthropic's September 2024 tests, contextual embeddings plus contextual BM25 and reranking cut retrieval failures from 5.7% to 1.9%, a 67% reduction.
  • A RAG project with a 30-50 question evaluation set and measured before-and-after results is far more convincing to employers than a simple PDF chatbot.

To get an AI job built around large language models, learn to build retrieval-augmented generation (RAG) systems end to end: embeddings, vector search, prompt assembly, evaluation and deployment. Employers hiring generative AI developers and AI engineers in India increasingly expect you to ship a working, measured LLM application, not just call a chatbot API.

What is RAG and why do employers care about it?

RAG is a pattern where an application first retrieves relevant information from a trusted knowledge source and then passes it to a language model so the answer is grounded in that information. It lets companies use their own documents with an LLM without retraining the model.

AWS describes RAG as a lower-cost alternative to retraining that helps keep answers current, lets users see where answers came from through citations, and gives developers control over which sources the model can use. The term comes from a 2020 research paper by Patrick Lewis and colleagues, which combined a text generator with a retriever over a dense index of Wikipedia (arXiv).

For employers, RAG matters because almost every enterprise LLM project in India, whether an internal HR assistant, a customer support bot or a document search tool for a bank, needs the model to answer from company data. That is why "RAG" appears so often in generative AI job postings. LinkedIn's Skills on the Rise 2026 report for India also listed LLMOps in its AI and automation cluster, according to afaqs! in March 2026.

How does a RAG pipeline work, step by step?

A RAG pipeline has two phases: indexing your data ahead of time, and retrieving plus generating at query time. AWS breaks it into four steps, which map neatly to the skills you need.

  1. Create the external data. Collect documents from files, databases or APIs, split them into chunks and convert each chunk into a numerical embedding stored in a vector database.
  2. Retrieve relevant information. Convert the user's question into an embedding and find the most similar chunks.
  3. Augment the prompt. Insert the retrieved chunks into the prompt with clear instructions, such as answering only from the provided context and citing sources.
  4. Keep the data updated. Refresh documents and their embeddings on a schedule or in real time so answers stay current.

On top of this, production systems add evaluation, monitoring, access control and cost tracking. Those additions are what separate a tutorial project from something an employer trusts.

Which RAG and LLM skills should you learn?

Learn the skills in layers, starting with Python and ending with evaluation and deployment. The table shows what each layer covers and what to build to prove it.

Skill layerWhat to learnProof project idea
FoundationsPython, REST APIs, JSON, basic SQLScript that calls a model API and parses structured output
PromptingSystem prompts, few-shot examples, structured output, chainingTicket classifier that returns JSON
EmbeddingsWhat embeddings are, similarity measures, embedding model choiceSemantic search over 500 FAQ entries
ChunkingChunk size, overlap, splitting by headings or sentencesCompare answer quality with two chunking strategies
Vector databasesIndexing, approximate vs exact search, metadata filtersStore chunks in pgvector or a managed vector DB
Retrieval qualityHybrid keyword + vector search, rerankingAdd BM25 and a reranker, measure the change
EvaluationTest sets, faithfulness, context relevance, answer relevance50-question eval with scores per version
ProductionAPIs, Docker, cloud deployment, logging, cost and latencyDeployed app with a usage dashboard

If you are starting from scratch, cover Python and basic machine learning first using our guides on how to learn Python for jobs and how to learn generative AI.

What should you know about embeddings and vector databases?

You should understand that an embedding turns text into a list of numbers so that similar meanings sit close together, and that a vector database finds the nearest ones quickly. You do not need to derive the maths, but you must be able to explain the trade-offs.

The pgvector extension for PostgreSQL is a good learning tool because it shows the trade-offs plainly. It supports distance measures including L2, inner product and cosine distance. By default it performs exact nearest-neighbour search with perfect recall; adding an HNSW or IVFFlat index switches to approximate search, which is faster but can miss some results. Its documentation notes that HNSW generally gives a better speed-recall balance but builds more slowly and uses more memory than IVFFlat.

Being able to say "I used HNSW because query speed mattered more than build time, and I checked recall on my test set" is exactly the kind of reasoning interviewers look for. If you already know SQL, pgvector also keeps your stack simple. Our guide on how to learn SQL for jobs helps here.

How do you improve retrieval quality?

Improve retrieval by combining keyword and vector search, reranking results, and giving each chunk enough context. Measure every change against a fixed test set.

Anthropic's September 2024 write-up on contextual retrieval is a useful case study with published numbers. It adds a short, chunk-specific explanation to each chunk before embedding it and before building a BM25 keyword index. In Anthropic's tests, measured as the share of relevant chunks missing from the top 20 results:

TechniqueRetrieval failure rateReduction vs baseline
Baseline (standard embeddings)5.7%-
Contextual embeddings3.7%35%
Contextual embeddings + contextual BM252.9%49%
Above + reranking1.9%67%

These figures come from Anthropic's own datasets, so your results will differ. The lesson for job seekers is the method: try hybrid search and reranking, then show the before-and-after numbers in your project README.

How do you evaluate an LLM or RAG application?

Evaluate it with a fixed set of questions and expected answers, and score both the retrieval step and the generated answer. Without evaluation, you cannot prove a change helped.

Open-source tools make this easier. The Ragas library lists RAG metrics such as context precision, context recall, response relevancy and faithfulness. In plain terms, you want to know whether the retriever found the right passages, whether the answer actually addresses the question, and whether every claim in the answer is supported by the retrieved text.

A simple evaluation routine:

  1. Write 30-50 realistic questions with reference answers and the source document for each.
  2. Run your pipeline and log retrieved chunks, final answers, latency and token cost.
  3. Score retrieval (did the right chunk appear?) and answers (correct, grounded, relevant).
  4. Change one thing at a time, such as chunk size or adding a reranker, and rerun.
  5. Record results in a table in your README.

What do most guides on RAG and LLM skills get wrong?

Most guides stop at "load a PDF, embed it, ask questions", which every applicant can now do in an afternoon. That demo alone rarely gets interviews.

Common gaps to fix:

  • No evaluation. A project without a test set and scores looks unfinished to hiring managers.
  • Vector search treated as magic. Pure embedding search can miss exact terms such as product codes, error numbers or policy IDs. Hybrid search with BM25 addresses this, as Anthropic's results show.
  • Chunking ignored. Chunk size and boundaries change answer quality. Try at least two strategies and report the difference.
  • Fine-tuning confused with RAG. Fine-tuning changes model behaviour; RAG supplies knowledge at query time. Many business problems need RAG first.
  • Cost and latency skipped. Indian employers, especially startups and IT services firms delivering to clients, care about cost per query. Log it.

How should you present RAG projects on your resume?

Describe the problem, the stack and the measured result in one bullet. For example: "Built a RAG assistant over 1,200 HR policy pages using pgvector and hybrid search; raised retrieval hit rate on a 50-question test set after adding reranking."

Name the specific tools that appear in job descriptions, such as LangChain, LlamaIndex, pgvector or a cloud AI service, so applicant tracking systems pick them up. Link the GitHub repository and, if possible, a short demo video. Our guide on how to list projects on a resume covers bullet writing, and the AI engineer skills guide shows where RAG sits among other expectations.

FAQ

What is RAG in simple terms?

Retrieval-augmented generation is a way of making a language model answer from specific documents. When a user asks something, the system first searches a knowledge base for relevant passages, then adds those passages to the prompt so the model answers using them. It lets companies use their own policies, manuals or records with an LLM without retraining the model, and makes it possible to cite sources.

Do I need machine learning knowledge to build RAG applications?

You need less than for traditional machine learning roles. Strong Python, APIs, SQL and an understanding of embeddings and similarity search will get you building working RAG systems. Some knowledge of how models are trained and evaluated helps you reason about failures, choose embedding models and design tests, so learning machine learning basics alongside is still worthwhile for long-term growth.

Which vector database should I learn first?

Start with whatever keeps your stack simple. pgvector inside PostgreSQL is a good first choice if you know SQL, because it shows exact versus approximate search and index trade-offs clearly. Once you understand the concepts, switching to a managed vector database or a cloud provider's vector search service is straightforward. Employers care more that you understand indexing and recall than which product you used.

How do I evaluate a RAG chatbot?

Create a test set of 30 to 50 realistic questions, each with a reference answer and the source document. Run your pipeline, then check whether the correct passage was retrieved, whether the answer is correct and relevant, and whether every claim is supported by the retrieved text. Libraries such as Ragas provide metrics like context precision, context recall and faithfulness to automate parts of this.

Is RAG better than fine-tuning?

They solve different problems. RAG supplies up-to-date knowledge at query time and makes answers traceable to sources, which suits company documents that change often. Fine-tuning adjusts how a model behaves, such as its format, tone or performance on a narrow task. Many business use cases start with RAG and only add fine-tuning when prompting and retrieval cannot reach the required quality.

What RAG project should I build for my portfolio?

Pick a real document set, such as public government scheme guidelines, a university handbook or an open-source project's documentation. Build question answering with citations, add hybrid keyword and vector search, create a 50-question evaluation set, and record how each change affected scores, latency and cost. Deploy it with a simple interface and document everything in the README.

Final thoughts

The RAG skills that get you hired are the unglamorous ones: chunking, hybrid search, evaluation and cost tracking, shown in one well-documented project. Once it is ready, check your resume against a real generative AI job description with the Asuraa AI resume reviewer.

Related articles

Share this article

Continue Reading

Data Science Career Paths

Explore different career trajectories in data science and find your perfect fit.

Read article →

Building Your DS Portfolio

Learn how to create projects that impress hiring managers and showcase your skills.

Read article →

Salary Negotiation Guide

Get the compensation you deserve with our proven negotiation strategies.

Review Your Resume →