Start Here

RivoKatalyst Roadmap

Two real routes into a data career in South Africa, plus skill-by-skill learning guides for each role. Pick what fits your situation.

Route 1

The University Path

A degree remains the most reliable route into graduate programmes at major South African employers. Here is what to study, where and what it takes to get in.

Best degrees for data careers

BSc Computer Science / Data Science

Strongest foundation for Data Engineering and AI/ML roles. Heavy on programming, algorithms and maths.

BSc Statistics / Mathematical Statistics

Excellent foundation for Data Science roles. Strong on probability, statistical modelling and analysis.

BCom Information Systems

A great blend of business and tech, strong for Data Analyst and Business Intelligence roles.

BCom Statistics & Finance / Quantitative Finance

Ideal if you are interested in data careers within banking, asset management or insurance.

South African universities to consider

University Relevant Programmes Typical APS Range* Notes
UCT BSc Computer Science, BSc Mathematical Statistics, BCom (Actuarial/IS) ~ 42-50 Highly competitive; strong Maths & Physical Science marks needed.
Wits BSc Computer Science, BSc Computational & Applied Maths, BCom IS ~ 38-45 Strong reputation for Computer Science and Statistics.
University of Pretoria (UP) BSc Information & Knowledge Systems, BSc Computer Science, BCom IS ~ 30-38 Well-regarded IT faculty; multiple data-relevant streams.
Stellenbosch BSc Computer Science, BSc Mathematical Statistics, BCom (Stats & Actuarial) ~ 32-42 Strong applied maths and statistics tradition.
University of Johannesburg (UJ) BSc Computer Science and Informatics specialising in AI, BCom Information Systems 34 One of the few SA undergrad degrees with a dedicated AI specialisation. Maths at 80%+ is non-negotiable for the CS AI degree.
UNISA BSc Computing, BSc Statistics, BCom Information Systems (distance learning) No APS, open admission criteria apply Great option if you are working, in a rural area, or need flexibility.

*APS ranges are approximate and change yearly. Always confirm current requirements directly with each university before applying.

Recommended matric subjects

Mathematics (not Maths Literacy) is essential for almost every programme above. Physical Sciences and/or Information Technology will strengthen your application, especially for Computer Science and Engineering-aligned degrees. English proficiency matters too - data careers involve a lot of written communication.

If You Choose University

Year-by-Year Guide

Year 1: Build the Foundations

Learn: Focus on passing your core maths, statistics and intro programming courses solidly - these are the foundation for everything else.

Projects: Nothing fancy yet, but start a habit of doing small personal coding exercises outside of coursework.

Networking: Join your university’s Data Science, Computer Science or Statistics society. Follow South African data professionals on LinkedIn.

Year 2: Start Building Real Skills

Learn: Deepen Python and SQL skills outside of class. Start learning Power BI or Tableau.

Projects: Build your first 1-2 portfolio projects using public South African datasets (see the Portfolio page).

Networking: Start attending data meetups, webinars and hackathons. Begin building a LinkedIn presence.

Internship strategy: Apply for vacation work or short internships. Even unpaid shadowing opportunities are valuable at this stage.

Year 3: Specialise & Apply for Internships

Learn: Choose a direction - analytics, data science, or engineering - and go deeper. Start learning cloud basics (Azure or AWS).

Projects: Build 2-3 more advanced projects aligned to your chosen specialisation. Put everything on GitHub.

Networking: Reach out directly to professionals for informal chats. Attend career fairs.

Internship strategy: This is the critical year. Apply broadly to internships and graduate programme pipelines - many open applications for the following year around this time.

Final Year & Beyond: Land Your First Role

Learn: Polish your strongest skills and start preparing for technical interviews (SQL tests, case studies, take-home assessments).

Projects: Have a polished portfolio of 4-5 strong projects with clear write-ups, ready to show recruiters.

Networking: Lean on your network for referrals. Most graduate roles get hundreds of applications, and referrals matter.

Internship strategy: Apply to graduate programmes (FNB, Standard Bank, Absa, Discovery, Deloitte, PwC and similar) as early as the application windows open - many close by mid-year for the following year’s intake.

Positioning Yourself, At Any Stage

Portfolio Projects

Real, end-to-end projects using South African data are the single best way to prove your skills. See the Portfolio page for ideas.

LinkedIn

Build a clean profile, post about what you are learning, and connect with data professionals - many SA opportunities are shared there first.

Communities

Join South African data communities (Discord servers, university societies) - opportunities and advice flow through these networks.

Hackathons

Hackathons (often run by banks, universities or organisations like EDSA) are a great way to build projects fast and meet employers directly.

Route 2

The No-Degree Path

Can you get a data job in South Africa without a degree?

Honestly, yes - but it is harder, and it takes longer. Most graduate programmes at large companies require a degree. But many smaller companies, startups and even some larger employers will hire based on a strong portfolio, certifications and demonstrated skills, especially for Data Analyst roles.

The reality of the market: Without a degree, you will need to work harder to prove your skills (a strong GitHub portfolio is essential), you may start in more junior or contract roles, and progress may rely more heavily on networking and referrals than formal applications. It is absolutely possible, but go in with realistic expectations and a plan to keep building skills and credentials over time.

Recommended platforms

DataCamp

Excellent for structured, hands-on Python, SQL and statistics learning at your own pace. Paid subscription, often with free trial periods.

Coursera

Home to respected certificates (e.g. Google Data Analytics, IBM Data Science). Many courses can be audited for free; certificates require payment.

ALX Africa

Intensive, project-based tech and data programmes designed for the African job market, with strong peer-learning communities.

HyperionDev

South African coding and data bootcamps with mentor support, often partnered with universities for accredited short courses.

Explore Data Science Academy (EDSA)

A well-regarded South African data science training provider, often linked with employer partnerships and bootcamp-style intakes.

Free options

YouTube (freeCodeCamp, Alex The Analyst), Kaggle Learn, and official documentation are excellent free starting points for SQL and Python.

Free vs Paid - what is worth it?

Free resources are genuinely good enough to learn the technical skills. SQL, Python, Excel and Power BI can all be learned to a job-ready level for free. Paid courses and bootcamps are worth it mainly for: structure and accountability (especially if you struggle with self-discipline), mentorship and career support, and recognised certificates that help your CV stand out when you have no other credentials.

Certifications employers recognise

  • Google Data Analytics Professional Certificate
  • Microsoft Power BI Data Analyst Associate (PL-300)
  • Microsoft Azure Data Fundamentals (DP-900)
  • IBM Data Science Professional Certificate
  • AWS Certified Cloud Practitioner
  • SQL certifications from accredited platforms (e.g. DataCamp, Coursera)

Realistic Timelines for the No-Degree Path

Everyone progresses at a different pace, but here is a realistic guide to what you should aim to achieve.

3

3 Months

Comfortable with Excel/Google Sheets and basic SQL. Completed an introductory course (free or paid). Started exploring public datasets.

6

6 Months

Confident with SQL and intermediate Excel. Built your first 1-2 portfolio projects with a public write-up on GitHub. Started learning Power BI and basic Python.

12

12 Months

3-4 polished portfolio projects covering SQL, Power BI and Python. At least one certification completed. Actively applying for junior Data Analyst roles and internships.

18

18 Months

Working in (or very close to landing) a junior data role. Beginning to specialise - deepening Python and statistics for Data Science, or SQL and cloud for Data Engineering.

Role Roadmap

Data Analyst Roadmap

Step-by-step guide covering what to learn, why it matters, what to build, and how to know when you are ready to move on.

Months 1-3Foundation

SQL, Excel and the Analyst Mindset

SQL is the analyst's first language. Master it before anything else. Pair it with Excel/Sheets for fast exploratory work. Everything in this role starts with understanding data shape and asking the right questions.

SQL (PostgreSQL / BigQuery)Excel / Google SheetsGoogle Analytics 4Basic Python (pandas)
  • Write 20 SQL queries against a public dataset (e.g. SA census data, Kaggle retail dataset)
  • Build a sales dashboard in Excel with pivot tables and charts
  • Complete Google Analytics Certification (free)
Interview tip:
  • Expect a live SQL test: GROUP BY, window functions, JOINs, subqueries
  • Be ready to explain how you find data quality issues (NULLs, duplicates, outliers)
Months 4-6Visualisation

Dashboards and Storytelling

Analysts live in BI tools. Learn Power BI first (dominant in SA corporates). Understand the difference between a chart that shows data and one that drives a decision.

Power BITableau (bonus)Looker / Looker StudioPython: matplotlib, seaborn
  • Build a multi-page Power BI dashboard connected to a real dataset
  • Create a storytelling deck from an analysis: one insight, one chart, one recommendation per slide
  • Replicate a Financial Mail or Stats SA chart in Python
Interview tip:
  • Interviewers show you a bad dashboard and ask how you'd improve it
  • Practice explaining a chart to a non-technical audience in under 60 seconds
Months 7-9Analytics and Stats

Python, A/B Testing and Statistical Thinking

Python unlocks automation and more rigorous analysis. Stats knowledge separates junior analysts from senior ones. Understand when a difference is real vs noise.

Python: pandas, numpy, scipyJupyter NotebooksA/B testing fundamentalsBasic regression
  • Analyse a marketing campaign dataset: segment users, measure conversion rates, run a chi-squared test
  • Build an automated monthly report in Python that exports to Excel
  • Implement a simple linear regression and interpret the output in plain English
Interview tip:
  • Be ready to explain p-values without jargon to a stakeholder
  • Common case: "We ran a promo. Did it work?" Walk through sample size, control group, statistical significance
Months 10-12Job Ready

Portfolio, GitHub and Job Applications

Three polished end-to-end projects beat ten half-finished ones. Each project should have: a business question, a dataset, an analysis, a visualisation and a recommendation.

GitHubPower BI / Tableau PublicLinkedInGoogle Data Analytics Certificate (optional)
  • Project 1: Customer segmentation analysis (RFM or clustering)
  • Project 2: Financial performance dashboard with month-on-month trends
  • Project 3: A/B test analysis for a fictional e-commerce promotion
Interview tip:
  • Apply to banks, retailers and telecoms in SA: FNB, Woolworths, MTN, Discovery all hire junior analysts
  • Your SQL test in interviews will cover: CTEs, window functions (ROW_NUMBER, LAG, RANK), date arithmetic
Role Roadmap

Data Scientist Roadmap

Step-by-step guide covering Python, statistics, machine learning, experimentation, and your path to a first data science role.

Months 1-3Python First

Python and Statistics: the Non-Negotiables

Every DS interview tests Python. Not SQL. Not Excel. Python. Get fluent in numpy, pandas and scipy before touching any ML library. Statistics is the foundation: distributions, hypothesis testing, correlation vs causation.

Python (numpy, pandas, scipy)Jupyter NotebooksGitStatistics: probability, distributions, CLT
  • Complete the entire pandas documentation exercises for groupby, merge and reshape
  • Implement Bayes' theorem from scratch on a real classification problem
  • Write a hypothesis test without using a library function; then verify with scipy
  • Study: z-test, t-test, chi-squared, ANOVA; know when to use each
Interview tip:
  • DS interviews open with: "Write a function to compute X in Python." No libraries.
  • Know time complexity. O(n) vs O(n^2) matters for large datasets.
  • Be able to explain the Central Limit Theorem in simple terms - it comes up constantly
Months 4-6Core ML

Machine Learning Fundamentals

Learn the algorithms from the ground up with scikit-learn. Understand what each algorithm assumes about data, when it breaks, and how to evaluate it properly. Interviewers probe here deeply.

scikit-learnLinear and Logistic RegressionDecision Trees / Random ForestsXGBoost / LightGBMCross-validation / GridSearchCV
  • Implement linear regression with gradient descent from scratch in numpy; then replicate with scikit-learn
  • Build a churn prediction model; use 5-fold cross-validation; tune hyperparameters
  • Compare 4 algorithms on the same problem and write a clear comparison of results
  • Handle an imbalanced dataset (see technique box below) and show metric improvement
Imbalanced datasets: what interviewers expect you to know
  • Why accuracy fails: 99% accuracy on 1% minority class means the model predicts majority every time
  • Metrics to use instead: Precision, Recall, F1, ROC-AUC, PR-AUC (know when each matters)
  • Resampling: SMOTE (synthetic oversampling), random undersampling, ADASYN
  • Algorithm-level fixes: class_weight="balanced" in scikit-learn, scale_pos_weight in XGBoost
  • Threshold tuning: shift the decision boundary away from 0.5 using the precision-recall curve
  • Always: stratify your train/test split to preserve class distribution
Interview tip:
  • Common question: "How would you handle a dataset where 2% of transactions are fraud?"
  • Walk through your entire pipeline: EDA, baseline model, metric choice, resampling strategy, evaluation
  • Never say "accuracy" as your primary metric for imbalanced data in an interview
Months 7-9Advanced ML

Ensemble Methods, Feature Engineering and Experimentation

The difference between a graduate and a mid-level data scientist is feature engineering and rigorous experiment design. Learn what makes features predictive. Learn how to run a proper A/B test.

XGBoost / LightGBMSHAP (model explainability)Feature selection: RFE, LASSOA/B testingPython: statsmodels
  • Engineer 10 features from a raw transactional dataset and show their predictive power
  • Use SHAP values to explain which features drive a model's predictions (critical for SA banking sector)
  • Design and analyse a simulated A/B test: power calculation, sample size, test duration, results
  • Implement a stacking ensemble from scratch
Interview tip:
  • Interview case: "You trained a model last year. It worked. Now performance has dropped. Why?"
  • Talk through: data drift, concept drift, schema changes, seasonality, upstream data pipeline issues
  • "What features would you engineer for credit risk?" is a common SA banking interview question
Months 10-12Deep Learning and NLP

Neural Networks, NLP and the ML Pipeline

Most SA data science roles are ML-first, not deep learning. But knowing the basics of neural networks and NLP makes you competitive. Focus on applying these tools, not re-implementing them.

TensorFlow / PyTorch (basics)scikit-learn PipelinesNLP: spaCy, NLTK, transformers (HuggingFace)MLflow (experiment tracking)SQL (secondary, data retrieval)
  • Train a feedforward neural network on a tabular dataset; compare to XGBoost
  • Build a sentiment classifier on SA product reviews using a pretrained transformer
  • Package a model into a scikit-learn Pipeline with preprocessing included
  • Track 3 experiment runs in MLflow: log parameters, metrics and artifacts
Interview tip:
  • Know the bias-variance tradeoff cold. Draw it. Explain it. Relate it to overfitting vs underfitting.
  • Know regularisation: L1 (LASSO, feature selection), L2 (Ridge, weight shrinkage), dropout in neural nets
  • "Walk me through a model you built end to end" is the most common DS interview question. Prepare one story
Months 12+Job Ready

Portfolio, SQL Basics and Applications

Your portfolio should show the full DS workflow on real problems. SQL is tested in every SA data science interview - window functions, CTEs, complex GROUP BY and date arithmetic are all fair game. Python and ML concepts dominate, but do not walk in without solid SQL.

GitHubKaggle profilePower BI (optional but valued in SA)SQL: window functions, CTEs, subqueries
  • Portfolio project 1: Full credit scoring model (EDA, feature engineering, XGBoost, SHAP, threshold optimisation)
  • Portfolio project 2: Time series forecasting (ARIMA or Prophet on SA economic data)
  • Portfolio project 3: NLP classifier or recommendation system with writeup
Interview tip:
  • SA companies hiring DS: FNB, Absa, Standard Bank, Discovery, Old Mutual, Capitec, MTN, Allan Gray, Coronation
  • SQL is tested at every SA bank DS interview: window functions (ROW_NUMBER, LAG, RANK, SUM OVER), CTEs, complex GROUP BY and joins
  • FirstRand Quant Graduate Programme (Workday): strong maths, stats and Python required
  • Discovery hiring: heavy on ML for insurance pricing; prepare actuarial-style case studies
Role Roadmap

Data Engineer Roadmap

Step-by-step guide covering SQL, Python, pipelines, cloud infrastructure and the tools that power modern data stacks.

Months 1-3SQL and Python

SQL Mastery and Python for Data

Data engineers write SQL all day. Master it at an advanced level: window functions, CTEs, query optimisation, partitioning. Python automates and transforms. Start here before touching any cloud or pipeline tool.

SQL (PostgreSQL, BigQuery)Python: pandas, sqlalchemyGitLinux command line basics
  • Write 30 SQL queries covering GROUP BY, window functions, recursive CTEs, performance tuning
  • Build a Python script that pulls data from a public API and loads it into a local PostgreSQL database
  • Practice query optimisation: EXPLAIN ANALYZE, indexes, partitioning
Interview tip:
  • DE interviews always include a SQL test. Advanced window functions are expected.
  • Know: ROW_NUMBER vs RANK vs DENSE_RANK, LAG/LEAD, running totals with SUM OVER
Months 4-6Data Modelling

Warehousing, dbt and Data Modelling

Learn how data is structured for analytical workloads: star schema, dimensional modelling. dbt is now standard for transformation in modern data stacks. Understand Kimball methodology.

dbt (data build tool)Star schema / dimensional modellingBigQuery / Snowflake / RedshiftData warehouse concepts
  • Set up a dbt project on a free BigQuery sandbox; write 5 models with tests and documentation
  • Design a star schema for a fictional e-commerce company (fact_orders, dim_customers, dim_products)
  • Implement incremental models and understand when to use them vs full refresh
Interview tip:
  • Know the difference between OLTP and OLAP and why it matters for schema design
  • Be ready to explain slowly changing dimensions (SCD Type 1 vs Type 2)
Months 7-9Pipelines

Orchestration, Spark and Batch Pipelines

Airflow is the industry standard orchestrator. Spark handles scale. Understand how to build reliable, idempotent pipelines that fail gracefully and recover cleanly.

Apache AirflowApache Spark (PySpark)DockerCloud storage: GCS / S3 / Azure Blob
  • Build an Airflow DAG that extracts data daily from an API, transforms it and loads it to BigQuery
  • Write a PySpark job that processes a 10M-row CSV and outputs aggregated results
  • Containerise a pipeline with Docker and docker-compose
Interview tip:
  • System design: "How would you build a pipeline that ingests 1M events per day?"
  • Walk through: ingestion layer, storage, transformation, scheduling, monitoring, failure handling
Months 10-12Streaming and Cloud

Kafka, Cloud Infrastructure and Real-time Data

Real-time data pipelines use Kafka or Pub/Sub. Cloud platforms (GCP, AWS, Azure) are where most production pipelines run. Learn one cloud well rather than all three superficially.

Apache KafkaGCP (BigQuery, Dataflow, Pub/Sub)Terraform (IaC basics)Great Expectations (data quality)
  • Set up a local Kafka cluster; produce and consume messages with Python
  • Deploy a Dataflow streaming pipeline on GCP free tier
  • Write data quality checks with Great Expectations on a pipeline output
Interview tip:
  • SA cloud hiring: most companies use AWS or Azure. GCP is common at fintechs.
  • Know the difference between batch and stream processing and when each applies
Role Roadmap

AI / ML Engineer Roadmap

Step-by-step guide covering deep learning, NLP, MLOps, model deployment and building production AI systems.

Months 1-3ML Foundations

Python, ML and the Math that Matters

ML engineering is applied ML. You need to understand models well enough to deploy, optimise and debug them in production. Start with the foundations: Python at a professional level, key algorithms, linear algebra and calculus intuition.

Python (numpy, pandas, scikit-learn)Linear algebra (vectors, matrices, dot products)Calculus (derivatives, chain rule, gradient descent)Git and virtual environments
  • Implement linear regression, logistic regression and a neural network from scratch in numpy
  • Study backpropagation: understand the chain rule computation graph
  • Complete fast.ai Practical Deep Learning for Coders (free, project-based)
Interview tip:
  • Interviewers ask you to derive gradient descent. Know it.
  • Be able to explain backpropagation to someone who has not heard of it
Months 4-6Deep Learning

Neural Networks, CNNs, RNNs and Transformers

Deep learning is the engine of modern AI. Understand architecture choices, training dynamics and common failure modes. PyTorch is preferred in research; TensorFlow / Keras in industry.

PyTorchTensorFlow / KerasHuggingFace TransformersW&B / MLflow (experiment tracking)CUDA basics (GPU training)
  • Train a CNN image classifier in PyTorch; log experiments with Weights and Biases
  • Fine-tune a HuggingFace BERT model on a text classification task
  • Train an LSTM for time-series prediction and compare to a Transformer baseline
  • Understand attention mechanism: implement scaled dot-product attention from scratch
Interview tip:
  • Know vanishing gradients: why they happen, how batch norm and residual connections fix them
  • Know attention in transformers at the level you can explain self-attention step by step
Months 7-9LLMs and GenAI

Large Language Models, RAG and Prompt Engineering

LLMs are now a core skill for AI engineers. Understand how to work with them via APIs, fine-tune smaller models and build retrieval-augmented generation (RAG) systems that ground responses in real data.

OpenAI / Anthropic APIsLangChain / LlamaIndexVector databases: Pinecone, Chroma, WeaviateRAG architecturePrompt engineering
  • Build a RAG system over a set of PDF documents using LangChain and Chroma
  • Fine-tune a small open-source LLM (Mistral, LLaMA) on a domain-specific dataset using LoRA
  • Evaluate LLM outputs: implement precision, recall and hallucination detection metrics
Interview tip:
  • Know the difference between fine-tuning and RAG: when each approach is appropriate
  • Interviewers ask: "How do you prevent hallucination in a production chatbot?"
Months 10-12MLOps

Model Deployment, Monitoring and Production ML

Getting a model to 80% accuracy is 20% of the work. Serving it reliably, monitoring it, retraining it when it drifts and scaling it to millions of requests is the other 80%. MLOps is what separates academic projects from production AI.

FastAPI (model serving)Docker and KubernetesMLflow / BentoML / SeldonPrometheus / Grafana (monitoring)CI/CD for ML (GitHub Actions)
  • Wrap a trained model in a FastAPI endpoint; containerise with Docker; deploy to a cloud VM
  • Set up model monitoring: log predictions, feature distributions and alert on drift
  • Build a CI/CD pipeline that retrains a model on new data and runs evaluation gates before deployment
  • Implement A/B testing between two model versions in production
Interview tip:
  • System design: "How would you serve a recommendation model to 1 million users?"
  • Cover: latency vs throughput, model caching, feature stores, canary deployments, rollback strategy

Not sure where to start?

Answer five quick questions and get a personalised step-by-step plan built around your background, current skills and time commitment.

Build My Career Plan