Your End-to-End ML Project, Deployed

The capstone project for Data Science & Machine Learning with Python: From pandas to a Deployed Model — a free Data & Tech course. Pass it and you earn a certificate anyone can verify.

Start the course free Download the project workbook

The project unlocks once you complete the lessons. The workbook is free to download now so you can see exactly what is expected.

What you will submit

One shared link (Google Doc, Notion page or the README itself) containing the live app URL, the public repository URL, and the responsible-ML note.

What to do, step by step

  1. Part 1 — Question and data: choose a real dataset with at least 1,000 rows and a question a Nigerian business, agency or community would pay to answer (predict, classify, segment or forecast). State the question, who would use the answer, and what decision it changes. If you cannot get real data, use the course's loan-repayments or lagos-rent-listings file and say so.
  2. Part 2 — Cleaning and exploration (notebook): document every cleaning step and why — missing values, types, duplicates, outliers, inconsistent categories — with row counts before and after. Produce at least five charts that each make one point, with a sentence of interpretation under each. State the base rate or the typical value of your target.
  3. Part 3 — Models: build a proper baseline (median or majority class), then at least two different model families in scikit-learn Pipelines with a train/validation/test split or cross-validation. Report the metrics that fit the problem (MAE/RMSE/R² for regression; precision, recall, F1, ROC-AUC or PR-AUC and a confusion matrix at a stated threshold for classification). Explain which model you chose and why, and show one check for leakage.
  4. Part 4 — Responsible-ML note (300-500 words): the lawful basis for the data under the NDPA, whether the model profiles people or makes automated decisions about them, which features could act as proxies for ethnicity, religion, gender or state of origin and what you did about it, and a plain-language explanation of how the model reaches a prediction.
  5. Part 5 — Deployment: save the trained pipeline, wrap it in a FastAPI endpoint or a Streamlit app, and deploy it to a free tier (Hugging Face Spaces, Streamlit Community Cloud or similar). Submit the live URL. It must accept new input and return a prediction with the model's version and a short caveat.
  6. Part 6 — Repository: a public GitHub repository with the notebook, the app code, a requirements file with pinned versions, and a README that states the question, the data source and licence, the cleaning summary, the final metrics with the baseline beside them, the deployment link, and what you would do next. No raw personal data and no secrets committed. Put the live app URL and the repository URL in one shared document and submit that link.

Files to work with

Project workbook (fill in, then submit)The whole brief, an evidence checklist, the grading rubric as a self-check and the link-sharing steps in one file. Opens in Word, Google Docs or WPS.Word · 7 KB Loan repayments (CSV, synthetic, messy) — fallback dataset for a classification projectSynthetic and deliberately messy: mixed date formats, naira strings, spelling variants, missing values, an impossible age, duplicates. Use only if you cannot get real data, and say so.CSV · 239 KB Lagos rent listings (CSV, synthetic, messy) — fallback dataset for a regression projectSynthetic and deliberately messy. Real listings scraped with permission or a business's own records make a far stronger project.CSV · 181 KB

How it is graded

CriterionWeight
A real, decision-relevant question on a dataset of 1,000+ rows, with the user and the decision named10%
Cleaning is documented step by step with before/after counts, and exploration produces five or more single-point charts with interpretations and a stated base rate or typical value20%
A baseline plus two model families in Pipelines, honest evaluation with problem-appropriate metrics and a stated threshold, a justified choice, and a leakage check30%
The responsible-ML note addresses lawful basis, profiling and automated decisions, proxy features and their mitigation, and gives a plain-language explanation15%
The deployed app is live, accepts new input, returns a prediction with a version and caveat15%
The repository is public, runnable from pinned requirements, free of secrets and raw personal data, with a README that reports metrics beside the baseline10%

You need 65% overall to pass. A failed submission comes back with feedback and can be revised and resubmitted.

Before you submit: make your link public

If your work lives in Google Drive or Google Docs, open Share → General access and change “Restricted” to “Anyone with the link” (Viewer). Then paste the link into a private browser window: if it opens without sign-in, you are ready. A private link cannot be graded — it is the single most common reason a good project fails.

Frequently asked questions

Do I have to finish Data Science & Machine Learning with Python: From pandas to a Deployed Model before submitting the project?

Yes. The submission screen opens once every lesson in Data Science & Machine Learning with Python: From pandas to a Deployed Model is marked complete. The lessons are where the methods, the Nigerian context and the worked examples the project depends on are taught.

How is the project graded?

An examiner scores each rubric criterion from 0 to 100 and weights them as shown on this page; you need 65% overall to pass. Most submissions are graded within minutes and you get written feedback on what was strong and what to improve.

Can I resubmit if I fail?

Yes. A failed project comes back with feedback; revise the weak parts and resubmit. A passed project is final — the certificate is issued and cannot be re-rolled.

What do I actually submit?

A write-up of what you did (under 5,000 characters) plus a public link to your work — a Google Drive folder, Google Doc, spreadsheet, GitHub repository or video. The link must open without sign-in; a private link cannot be graded. The free project workbook on this page walks you through all 6 steps.

Is the certificate real?

Yes. Passing this project issues a certificate with a unique verification code. Anyone — an employer, a client, a school — can open the verification page and see that the certificate is genuine and which project earned it.

More Data & Tech projects

Ready to earn this certificate?

Data Science & Machine Learning with Python: From pandas to a Deployed Model is free and self-paced. Finish the lessons, complete this project, and the certificate is yours to share.

Start learning free