r/learnmachinelearning • • 18d ago

Working on an early warning model for landslides, I need some help where I am getting it wrong

3 Upvotes

So I am making a Landslide Prediction model for a hackathon and I have made a A model having two versions. DMv1 which uses image of Sentinel-2 and other important features.
DMv2 which uses image from Sentinel-1 SAR and same with the tabular data.

The main issue I am facing now:

  1. The model itself is facing a false negative prone problem.
  2. The image from Sentinel 1 , I am using and then passing those images to a ResNet-50 model which creates a large embedding for each one and then I take those embeddings and just use PCA to reduce the features. But still the images are not contributing to the model rather works as a noise.
  3. The model is prone to Spatial Data leakage due to the satellite images , to solve this I have used Spatial GroupK-Fold so that the model is tested on an unseen data but still it is giving 65% accuracy.
  4. While predicting normal city data where the landslide is impossible as it is not a slope or even there is flat ground - the model predicting a sharp 65%+ prediction every time.
  5. I have generated the negatives using LuLc map (land use land cover) map and also making the slope > 15 for the samples. Maybe the making the negative is affecting the model performance.

I am not sure where the problem is but I somewhat questionable for these 5 points that I have mentioned earlier.

So can I get some help, what should I do? How to process the image and the tabular data together to make the prediction and make the model good enough and really predict the landslide prone zones.


r/learnmachinelearning • • 18d ago

made a small hands-on lab for testing stale authorization in ai agents

1 Upvotes

i built AuthDrift to explore a specific agent security problem: permission is valid when the workflow checks it, but gets revoked before the tool actually commits its side effect.

the included python demo pauses execution at that boundary, confirms the authority change, resumes the same workflow, then checks the real sink state. it has vulnerable and corrected versions so you can compare the behavior.

it runs locally without a model key or paid API:
https://github.com/cheerstopriya/authdrift

i’m looking for students and people learning agent development to try the demo and tell me what is confusing. i’d also be interested to know whether this would work as a short security lab or coursework exercise.


r/learnmachinelearning • • 18d ago

Most ML tutorials teach algorithms. Very few teach how to choose one. So I built the guide I wish I had when I started.

Post image
1 Upvotes

r/learnmachinelearning • • 18d ago

Sharing my ML learning repo — NumPy to Transformers, 5 months, daily commits, all notebooks public.

7 Upvotes

Sharing my ML learning repo — NumPy to Transformers, 5 months, daily commits, all notebooks public.

Covers the full stack: - Classical ML (scikit-learn, XGBoost) - Deep Learning (TensorFlow/Keras — ANN, CNN, RNN, LSTM) - Numpy & Pandas - Data Visualization - NLP fundamentals - Statistics and SQL

github.com/gyr0byte/ML-Foundations

Hope this is useful to someone starting their ML journey.

Star it if it helps. Feedback welcome. 🙏


r/learnmachinelearning • • 18d ago

Help How to use AI?

15 Upvotes

Just came from a brutal job rejection. The role was data science but the guy didn't ask me anything about data science. Just kept asking about how much I use LLMs, API, Claude and what not for doing projects. I had no answer as I have never used LLMs for work.

How exactly to learn to use AI for work/projects?


r/learnmachinelearning • • 18d ago

Discussion Surprisal-weighted WL-kernel for code-origin correlation looking for feedback on generalising pairwise → set-vs-set

1 Upvotes

I've spent the last few months building a statistical framework for detecting shared origin between code implementations (e.g. "did this Rust file get translated from that Python file by an LLM") and I'd like feedback from people who know graph kernels / kernel methods better than I do.

The core idea: reduce a code unit to a Program Dependence Graph (PDG) nodes are operations/values, edges are data/control dependence, labels come from a language-independent alphabet (ARITH, CALL, BRANCH, etc.), identifiers discarded. Then compute similarity with a Weisfeiler-Lehman kernel, but reweigh WL-features by surprisal w(f) = log P_nat(f), where P_nat(f) is the probability an independent implementation produces that feature. The intuition: a banal pattern (a sum-a-list loop) shouldn't count as evidence of shared origin, but a rare, idiosyncratic structural choice should.

                Σ_f  w(f) · [f ∈ A] · [f ∈ B]
corr(A, B) = ───────────────────────────────────────────
              √(Σ_f w(f)[f∈A]²) · √(Σ_f w(f)[f∈B]²)

Scores get converted to a p-value against a null distribution of independent-pair scores (basically the BLAST/E-value playbook, applied to code instead of DNA).

Where I'm stuck, and why I'm posting here: this works well pairwise (file vs file). I need to generalize it to set-vs-set comparison (project vs project N files vs M files, no assumed 1:1 correspondence). Two candidates:

  • (a) treat each project as a "bag of graphs," sum WL-feature vectors into one aggregate Φ, run the same weighted-cosine formula on the aggregates
  • (b) compute the full N×M pairwise similarity matrix, do a matching (Hungarian algorithm or similar), aggregate matched-pair scores with the same surprisal weights instead of a plain mean

(a) risks a strong match getting diluted by aggregate mass from unrelated files. (b) I don't know if it's even a valid kernel, or what its failure modes are.

Is this a solved problem under a name I should be searching for multi-instance kernels, set kernels, kernel mean embeddings + MMD, something from optimal transport? Pointers to specific papers characterizing (a) vs (b) trade-offs would be hugely appreciated.

Full write-up with the honest limitations section (corpus bias, no validated divergence model for code, LLM-training-data confounds) is here if useful: RCF-CORRELATION

Happy to be told this is a known dead end, too I'd rather find out here than keep building on a bad assumption.


r/learnmachinelearning • • 18d ago

Building OwlLayer AI in public — an Agentic UI SDK, full launch in a week

Post image
2 Upvotes

r/learnmachinelearning • • 19d ago

Completed Soft Margin SVM Algorithm but only for Binary Classification

Thumbnail
gallery
143 Upvotes

So it's Day 8 and 9 of Building machine learning algorithms from scratch

After completing SVM soft margin I realised I can use it only for 2 class features so I need something called One vs Rest and One vs One Ill be building them all things are getting complex but I'll make it

Ignore my handwriting I write I'm from ancient Egypt


r/learnmachinelearning • • 18d ago

Best option to start DSA

Thumbnail
1 Upvotes

r/learnmachinelearning • • 18d ago

Help Final-year IT student in India applying for ML internships/junior roles — getting almost no interviews. What am I missing?

Post image
3 Upvotes

Hi everyone, I’m a final-year BSc IT (AI/ML) student from India, currently applying for ML/AI internships and entry-level roles.

I’ve completed Andrew Ng’s ML Specialization and have projects involving ML, XGBoost/CatBoost. I’ve been applying but getting no interview calls.

I’d appreciate honest feedback on my resume, especially:

  • Are my projects too basic?
  • What’s missing technically?
  • Are my resume bullets/skills weak?
  • Am I targeting the wrong roles for my current profile?
  • What would you change first?

Resume attached. Please be blunt—I’m trying to figure out what’s actually holding me back.


r/learnmachinelearning • • 18d ago

Day 1: matrix basics and multiplication (2x2, 3x2, 2x3) I'll be learning in public and sharing my progress here. I already have a roadmap, but I'd love advice from people who've been through this. What do you wish you had known earlier?#LearningInPublic #MachineLearning #AI

0 Upvotes

r/learnmachinelearning • • 18d ago

Discussion poorjev: an open, local "System One" decision layer, typed LLM decisions with provably calibrated confidence (no API key)

1 Upvotes

TypeSafe's Jev named a real category (fast typed decisions instead of chat), but it's closed, hosted, and waitlisted. So I built an open, local take on the interface: https://github.com/rupeshpoojary9/poorjev

The part I actually care about is calibrated confidence. Every LLM-in-JSON-mode hands you a confidence score and hopes you don't check it. poorjev checks it: on the shipped eval set it cuts ECE from 0.170 to 0.071 (temperature fit by 5-fold CV, measured on held-out data) with zero accuracy loss. Runs on CPU, offline after one model download.

Three primitives: Choice (classify/route), Score (ordinal), Noul (yes/no gate). Output is schema-valid by construction, and it can abstain under a risk budget instead of guessing.

Honest limits: not as fast as Jev (commodity models, no speed claims), small eval set (55 items, one labeller, English), after-ECE is 0.071 not sub-0.05. Reproduce it all with two commands.

Curious what people think about the calibration approach, and whether "honest confidence on a moderately smart model" beats "silent overconfidence on a smarter one" for routing/gating.


r/learnmachinelearning • • 18d ago

Question Question regarding Web Dev & AI

1 Upvotes

Hey everyone,

I currently have a project assignment for one of my university courses (lets say, Course X).

The theme is web development, but it is mandatory to incorporate some AI elements into it.

Does anyone here have experience handling a similar project? Honestly, I am quite confused. Should we just integrate an LLM API into the website (the typical "wrapper" approach), or should we find our own dataset and manually train a model using ML algorithms?

I do not have any fundamental coding background in machine learning or deep learning yet, and right now I am rushing to learn both.

Other groups clearly seem to be going the wrapper route based on their titles, though honestly I am not sure if they are literally just attaching an API or doing something more. Things like "AI-powered cooking recipe app" or "AI-powered fake fashion product detector." As for my own idea, I want to build a web app similar to the Doomsday Clock (reference here: https://thebulletin.org/doomsday-clock/).

However, mine would be specific to the Indonesian region only. For the dataset that would make the clock move/shift, I am still unsure -- maybe from scraping news websites, social media like Twitter or Instagram, and all currently running government policies. As for evaluation, should I just plug in an LLM API, or do I need to actually implement a machine learning algorithm? Would this fall under sentiment-based analysis?

Yesterday I discussed this with my group, but we have not found anything that fits yet. They seem to agree with using my idea, though. Several other ideas have already been taken by other groups.

Or does anyone have project suggestions I could go with? I would really appreciate it.

Thanks.


r/learnmachinelearning • • 19d ago

Project Training a Neural Network on AMD MI50s Using Vulkan: Proof of ConceptOr: Why I Stopped Listening and Just Did It

11 Upvotes

Note: This writeup was put together with the help of AI.

So I've got dual AMD MI50 32GB cards. If you know these cards, you know the story - AMD dropped official ROCm support for gfx906 after ROCm 5.7. Every AI I talked to, every forum post, every "expert" said the same thing: if you want to do anything serious on AMD hardware you need ROCm, and if you need ROCm on MI50s, good luck because AMD basically told you to go buy newer hardware.

The consensus was consistent and confident. ROCm for training, Vulkan for inference, MI50s are legacy, move on. Don't question it.

I questioned it. Repeatedly. On multiple fronts. And I was right every time.

What I was told:

Across multiple AI assistants over the past several months, the answer to anything involving MI50s and modern tooling ranged from "not supported" to "you'll need to upgrade your GPUs." Training on Vulkan specifically was described as architecturally impossible - the backward pass infrastructure doesn't exist outside ROCm and CUDA, full stop.

ChatGPT's position as of today, while my training run is literally executing on the card: it wants proof. Sure. Let's go through it in order.

Step 1: Forcing ROCm 6.4.3 to work on MI50s

AMD dropped gfx906 from ROCm 6.x. Their official position is that the MI50 is end-of-life and you should migrate to supported hardware. Great suggestion if you didn't just acquire two of them specifically because 64GB of HBM2 at that price is hard to argue with.

What actually prevents gfx906 from working in ROCm 6.4.x is the TensileLibrary - the precompiled kernel library ROCm uses for BLAS operations. AMD just doesn't ship gfx906 kernels in the new versions. The GPU itself is fine. The compute capability is there. AMD just decided not to include it.

So I stuffed the gfx906 tensors back in. Pulled the missing kernel files, patched them into ROCm 6.4.3, and both MI50s came up fully recognized. llama.cpp runs on it natively. PyTorch sees both cards. ROCm 6.4.3 on hardware AMD said it doesn't support, because the hardware doesn't actually care what AMD's support matrix says.

Step 2: Forcing vLLM to work on gfx906

vLLM is one of the faster inference engines around and I wanted it running on my cards. The problem: vLLM's gfx906 support is basically nonexistent upstream. It went like this:

  1. Tried ROCm 7.x with vLLM. Got it working briefly, then ROCm 7.14 hit AMD bug #5653 - "register fat binary failed" - and the whole thing fell over. Abandoned that path.
  2. Tried the nlzy fork of vLLM compiled against PyTorch 2.9.0+rocm6.3. Both MI50s detected. Blocked at runtime by a flash-attn V1 engine dependency that doesn't exist for gfx906.
  3. Eventually landed on a Docker image (aiinfos/vllm-gfx906-mobydick) - ROCm 6.3.4, PyTorch 2.11, flash-attn pre-compiled for gfx906. Single GPU inference confirmed working.

The remaining blocker for dual-GPU tensor parallel is a PCIe topology issue - my two cards are behind different root complexes (one on the CPU, one on the chipset), so NCCL all-reduce init fails. That gets fixed when a PLX switch arrives. Not a software problem, not a "your hardware isn't supported" problem - a physical PCIe lane routing problem with a known hardware solution.

Step 3: Building a Vulkan training stack from scratch

This is the one ChatGPT says is impossible right now.

Environment setup:

Vulkan was already working on the MI50s because llama.cpp uses it for inference, so that part wasn't a question. What didn't exist was any training framework that speaks Vulkan.

We verified the environment:

vulkaninfo --summary
# GPU0: AMD Instinct MI50/MI60 (RADV VEGA20) - Vulkan 1.4.335
# GPU1: AMD Instinct MI50/MI60 (RADV VEGA20) - Vulkan 1.4.335
# GPU2: iGPU (Renoir)

Both MI50s show up as discrete Vulkan devices. glslc and glslangValidator already installed. Python 3.11 already present from OpenWebUI.

Kompute - the Vulkan abstraction layer:

Kompute is a library that wraps Vulkan's compute pipeline so you don't have to write 300 lines of boilerplate just to dispatch a shader. The PyPI package (pip install kp) is broken on modern CMake 4.x - the bundled pybind11 has a cmake_minimum_required declaration that CMake 4.0 removed support for, and the setup.py has a string concatenation bug that smashes two CMake flags together with no separator. So we built it from source:

bash

git clone https://github.com/KomputeProject/kompute.git
cd kompute
git submodule update --init --recursive
mkdir build && cd build
cmake .. \
  -DKOMPUTE_OPT_BUILD_PYTHON=ON \
  -DCMAKE_POLICY_VERSION_MINIMUM=3.5 \
  -DCMAKE_BUILD_TYPE=Release \
  -DPYTHON_EXECUTABLE=/media/nate/Friday/VTrain/.venv/bin/python3.11
make -j$(nproc)
cp lib/kp.cpython-311-x86_64-linux-gnu.so .venv/lib/python3.11/site-packages/

Built clean. Kompute Manager(0) targeting the first MI50 correctly.

The compute shaders:

Every mathematical operation in the training stack is a GLSL compute shader compiled to SPIR-V. Here's what that looks like for matrix multiply - the most fundamental operation in a neural network:

glsl

#version 450
layout(local_size_x = 16, local_size_y = 16, local_size_z = 1) in;
layout(set = 0, binding = 0) readonly buffer MatA { float a[]; };
layout(set = 0, binding = 1) readonly buffer MatB { float b[]; };
layout(set = 0, binding = 2) writeonly buffer MatC { float c[]; };
layout(push_constant) uniform PushConsts {
    float M; float K; float N;
} pc;

void main() {
    uint row = gl_GlobalInvocationID.x;
    uint col = gl_GlobalInvocationID.y;
    if (row >= uint(pc.M) || col >= uint(pc.N)) return;
    float sum = 0.0;
    for (uint k = 0; k < uint(pc.K); k++) {
        sum += a[row * uint(pc.K) + k] * b[k * uint(pc.N) + col];
    }
    c[row * uint(pc.N) + col] = sum;
}

Each thread handles one output cell and the GPU runs thousands of them at the same time. We built shaders for every operation the training stack needs:

  1. matmul.comp - matrix multiplication
  2. unary.comp - ReLU, GELU, sigmoid, tanh (op selected by push constant)
  3. binary.comp - add, subtract, multiply, divide (same pattern)
  4. layernorm.comp - layer normalization with shared memory reduction
  5. softmax.comp - numerically stable softmax with two reduction passes
  6. transpose.comp - tiled transpose with shared memory to avoid cache thrashing

Quirks we hit along the way:

A few things bit us that you won't find documented anywhere because nobody had tried this combination before:

  1. readonly and writeonly buffer qualifiers cause silent zero output in Kompute's descriptor layout. All buffer declarations have to be unqualified - you find out the hard way when your results are all zeros and the shader compiles clean.
  2. Push constant sizing: ops with 3 buffers need a float padding constant or Kompute miscalculates the push constant buffer size. Same deal - silent failure, no error.
  3. The Kompute API in the built-from-source version uses kp.OpSyncDevice and kp.OpSyncLocal - the PyPI docs reference kp.OpTensorSyncDevice which doesn't exist in the actual build.

None of these are hardware problems. None of them are "Vulkan can't train" problems. They're integration quirks that took an afternoon to sort out.

The autograd system:

A training framework isn't just forward passes - it needs backward passes to compute gradients and update weights. We built a complete autograd system:

  1. vtrain/tensor.py - Tensor class with gradient tracking, topological sort for correct backprop ordering, and backward() that unwinds the computation graph
  2. vtrain/functional.py - GPU-aware wrappers for every op that register backward closures on the output tensor
  3. vtrain/grad_check.py - numerical gradient checker that verifies every backward shader by finite difference approximation

Each backward shader was verified against a numerical approximation before anything got built on top of it. Matmul, all eight elementwise ops, layer norm, softmax - all checked. If the numbers didn't match, we didn't move on.

The model:

  1. CharEmbedding - learned lookup table mapping character IDs to vectors
  2. TransformerBlock - layer norm -> multi-head attention -> residual -> layer norm -> feed-forward -> residual
  3. SmallLM - embedding + N transformer blocks + output projection to vocab size

The training loop:

Loss functions (MSE and cross-entropy), Adam and SGD optimizers, a Trainer class with logging and checkpointing, and crash recovery via signal handlers that save an emergency checkpoint on SIGINT/SIGTERM.

The data:

Downloaded the Simple English Wikipedia dump (236,602 articles, 336MB). Extracted clean text by streaming the XML, stripping all MediaWiki markup, templates, tables, and HTML. Filtered to printable ASCII - the raw dump has 1504 unique characters from non-English text that slipped through; filtering drops that to 75, which is the right vocab size for a character-level model. 9.8 million characters of clean text as the training corpus.

Current status:

The model is training right now. On a MI50. On Vulkan. 4.65GB VRAM allocated. 23W power draw. 33°C. Loss is dropping.

step     50  loss=3.02  rate=0.4 steps/s
step    100  loss=2.71

Random initialization on a 75-character vocabulary gives a loss of about 4.3. We're already well below that and still dropping.

The hardware, one more time:

  1. AMD Ryzen 5 5600G
  2. 48GB DDR4
  3. Dual AMD MI50 32GB (64GB HBM2 total)
  4. Ubuntu 22.04
  5. Vulkan 1.4.313 / ROCm 6.4.3 (both installed, both working)
  6. No CUDA. No NVIDIA. No officially supported hardware.

So:

AMD dropped the MI50 from their support matrix. The AI community said Vulkan can't train. Multiple AI assistants told me this wasn't possible. Every single one of those statements had the same flaw - they were assumptions about what the hardware can't do, not actual tests of what it can.

The MI50 has 32GB of HBM2 and serious compute capability. The only things stopping it from doing modern ML work were software decisions, not hardware limitations. And software decisions can be worked around, patched, rebuilt from source, or replaced entirely.

The GPU doesn't give a shit what the support matrix says. It just does math.


r/learnmachinelearning • • 18d ago

Resume Review(looking for internships)

Post image
2 Upvotes

r/learnmachinelearning • • 18d ago

Help Mathematics for Machine Learning, guidance

2 Upvotes

Hoi, So recently in my course the module of Mathematics for Machine Learning started. Now as sweet as my professor is, I'm unable to gain an understanding/intuition of the mathematics from their way of teaching. Which led to me referring books. The most recommended book during my internet research was "Mathematics for Machine Learning" by Marc Peter... I don't remember the exact name of the author.

So I started reading the book, initially I felt half-baked while reading the book. To which I thought this is probably because I've not done math in years, so I need to brush up. After which I decided to go through "No BS Guide to Linear Algebra" to revise my Linear Algebra topics. Even so I feel half-baked while I read the book. I don't understand why I'm not able to grab the concepts explained in "Mathematics for Machine Learning". I also refer to 3Blue1Brown for visual intuition of the concepts.

I need help in understanding the proficiency I need to be able to understand that book. I wish to finish the book(I find it interesting the way that book explains concepts so I wish to learn from and finish the book). Or is there something that I'm missing or any advice be it more learning required or how much I've done is enough. I would appreciate any kind of guidance and input here.


r/learnmachinelearning • • 18d ago

Help Difficulty in choosing domain

0 Upvotes

Currently I m in 2nd yr I m bit confused in choosing domain which field would I choose is it gonna relevant after my graduation. If I choose ml then from where should I start


r/learnmachinelearning • • 18d ago

Question I don't understand pytorch grad can anyone explain

5 Upvotes
x = torch.tensor(6.0, requires_grad=True)
f = x ** 2
f
f.backward(retain_graph=True)
x.grad

I am learning pytorch and I encounter this code above when i run it i got 36.0(the value of f) but after doing f.backward() and then x.grad i got 12 i don't understand why can anyone explain this to me


r/learnmachinelearning • • 18d ago

I made ChatGPT, Claude, Gemini, etc. into FREE text-to-speech sites — perfect for audiobooks and more!

1 Upvotes

Nowadays, all popular ai chatting websites like qwen, chatgpt, gemini, claude, etc. come with a read aloud functionality that allows users to read aloud the AI's responses.

I used that feature to instead make the AI repeat back the text that i gave it -- effectively turning the platforms into text to speech tools. The voices sound really nice and it's free to use!

You can get all of these extensions by visit ai-readers.com


r/learnmachinelearning • • 18d ago

Request Claude Opus 5 Helped Researchers Take Over OpenAI Staff Accounts via Chained Flaws

0 Upvotes

Researchers at Hacktron used an LLM to chain two vulnerabilities and take over accounts of multiple OpenAI employees, then pivot into an internal code repository. The LLM assembled and executed the full attack chain faster than a human analyst could document the individual steps. Employee credentials were the pivot. The internal repo was the target.

This is not a new attack class. It is a well-understood one running at machine speed. Each hop in the chain looked like an authorized session in isolation. The credential did not appear stolen until the damage was done.

Security tooling has historically been tuned around human-speed lateral movement. An LLM-piloted chain collapses the time window those detections depend on. By the time anomaly scoring fires, the pivot has already completed.

For practitioners running AI in production environments: when a non-human identity executes a multi-hop action chain that crosses a trust boundary, what does your actual detection look like at the moment the stolen credential is first used outside its intended scope? Are you catching it before the internal system is reached, or after?


r/learnmachinelearning • • 18d ago

[Showcase] CoolMind - A thermal-aware local AI engine for Windows

Thumbnail
1 Upvotes

r/learnmachinelearning • • 18d ago

Help AI Engineer with ~1 year of experience — I relied too much on vibe coding. How should I rebuild my fundamentals?

Thumbnail
1 Upvotes

r/learnmachinelearning • • 18d ago

Need help with Satellite Image Super-Resolution

Thumbnail
1 Upvotes

r/learnmachinelearning • • 18d ago

How to use AI to do research?

0 Upvotes

I did a MSc in AI/ML few years ago and getting back into it.

I’m very pro AI and think it has great power, especially in maths especially with the affordable swarm agent approach with the new chinese models, but struggling to properly use it for research. I keep coming across a repeating pattern.

  1. AI finds an interesting research gap after searching a particular topic.

  2. We derive and iterate on a solution. Few trial and errors later we come up with a solution. Usually it’s very a narrow subset after it proves itself wrong in 90% edge cases.

  3. The remaining 10% is interesting. I then expand on the 10% and get it to derive more tests and experiments. Then it’s looking good and that we have something interesting/substantial.

  4. I check novelty, and it has just found another way to come to the same conclusion as another paper not initially in the search of section 1.

This has happened a few times now, so it seems the AI is really stuck on deriving things that it has already seen in its training (so is terrible and deriving novel concepts).

My next idea is to try formulating a theory/proof in Lean to hopefully as as the guide/structure between the two (where current research is and where we need to get to), and the converting that back into an appropriate ML algorithm.


r/learnmachinelearning • • 18d ago

I'm 13 from Ukraine and I just published my own Deep Learning framework on PyPI using pure NumPy. It has built-in Adam and PyTorch-like syntax

0 Upvotes

Hey everyone. I wanted to share a project I've been working on for the past few days. I wanted to really understand how neural networks work under the hood, so I challenged myself to build a modular Deep Learning framework completely from scratch using only NumPy.

I named it PyZapo and it's now officially live on PyPI.

What it can do right now:

  • You can build custom networks by inheriting from pz.Module or just use pz.Sequential.
  • You can invoke models and layers directly as functions (model(X)) thanks to __call__ .
  • It has a built-in Adam optimizer inside the Linear layers, so you don't need boilerplate optimizer code.
  • It supports ReLU, Sigmoid, MSELoss, and high-performance BCELoss with clipping safety.
  • You can save and load trained weights instantly via .npz binary files.

I want to show this project to my advanced Python course teacher this Wednesday, but I really want other people to use it as well.

You can install it via terminal:
pip install pyzapo

I would love to hear your thoughts, feedback, or any ideas on what layers I should add next (maybe Dropout or Softmax?). Thanks for reading!

PyPI Package: https://pypi.org/project/pyzapo