Andrej Karpathy

The spelled-out intro to neural networks and backpropagation: building micrograd

1913 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: Modern neural-network training reduces to constructing a differentiable computation graph, applying local derivatives backward via the chain rule, and iteratively updating parameters in the loss-reducing direction.
  • Why it matters: It gives a first-principles model of autograd, backpropagation, gradient accumulation, and training-loop failure modes that underlie PyTorch, JAX, and the model-training stack beneath contemporary AI systems.
  • Best use: Use it as a rigorous implementation-oriented refresher before building custom differentiable components, debugging model-training code, or evaluating claims about neural-network training infrastructure; the full lecture is more pedagogical than operational.

Executive Summary

Andrej Karpathy builds micrograd, a minimal scalar automatic-differentiation engine, from a blank notebook to a trainable multilayer perceptron. His central reframing is that neural networks are not mysterious special objects: they are ordinary mathematical expressions over data and learnable weights, followed by a loss function. Backpropagation is a generic algorithm for differentiating such expressions, not an algorithm unique to neural networks.

The implementation rests on a Value object that stores a scalar's forward value, its gradient, its parent nodes, the operation that produced it, and a closure that knows the operation's local backward rule. The engine performs a topological sort of the computation graph and traverses the result in reverse, seeding the final output gradient at 1 and recursively applying the chain rule. Addition routes upstream gradient unchanged; multiplication multiplies it by the other operand; tanh uses 1 - tanh(x)^2.

The lecture then layers a Neuron, Layer, and MLP abstraction on top of that engine. A small 3-input, two-hidden-layer MLP with 41 parameters is trained on four binary-classification examples using squared-error loss and gradient descent. Each cycle is forward pass, reset gradients, backward pass, and parameter update in the negative-gradient direction; minimizing the designed loss makes predictions approach their labels.

Its most useful practical contribution is the debugging discipline embedded in the build. Gradients must accumulate when a node contributes through multiple paths, but they must also be reset before each new backward pass. Karpathy demonstrates both bugs directly: overwriting instead of accumulating gives incorrect derivatives for expressions such as a + a, while forgetting zero_grad silently compounds gradients across training iterations and can appear to work on a trivial problem while failing on a realistic one.

Key Takeaways

  • Claim: Autograd is best understood as differentiation of an arbitrary computation graph; neural-network training is one application of that general machinery. | Evidence: micrograd can build an arbitrary scalar expression from inputs a and b, evaluate forward output G = 24.7, then compute G's derivatives with respect to a and b (138 and 645). Karpathy then maps the same structure to network inputs, weights, predictions, and loss. | Implication: When assessing or designing ML systems, separate the conceptual autograd/training mechanism from implementation concerns such as tensorization, GPU kernels, batching, and distributed efficiency. | Caveat: micrograd is scalar-valued for pedagogy; production frameworks use tensors and parallel hardware, but the underlying differentiation math is unchanged.
  • Claim: Backpropagation is a reverse traversal that composes local derivatives using the chain rule rather than symbolically differentiating an entire network. | Evidence: For L = d × f, the local derivatives are dL/dd = f and dL/df = d. For an upstream path through d = c + e, the gradient to each input is the upstream gradient multiplied by 1; for e = a × b, gradients are multiplied by b and a respectively. | Implication: A custom differentiable operation only needs a correct forward computation and a correct local backward rule; the autograd engine can compose it into arbitrarily large graphs.
  • Claim: The essential autograd-engine architecture is small: each value records graph lineage and a local backward closure, while a reverse topological ordering ensures dependencies are processed in the correct sequence. | Evidence: The Value object holds data, grad, _prev parent references, _op metadata, and _backward. Calling backward first constructs a topological ordering from the output node, seeds output.grad = 1, and invokes _backward on nodes in reverse order. | Implication: For any internal differentiable-programming or optimization engine, prioritize explicit provenance, deterministic graph ordering, and operation-local gradient logic over monolithic symbolic derivatives. | Caveat: This topological-sort approach assumes a directed acyclic computation graph, which is the normal representation for a single forward pass.
  • Claim: Gradient contributions must be accumulated, not overwritten, whenever a value appears on multiple computational paths. | Evidence: The initial implementation incorrectly gives a gradient of 1 for b = a + a because both routes assign rather than add the same gradient. Changing backward rules from assignment to += produces the correct derivative of 2; the same issue appears in branched expressions such as f = (a + b)(ab). | Implication: Gradient accumulation is a correctness invariant, not an optimization detail. Reused parameters, skip connections, shared modules, and multi-loss objectives all require summed gradient contributions.
  • Claim: Neural-network training is an optimization loop in which the loss converts prediction quality into a differentiable scalar objective and parameter gradients prescribe local updates. | Evidence: The lecture trains a 3-input MLP with layer sizes [4, 4, 1] and 41 parameters on four labeled examples. It computes summed squared errors, calls loss.backward(), and updates every parameter as p.data += -learning_rate × p.grad, reducing loss from roughly 7 to near zero and driving outputs toward 1 or -1. | Implication: The operative design question is whether the loss faithfully represents the behavior desired from the system; gradient descent only optimizes the objective supplied. | Caveat: The squared-error loss and full-dataset updates are pedagogical choices; classification systems commonly use cross-entropy and larger datasets use mini-batches.
  • Claim: Training code can seem successful while containing serious gradient-state or step-size bugs, so validation must include explicit gradient checks and controlled optimization behavior. | Evidence: Karpathy numerically checks analytic gradients by perturbing a value by a small h and comparing output change over h. He also exposes a common training-loop error: failing to reset parameter gradients before backward causes gradients to accumulate across iterations, effectively producing unstable and unintended update sizes. | Implication: Treat zero_grad, finite-difference tests for new operators, loss-curve monitoring, and learning-rate stability checks as baseline safeguards rather than optional debugging steps. | Caveat: Numerical finite-difference checks are approximate and become unreliable when h is too small because of floating-point precision.
  • Claim: PyTorch differs from micrograd primarily in scale and efficiency, while exposing the same conceptual interface: tensors retain data and gradients, operations create graphs, and backward propagates derivatives. | Evidence: A one-element PyTorch tensor version of the neuron reproduces micrograd's forward output of approximately 0.707 and gradients of 0.5, 0, -1.5, and 1. PyTorch requires requires_grad=True for leaves whose gradients are needed and uses tensor operations to exploit parallelism. | Implication: Use compact systems such as micrograd to inspect correctness and learn mechanisms; use PyTorch/JAX for real workloads, where tensorized execution and mature kernels dominate. | Caveat: Production-library internals are substantially more complex because they support CPU/GPU kernels, multiple data types, complex values, and performance constraints; Karpathy notes difficulty locating PyTorch's tanh backward implementation among thousands of search hits.

Detailed Brief

Operation design and extensibility

  • Claims: The granularity of a differentiable primitive is a software-design choice rather than a mathematical requirement.; A composite function can be represented as one custom operation or decomposed into more elementary operations, provided its forward and backward behavior is correct.; Python operator ergonomics require handling both Value-and-Value operations and Value-and-scalar operations, including reverse operations such as 2 × a.
  • Evidence: tanh is first implemented directly with its derivative 1 - tanh(x)^2, then re-expressed using exponentiation, addition, subtraction, multiplication, and division; both versions produce identical forward outputs and leaf gradients.; Division is implemented as multiplication by other raised to -1, while x raised to a constant k uses the local derivative k × x^(k-1).; PyTorch's custom torch.autograd.Function pattern similarly requires defining forward and backward methods for a new operation, illustrated through a Legendre-polynomial example.
  • Caveats: The lecture's power implementation is restricted to constant exponents; differentiating a variable exponent requires a different derivative rule.; Equivalent mathematical formulations can differ materially in numerical stability and performance in production systems, a topic not covered here.
  • Implications: Custom model layers, objective functions, or differentiable simulators can be added compositionally, but their backward contracts should be tested independently before inclusion in larger models.; Abstraction boundaries for ops should be selected based on maintainability, numerical behavior, and runtime characteristics—not because backprop requires a particular decomposition.

What changes in realistic training

  • Claims: The core loop remains forward, backward, reset-gradient, and update even as models scale from a 41-parameter MLP to large language models.; Large datasets are normally processed in batches rather than in one full-dataset graph.; Learning-rate scheduling and regularization are practical extensions needed beyond the toy example.
  • Evidence: The micrograd demo notebook uses a larger two-dimensional red/blue classification set, mini-batches, max-margin loss, L2 regularization, and a learning rate that decays over iterations.; Karpathy contrasts the toy squared-error setup with GPT-style next-token prediction, where cross-entropy is the appropriate loss and models have hundreds of billions of parameters.
  • Caveats: The lecture does not cover optimizer variants, normalization, initialization strategy, distributed training, validation methodology, or the mechanics behind generalization.; The claim that large-model training follows the same fundamentals is conceptually correct, but operational complexity at scale is not 'just efficiency'; it includes stability, memory, data, systems, and evaluation challenges.
  • Implications: Use this as a foundation, not as a complete production training playbook.; For architecture or investment diligence, distinguish the invariant core of differentiation from the substantial systems layer required to train and serve modern models reliably.

Notable Concepts & Terms

  • Autograd: Automatic differentiation that records computations and computes derivatives programmatically; micrograd is a minimal autograd engine.
  • Computation graph: The directed graph of values and operations created during the forward pass; it provides the dependency structure used by backpropagation.
  • Chain rule: The calculus rule that multiplies an upstream derivative by an operation's local derivative to pass gradient information backward.
  • Local derivative: An operation-specific sensitivity of its output to each direct input, independent of the larger graph in which the operation sits.
  • Topological sort: An ordering of graph nodes that ensures reverse backpropagation processes each node only after all downstream gradient contributions are available.
  • Gradient accumulation: Summing, via +=, all gradient contributions reaching a reused value or parameter through multiple graph paths.
  • zero_grad: Resetting accumulated parameter gradients before a fresh backward pass; omission creates incorrect cross-step accumulation.
  • Gradient descent: The iterative parameter update p ← p - learning_rate × gradient used to locally minimize a loss.

Operator Notes / Why Ken Should Care

  • If commissioning or reviewing any custom differentiable operation, require a finite-difference gradient check alongside unit tests for forward outputs.
  • Add an explicit gradient-lifecycle check to training-code reviews: gradients should accumulate within one backward graph and be reset exactly once before the next optimization step.
  • When evaluating an ML-training platform, separate its autograd correctness model from the harder operational requirements: tensor compilation, hardware utilization, batching, memory management, numerical stability, and distributed execution.
  • Use the micrograd codebase as a compact reference implementation when diagnosing opaque PyTorch/JAX gradient behavior, rather than trying to infer core semantics from production framework internals.

Source/Metadata

  • Title: The spelled-out intro to neural networks and backpropagation: building micrograd
  • Transcript words: 28189
  • Duration seconds: 8752
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.
Full transcript 23618 words · 124 min read
0:00

Hello, my name is Andrei, and I've been training deep neural networks for a bit more than a decade. And in this lecture, I'd like to show you what neural network training looks like under the hood. So in particular, we are going to start with a blank Jupyter notebook, and by the end of this lecture, we will define and train a neural net, and you'll get to see everything that goes on under the hood and exactly how that works on an intuitive level. Now, specifically, what I would like to do is I would like to take you through building of micrograd. Now micrograd is this library that I released on GitHub about two years ago,

0:31

but at the time, I only uploaded the source code, and you'd have to go in by yourself and really figure out how it works. So in this lecture, I will take you through it step by step and comment on all the pieces of it. So what's micrograd and why is it interesting? Micrograd is an autograd engine. Autograd is short for automatic gradient. And really what it does is it implements backpropagation. Now, backpropagation is this algorithm that allows you to efficiently evaluate the gradient of some kind of a loss function with respect to the weights of a neural network. And what that allows us to do then is we can iteratively tune the weights of that neural network

1:12

to minimize the loss function and therefore improve the accuracy of the network. So backpropagation would be at the mathematical core of any modern deep neural network library, like say PyTorch or JAX. So the functionality of micrograd is, I think, best illustrated by an example. So if we just scroll down here, you'll see that micrograd allows you to build out mathematical expressions. And here what we are doing is we have an expression that we're building out where you have two inputs, A and B. And you'll see that A and B are negative four and two, but we are wrapping those values into this value object that we are going to build out as part of

1:50

micrograd. So this value object will wrap the numbers themselves. And then we are going to build out a mathematical expression here where A and B are transformed into C, D, and eventually E, F, and G. And I'm showing some of the functionality of micrograd and the operations that it supports. So you can add two value objects, you can multiply them, you can raise them to a constant power, you can offset by one, negate, squash at zero, square, divide by constant, divide by hit, etc. And so we're building out an expression graph with these two inputs A and B, and we're creating an output value of G. And micrograd will, in the background, build out this entire mathematical

2:34

expression. So it will, for example, know that C is also a value. C was a result of an addition operation. And the child nodes of C are A and B because it will maintain pointers to A and B value objects. So it will know exactly how all of this is laid out. And then not only can we do what we call the forward pass, where we actually look at the value of G, of course, that's pretty straightforward. We will access that using the dot data attribute. And so the output of the forward pass, the value of G, is 24.7, it turns out. But the big deal is that we can also take this G value object and we can call dot backward. And this will initialize backpropagation at the node G.

3:19

And what backpropagation is going to do is it's going to start at G, and it's going to go backwards through that expression graph, and it's going to recursively apply the chain rule from calculus. And what that allows us to do then is we're going to evaluate the derivative of G with respect to all the internal nodes like E, D, and C, but also with respect to the inputs A and B. And then we can actually query this derivative of G with respect to A. For example, that's A.grad. In this case, it happens to be 138. And the derivative of G with respect to B, which also happens to be here, 645. And this derivative we'll see soon is very important information,

4:00

because it's telling us how A and B are affecting G through this mathematical expression. So in particular, A.grad is 138. So if we slightly nudge A and make it slightly larger, 138 is telling us that G will grow, and the slope of that growth is going to be 138. And the slope of growth of B is going to be 645. So that's going to tell us about how G will respond if A and B get tweaked a tiny amount in a positive direction. Okay? Now, you might be confused about what this expression is that we built out here. And this expression, by the way, is completely meaningless. I just made it up. I'm just flexing about the kinds of operations that are supported by micrograd.

4:44

What we actually really care about are neural networks. But it turns out that neural networks are just mathematical expressions, just like this one, but actually slightly less crazy even. Neural networks are just a mathematical expression. They take the input data as an input, and they take the weights of a neural network as an input. And it's a mathematical expression, and the output are your predictions of your neural net, or the loss function. We'll see this in a bit. But neural networks just happen to be a certain class of mathematical expressions. But backpropagation is actually significantly more general. It doesn't actually care about

5:17

neural networks at all. It only cares about arbitrary mathematical expressions. And then we happen to use that machinery for training of neural networks. Now, one more note I would like to make at this stage is that, as you see here, micrograd is a scalar-valued autograd engine. So it's working on the level of individual scalars, like negative 4 and 2. And we're taking neural nets, and we're breaking them down all the way to these atoms of individual scalars, and all the little pluses and times, and it's just excessive. And so obviously, you would never be doing any of this in production.

5:47

It's really just done for pedagogical reasons, because it allows us to not have to deal with these n-dimensional tensors that you would use in modern deep neural network library. So this is really done so that you understand and factor out backpropagation and chain rule and understanding of neural training. And then if you actually want to train bigger networks, you have to be using these tensors, but none of the math changes. This is done purely for efficiency. We are taking all the scalar values, packaging them up into tensors, which are just arrays of these scalars. And then because we have these large arrays, we're making operations on those

6:23

large arrays that allow us to take advantage of the parallelism in a computer. And all those operations can be done in parallel, and then the whole thing runs faster. But really, none of the math changes, and that's done purely for efficiency. So I don't think that it's pedagogically useful to be dealing with tensors from scratch. And I think that's why I fundamentally wrote micrograd, because you can understand how things work at the fundamental level, and then you can speed it up later. Okay, so here's the fun part. My claim is that micrograd is what you need to train neural networks, and everything else is just

6:53

efficiency. So you'd think that micrograd would be a very complex piece of code. And that turns out to not be the case. So if we just go to micrograd, you will see that there's only two files here in micrograd. This is the actual engine. It doesn't know anything about neural nets. And this is the entire neural nets library on top of micrograd. So engine and nn.py. So the actual backpropagation autograd engine that gives you the power of neural networks is literally 100 lines of code of very simple Python, which we'll understand by the end of this lecture. And then nn.py, this neural network

7:34

library built on top of the autograd engine, is like a joke. It's like, we have to define what is a neuron, and then we have to define what is a layer of neurons. And then we define what is a multilayer perceptron, which is just a sequence of layers of neurons. And so it's just a total joke. So basically, there's a lot of power that comes from only 150 lines of code. And that's all you need to understand, to understand neural network training, and everything else is just efficiency. And of course, there's a lot to efficiency. But fundamentally, that's all that's happening. Okay, so now let's

8:08

dive right in and implement micrograd step by step. The first thing I'd like to do is I'd like to make sure that you have a very good understanding intuitively of what a derivative is, and exactly what information it gives you. So let's start with some basic imports that I copy-paste in every Jupyter notebook always. And let's define the function, scalar-valued function f of x, as follows. So I just made this up randomly, I just wanted to scale a valid function that takes a single scalar x and

8:36

there's a lot of power that comes from only 150 lines of code. And that's all you need to understand, to understand neural network training, and everything else is just efficiency. And of course, there's a lot to efficiency. But fundamentally, that's all that's happening. Okay, so now let's dive right in and implement micrograd step by step. The first thing I'd like to do is make sure that you have a very good understanding intuitively of what a derivative is, and exactly what information it gives you. So let's start with some basic imports that I copy-paste in every Jupyter notebook. And let's define the function, scalar-valued function f of x, as follows. So I just made this up randomly. I just wanted a scalar-valued function that takes a single scalar x and returns a single scalar y. And we can call this function, of course. So we can pass in, say, 3.0 and get 20 back. Now, we can also plot this function to get a sense of its shape. You can tell from the mathematical expression that this is probably a parabola. It's a quadratic. And so if we just create a set of scalar values that we can feed in using, for example, a range from negative 5 to 5 in steps of 0.25. So x is just from negative 5 to 5, not including 5, in steps of 0.25. And we can actually call this function on this NumPy array as well. So we get a set of y's if we call f on x. And these y's are basically also applying the function on every one of these elements independently. And we can plot this using MathPlotlib. So plt.plot, x's and y's, and we get a nice parabola. So previously here, we fed in 3.0 somewhere here, and we received 20 back, which is here the y-coordinate.

8:42

So now I'd like to think through what is the derivative of this function at any single input point x. So what is the derivative at different points x of this function? Now, if you remember back to your calculus class, you've probably derived derivatives. So we take this mathematical expression, 3x squared minus 4x plus 5, and you would write it out on a piece of paper, and you would apply the product rule and all the other rules and derive the mathematical expression of the derivative of the original function. And then you could plug in different x's and see what the derivative is. We're not going to actually do that, because no one in neural networks actually writes out the expression for the neural net. It would be a massive expression. It would be thousands, tens of thousands of terms. No one actually derives the derivative, of course. And so we're not going to take this symbolic approach. Instead, what I'd like to do is look at the definition of derivative and just make sure that we really understand what the derivative is measuring, what it's telling you about the function. And so if we just look up derivative, we see that, okay, this is not a very good definition of derivative. This is a definition of what it means to be differentiable. But if you remember from your calculus, it is the limit as h goes to 0 of f of x plus h minus f of x over h. So what it's saying is if you slightly bump up, you're at some point x that you're interested in, or a, and if you slightly increase it by small number h, how does the function respond? With what sensitivity does it respond? What's the slope at that point? Does the function go up or does it go down, and by how much? And that's the slope of that function, the slope of that response at that point. And so we can basically evaluate the derivative here numerically by taking a very small h. Of course, the definition would ask us to take h to 0. We're just going to pick a very small h, 0.001. And let's say we're interested in point 3.0. So we can look at f of x, of course, is 20. And now f of x plus h. So if we slightly nudge x in a positive direction, how is the function going to respond? And just looking at this, do you expect f of x plus h to be slightly greater than 20, or do you expect it to be slightly lower than 20? And so since 3 is here, and this is 20, if we slightly go positively, the function will respond positively. So you'd expect this to be slightly greater than 20. And now by how much? So f of x plus h minus f of x, this is how much the function responded in a positive direction. And we have to normalize by the run. So we have the rise over run to get the slope. So this, of course, is just a numerical approximation of the slope because we have to make h very, very small to converge to the exact amount. Now, if I'm doing too many zeros, at some point I'm going to get an incorrect answer because we're using floating point arithmetic. And the representations of all these numbers in computer memory are finite. And at some point, we get into trouble. So we can converge towards the right answer with this approach. But basically, at three, the slope is 14. And you can see that by taking 3x squared minus 4x plus 5, and differentiating it in our head. So 3x squared would be 6x minus 4. And then we plug in x equals three. So that's 18 minus 4 is 14. So this is correct.

8:47

So that's at three. Now, how about the slope at, say, negative three? What would you expect for the slope? Now, telling the exact value is really hard. But what is the sign of that slope? So at negative three, if we slightly go in the positive direction at x, the function would actually go down. And so that tells you that the slope would be negative. So we'll get a slight number below 20. And so if we take the slope, we expect something negative, negative 22. Okay. And at some point here, of course, the slope would be zero. Now for this specific function, I looked it up previously, and it's at point two over three. So at roughly two over three, that's somewhere here, this derivative would be zero. So basically, at that precise point, if we nudge in a positive direction, the function doesn't respond. This stays the same almost. And so that's why the slope is zero.

8:53

Okay, now let's look at a bit more complex case. So we're going to start complexifying a bit. So now we have a function here with output variable d that is a function of three scalar inputs, a, b, and c. So a, b, and c are some specific values, three inputs into our expression graph, and a single output, d. And so if we just print d, we get 4. And now what I'd like to do is again look at the derivatives of d with respect to a, b, and c. And think through, again, just the intuition of what this derivative is telling us. So in order to evaluate this derivative, we're going to get a bit hacky here. We're going to again have a very small value of h. And then we're going to fix the inputs at some values that we're interested in. So this is the point a, b, c at which we're going to be evaluating the derivative of d with respect to a, b, and c at that point. So there are the inputs, and now we have d1 is that expression. And then we're going to, for example, look at the derivative of d with respect to a. So we'll take a and we'll bump it by h, and then we'll get d2 to be the exact same function. And now we're going to print f1, d1 is d1, d2 is d2, and print slope. So the derivative or slope here will be, of course, d2 minus d1 divide h. So d2 minus d1 is how much the function increased when we bumped the specific input that we're interested in by a tiny amount. And this is normalized by h to get the slope.

9:01

So, yeah, I just found this. We're going to print d1, which we know is 4. Now d2 will be bumped. a will be bumped by h. So let's just think through a little bit what d2 will be printed out here. In particular, d1 will be 4. Will d2 be a number slightly greater than 4 or slightly lower than 4? And that's going to tell us the sign of the derivative. So we're bumping a by h, So the derivative, or slope, here will be, of course, d2 minus d1 divided by h. So d2 minus d1 is how much the function increased when we bumped the specific input that we're interested in by a tiny amount. And this is normalized by h to get the slope.

9:17

So I just found this. We're going to print d1, which we know is 4. Now d2 will be bumped, a will be bumped by h. So let's just think through a little bit what d2 will be printed out here. In particular, d1 will be 4. Will d2 be a number slightly greater than 4 or slightly lower than 4? And that's going to tell us the sign of the derivative. So we're bumping a by h, b is minus 3, c is 10. So you can intuitively think through this derivative and what it's doing. a will be slightly more positive, but b is a negative number. So if a is slightly more positive, because b is negative 3, we're actually going to be adding less to d. So you'd actually expect that the value of the function will go down. So let's just see this. Yeah. And so we went from 4 to 3.9996. And that tells you that the slope will be negative and will be a negative number because we went down. And then the exact number of slope, the exact amount of slope, is negative 3. And you can also convince yourself that negative 3 is the right answer mathematically and analytically, because if you have a times b plus c and you have calculus, then differentiating a times b plus c with respect to a gives you just b. And indeed, the value of b is negative 3, which is the derivative that we have. So you can tell that that's correct.

9:25

So now if we do this with b, if we bump b by a little bit in a positive direction, we'd get different slopes. So what is the influence of b on the output d? So if we bump b by a tiny amount in a positive direction, then because a is positive, we'll be adding more to d, right? And now what is the sensitivity? What is the slope of that addition? And it might not surprise you that this should be 2. And why is it 2? Because d of d by db, differentiating with respect to b, would give us a. And the value of a is 2. So that's also working well.

9:32

And then if c gets bumped a tiny amount by h, then of course a times b is unaffected. And now c becomes a slightly higher bit. What does that do to the function? It makes it a slightly higher bit because we're simply adding c. And it makes it slightly higher by the exact same amount that we added to c. And so that tells you that the slope is 1. That will be the rate at which d will increase as we scale c. Okay. So we now have some intuitive sense of what this derivative is telling you about the function.

9:38

And we'd like to move to neural networks. Now, as I mentioned, neural networks will be pretty massive expressions, mathematical expressions. So we need some data structures that maintain these expressions. And that's what we're going to start to build out now. So we're going to build out this value object that I showed you in the readme page of micrograd. So let me copy-paste a skeleton of the first very simple value object.

9:44

So class value takes a single scalar value that it wraps and keeps track of. And that's it. So we can, for example, do value of 2.0, and then we can look at its content. And Python will internally use the repr function to return this string. Oops. Like that. So this is a value object with data equals 2 that we're creating here. Now what we'd like to do is we'd like to be able to have not just two values, but we'd like to do a plus b, right? We'd like to add them. So currently you would get an error because Python doesn't know how to add two value objects. So we have to tell it. So here's addition.

9:54

So you have to use these special double underscore methods in Python to define these operators for these objects. So if we use this plus operator, Python will internally call a dot add of b. That's what will happen internally. And so b will be the other and self will be a. And so we see that what we're going to return is a new value object. It's going to be wrapping the plus of their data. But remember, because data is the actual numbered Python number, this operator here is just the typical floating point plus addition. Now it's not an addition of value objects, and we'll return a new value. So now a plus b should work, and it should print value of negative one because that's two plus minus three. There we go.

9:59

Okay. Let's now implement multiply so we can recreate this expression here. So multiply, I think it won't surprise you, will be fairly similar. So instead of add, we're going to be using mul, and then here, of course, we want to do times. And so now we can create a c value object, which will be 10.0. And now we should be able to do a times b. Let's just do a times b first. That's value of negative six now.

10:03

And by the way, I skipped over this a little bit. Suppose that I didn't have the repr function here. Then you'll get some kind of an ugly expression. So what repr is doing is it's providing us a way to print out a nicer-looking expression in Python, so we don't just have something cryptic. We actually have value of negative six. So this gives us a times. And then we should now be able to add c to it because we've defined and told Python how to do a mul and add. And so this will call, this will be equivalent to a dot mul of b, and then this new value object will be dot add of c. And so let's see if that worked. Yep. So that worked well. That gave us four, which is what we expect from before. And I believe we can just call them manually as well. There we go. So yeah.

10:08

Okay. So now what we are missing is the connected tissue of this expression. As I mentioned, we want to keep these expression graphs. So we need to know and keep pointers about what values produce what other values. So here, for example, we are going to introduce a new variable, which we'll call children. And by default, it will be an empty tuple. And then we're actually going to keep a slightly different variable in the class, which we'll call underscore prev, which will be the set of children. This is how I did it in the original micrograd, looking at my code here. I can't remember exactly the reason. I believe it was efficiency. But this underscore children will be a tuple for convenience. But then when we actually maintain it in the class, it will be this set, I believe, for efficiency.

10:14

So now when we are creating a value like this with a constructor, children will be empty and prev will be the empty set. But when we are creating a value through addition or multiplication, we're going to feed in the children of this value, which in this case is self and other. So those are the children here. So now we can do d dot prev. And we'll see that the children of d, we now know, are this value of negative six and value of 10. And this, of course, is the value resulting from a times b and the c value, which is 10.

10:19

Now, the last piece of information we don't know: we know now the children of every single value, but we don't know what operation created this value. So we need one more element here. Let's call it underscore op. And by default, this is the empty set for leaves. And then we'll just maintain it here. And now the operation will be just a simple string. In the case of addition, it's plus. In the case of multiplication, it's times. So now we not just have d dot prev, we also have a d dot op. And we know that d was produced by an addition of those two values. And so now we have the full mathematical expression. And we're building out this data structure, and we know exactly how each value came to be, by what expression and from what other values.

10:24

Now, because these expressions are about to get quite a bit larger, we'd like a way to nicely visualize these expressions that we're building out. So for that, I'm going to copy-paste a bunch of slightly scary code that's going to visualize these expression graphs for us. So here's the code, and I'll explain it in a bit. But first, let me just show you what this code does. Basically, what it does is it creates a new function draw dot that we can call on some root node. And then it's going to visualize it. So if we call draw dot on d,

10:28

not just have d dot prep, we also have a d dot up. And we know that d was produced by an addition of those two values. And so now we have the full mathematical expression. And we're building out this data structure. And we know exactly how each value came to be, by what expression, and from what other values. Now, because these expressions are about to get quite a bit larger, we'd like a way to nicely visualize these expressions that we're building out. So for that, I'm going to copy-paste a bunch of slightly scary code that's going to visualize these expression graphs for us. So here's the code, and I'll explain it in a bit. But first, let me just show you what this code does. What it does is it creates a new function draw dot that we can call on some root node. And then it's going to visualize it. So if we call draw dot on d, which is this final value here, that is a times b plus c, it creates something like this. So this is d, and you see that this is a times b, creating an interpret value, plus c gives us this output node d. So that's draw dot of d. And I'm not going to go through this in complete detail. You can take a look at GraphVis and its API. GraphVis is an open source graph visualization software. And what we're doing here is we're building out this graph in GraphVis API. And you can see that trace is this helper function that enumerates all the nodes and edges in the graph. So that just builds a set of all the nodes and edges. And then we iterate through all the nodes, and we create special node objects for them using dot node. And then we also create edges using dot edge. And the only thing that's slightly tricky here is you'll notice that I add these fake nodes, which are these operation nodes. So, for example, this node here is just a plus node. And I create these special op nodes here, and I connect them accordingly. So these nodes, of course, are not actual nodes in the original graph. They're not actually a value object. The only value objects here are the things in squares. Those are actual value objects, or representations thereof. And these op nodes are just created in this draw dot routine so that it looks nice. Let's also add labels to these graphs so we know what variables are where. So let's create a special underscore label. Or let's just do label equals empty by default and save it in each node. And then here we're going to do label is a, label is b, label is c, a label. And then let's create a special e equals a times b. And e dot label will be e. It's kind of naughty. And e will be e plus c. And a d dot label will be b. Okay, so nothing really changes. I just added this new e function, new e variable. And then here, when we are printing this, I'm going to print the label here. So this will be a percent s bar. And this will be end dot label.

10:32

And so now we have the label on the lot here. So this is a b creating e, and then e plus c creates d, just like we have it here. And finally, let's make this expression just one layer deeper. So d will not be the final output node. Instead, after d, we are going to create a new value object called f. We're going to start running out of variables soon. f will be negative 2.0. And its label will, of course, just be f. And capital L will be the output of our graph. And l will be d times f. Okay, so l will be negative 8 as the output. So now we don't just draw d, we draw l. Okay. And somehow the label of l was undefined. Oops. Although that label has to be explicitly given to it. There we go. So l is the output. So let's quickly recap what we've done so far. We are able to build out mathematical expressions using only plus and times so far. They are scalar-valued along the way. And we can do this forward pass and build out a mathematical expression. So we have multiple inputs here, a, b, c, and f, going into a mathematical expression that produces a single output l. And this here is visualizing the forward pass. So the output of the forward pass is negative 8. That's the value. Now, what we'd like to do next is we'd like to run back propagation. And in back propagation, we are going to start here at the end. And we're going to reverse and calculate the gradient along all these intermediate values. And really what we're computing for every single value here, we're going to compute the derivative of that node with respect to l. So the derivative of l with respect to l is just one. And then we're going to derive what is the derivative of l with respect to f, with respect to d, with respect to c, with respect to e, with respect to b, and with respect to a. And in a neural network setting, you'd be very interested in the derivative of this loss function l with respect to the weights of a neural network. And here, of course, we have just these variables a, b, c, and f. But some of these will eventually represent the weights of a neural net. And so we'll need to know how those weights are impacting the loss function. So we'll be interested in the derivative of the output with respect to some of its leaf nodes. And those leaf nodes will be the weights of the neural net. And the other leaf nodes, of course, will be the data itself. But usually we will not want or use the derivative of the loss function with respect to data, because the data is fixed, but the weights will be iterated on using the gradient information. So next, we are going to create a variable inside the value class that maintains the derivative of l with respect to that value. And we will call this variable grad. So there's a dot data, and there's a self.grad. And initially, it will be zero. And remember that zero basically means no effect. So at initialization, we're assuming that every value does not impact, does not affect the output. Because if the gradient is zero, that means that changing this variable is not changing the loss function. So by default, we assume that the gradient is zero. And then, now that we have grad, and it's 0.0, we are going to be able to visualize it here after data. So here, grad is 0.4F. And this will be in dot grad. And now we are going to be showing both the data and the grad initialized at zero. And we are just about getting ready to calculate the backpropagation. And of course, this grad, again, as I mentioned, is representing the derivative of the output, in this case, l, with respect to this value. So with respect to, this is the derivative of l with respect to f, with respect to d, and so on. So let's now fill in those gradients and actually do backpropagation manually. So let's start filling in these gradients and start all the way at the end, as I mentioned here. First, we are interested to fill in this gradient here. So what is the derivative of l with respect to l? In other words, if I change l by a tiny amount h, how much does l change? It changes by h. So it's proportional, and therefore the derivative will be one. We can, of course, measure these or estimate these numerical gradients numerically, just like we've seen before. So if I take this expression, and I create a def lol function here, and put this here. Now, the reason I'm creating a gating function lol here is because I don't want to pollute or mess up the global scope here. This is just a little staging area. And as you know, in Python, all of these will be local variables to this function. So I'm not changing any of the global scope here. So here, l1 will be l. And then, copy-pasting this expression, we're going to add a small amount h in, for example, a, right? And this would be measuring the derivative of l with respect to a. So here, this will be l2. And then we want to print that derivative. So print l2 minus l1, which is how much l changed, and then normalize it by h. So this is the rise over run. And we have to be careful because l is a value node. So we actually want its data so that these are floats, dividing by h. And this should print the derivative of l with respect to a, because a is the one that we bumped a little bit by h. So what is the derivative of l with respect to a? It's six. Okay. And obviously, if we change l by h, then that would be here effectively. This looks really awkward, but changing l by h, you see the derivative here is one. That's the base case of what we are doing here. So basically, we can come up here, and we can manually set l.grad to one. This is our manual backpropagation. l.grad is one. And let's redraw. And we'll see that we filled in grad

10:42

which is how much l changed, and then normalize it by h. So this is the rise over run. And we have to be careful because l is a value node. So we actually want its data. So these are floats, dividing by h. And this should print the derivative of l with respect to a, because a is the one that we bumped a little bit by h. So what is the derivative of l with respect to a? It's six. Okay. And obviously, if we change l by h, then that would be here effectively. This looks really awkward, but changing l by h, you see the derivative here is one. That's the base case of what we are doing here. So we can come up here, and we can manually set l.grad to one. This

11:22

is our manual backpropagation. l.grad is one. And let's redraw. And we'll see that we filled in grad is one for l. We're now going to continue the backpropagation. So let's look at the derivatives of l with respect to d and f. Let's do d first. So what we are interested in, if I create a markdown on here, is we'd like to know, we have that l is d times f. And we'd like to know what is d l by d d. What is that? And if you know your calculus, l is d times f. So what is d l by d d? It would be f. And if you don't believe me, we can also just derive it because the proof would be fairly

11:59

straightforward. We go to the definition of the derivative, which is f of x plus h minus f of x divide h as a limit. Limit of h goes to zero of this kind of expression. So when we have l is d times f, then increasing d by h would give us the output of d plus h times f. That's f of x plus h, right? Minus d times f, and then divide h. And symbolically, expanding out here, we would have d times f plus h times f minus d times f divide h. And then you see how the df minus df cancels. So you're left with h times f, divide h, which is f. So in the limit as h goes to zero of the derivative definition, we just get f in the case of d times f. So symmetrically, d l by d

12:42

f will just be d. So what we have is that f.grad, we see now, is just the value of d, which is four. And we see that d.grad is just the value of f. And so the value of f is negative two. So we'll set those manually. Let me erase this markdown node, and then let's redraw what we have. Okay? And let's just make sure that these were correct. So we seem to think that d l by d d is negative two. So let's double-check. Let me erase this plus h from before. And now we want the derivative with respect to f. So let's just come here when I create f and let's do a plus h here. And this should print a derivative of l with

13:30

respect to f. So we expect to see four. Yeah, and this is four up to floating point funkiness. And then d l by d d should be f, which is negative two. Grad is negative two. So if we again come here and we change d, d.data plus equals h right here. So we expect, we've added a little h and then we see how l changed, and we expect to print negative two. There we go. So we've numerically verified. What we're doing here is kind of like an inline gradient check. Gradient check is when we are deriving this backpropagation and getting the derivative we expect to all the intermediate results. And then numerical gradient is just

14:16

estimating it using small step size. Now we're getting to the crux of backpropagation. So this will be the most important node to understand, because if you understand the gradient for this node, you understand all of backpropagation and all of training of neural nets. So we need to derive d l by d c. In other words, the derivative of l with respect to c, because we've computed all these other gradients already. Now we're coming here and we're continuing the backpropagation manually. So we want d l by d c, and then we'll also derive d l by d e. Now here's the problem. How do we derive d l

14:58

by d c? We actually know the derivative of l with respect to d. So we know how l is sensitive to d. But how is l sensitive to c? So if we wiggle c, how does that impact l through d? So we know d l by d c. And we also know here how c impacts d. And so just very intuitively, if you know the impact that c is having on d and the impact that d is having on l, then you should be able to somehow put that information together to figure out how c impacts l. And indeed, this is what we can actually do. So in particular, we know, just concentrating on d first, let's look at what is the derivative of d with respect to c. So in other words, what is d d by d c?

15:55

So here we know that d is c times c plus e. That's what we know. And now we're interested in d d by d c.

16:03

If you know your calculus again, and you remember that differentiating c plus e with respect to c, you know that that gives you 1.0. And we can also go back to the basics and derive this. Because again, we can go to our f of x plus h minus f of x divided by h. That's the definition of a derivative as h goes to 0. And so here, focusing on c and its effect on d, we can do the f of x plus h, which will be c incremented by h plus e. That's the first evaluation of our function minus c plus e, and then divide h. And so what is this? Just expanding this out, this will be c plus h plus e minus c minus e,

17:00

divide h. And then you see here how c minus c cancels, e minus e cancels. We're left with h over h, which is 1.0. And so by symmetry also, d by d e will be 1.0 as well. So the derivative of a sum expression is very simple. And this is the local derivative. So I call this the local derivative because we have the final output value all the way at the end of this graph. And we're now at a small node here. And this is a little plus node. And the little plus node doesn't know anything about the rest of the graph that it's embedded in. All it knows is that it's a plus. It took a c and an e, added them, and created a d. And this plus node also knows

17:52

the local influence of c on d, or rather the derivative of d with respect to c. And it also knows the derivative of d with respect to e. But that's not what we want. That's just a local derivative. What we actually want is d l by d c. And l is here just one step away. But in the general case, this little plus node could be embedded in a massive graph. So again, we know how l impacts d, and now we know how c and e impact d. How do we put that information together to write d l by d c? And the answer, of course, is the chain rule in calculus. And so I pulled up a chain rule here from

18:35

Wikipedia. And I'm going to go through this very briefly. So chain rule... Wikipedia sometimes can be very confusing, and calculus can be very confusing. This is the way I learned chain rule, and it was very confusing. What is happening? It's just complicated. So I like this expression much better. If a variable z depends on a variable y, which itself depends on the variable x, then z depends on x as well, obviously, through the intermediate variable y. And in this case, the chain rule is expressed as if you want d z by d x, then you take the d z by d y and you multiply it by d y by d x. So the chain rule fundamentally is telling you how we chain these

19:14

derivatives together correctly. So to differentiate through a function composition, we have to apply a multiplication of those derivatives. So that's really what chain rule is telling us. And there's a nice little intuitive explanation here, which I also think is cute. The chain rule states that knowing the instantaneous rate of change of z with respect to y and y relative to x allows one to calculate the instantaneous rate of change of z relative to x as a product of those two rates of change, simply the product of those two. So here's a good one. If a car travels twice

20:00

as fast as a bicycle, and the bicycle is four times as fast as walking men, then the car travels two times four, eight times as fast as a man. And so this makes it very clear that the correct thing to do is to multiply. So a car is twice as fast as a bicycle, and a bicycle is four times as fast as a man. So the car will be eight times as fast as the man. And so we can take these intermediate rates of change, if you will, And there's a nice little intuitive explanation here, which I also think is cute. The chain rule states that knowing the instantaneous rate of change of z with respect to y and y

20:40

relative to x allows one to calculate the instantaneous rate of change of z relative to x as a product of those two rates of change, simply the product of those two. So here's a good one. If a car travels twice as fast as a bicycle, and the bicycle is four times as fast as a walking man, then the car travels two times four, eight times as fast as a man. And so this makes it very clear that the correct thing to do is to multiply. So a car is twice as fast as a bicycle, and a bicycle is four times as fast as a man. So the car will be eight times as fast as the man. And so we can take these intermediate rates of change, if you will,

21:22

and multiply them together. And that justifies the chain rule intuitively. So have a look at chain rule. But here, really what it means for us is there's a very simple recipe for deriving what we want, which is dL by dc. And what we have so far is we know what we want, and we know what is the impact of d on l. So we know dL by dd, the derivative of l with respect to dd. We know that that's negative two. And now because of this local reasoning that we've done here, we know dd by dc. So how does c impact d? And in particular, this is a plus node. So the local derivative is simply 1.0. It's very simple. And so the chain rule tells us that dL by dc,

22:01

going through this intermediate variable, will just be simply dL by dd times dd by dc. That's chain rule. So this is identical to what's happening here, except z is rl, y is rd, and x is rc. So we literally just have to multiply these. And because these local derivatives like dd by dc are just 1, we just copy over dL by dd because this is just times 1. So because dL by dd is negative 2, what is dL by dc? Well, it's the local gradient 1.0 times dL by dd, which is negative 2. So literally what a plus node does, you can look at it that way, is it literally just routes the gradient

22:41

because the plus node's local derivatives are just 1. And so in the chain rule, 1 times dL by dd is just dL by dd. And so that derivative just gets routed to both c and to e in this case. So we have that e.grad, or let's start with c since that's the one we looked at, is negative 2 times 1, negative 2. And in the same way, by symmetry, e.grad will be negative 2. That's the claim. So we can set those. We can redraw. And you see how we just assign negative 2, negative 2. So this backpropagating signal, which is carrying the information of what is the derivative of L with respect to all the

23:14

intermediate nodes, we can imagine it almost like flowing backwards through the graph. And a plus node will simply distribute the derivative to all the leaf nodes, sorry, to all the children nodes of it. So this is the claim. And now let's verify it. So let me remove the plus h here from before. And now instead, what we're going to do is we want to increment c. So c.data will be incremented by h. And when I run this, we expect to see negative 2. Negative 2. And then of course for e. So e.data plus equals h. And we expect to see negative 2. Simple. So those are the derivatives of these internal nodes. And now we're going to recurse our way backwards

24:00

again. And we're again going to apply the chain rule. So here we go. Our second application of chain rule. And we will apply it all the way through the graph. We just happen to only have one more node remaining. We have that dl by dE, as we have just calculated, is negative 2. So we know that. So we know the derivative of L with respect to e. And now we want dl by dA, right? And the chain rule is telling us that that's just dl by dE, negative 2, times the local gradient. So what is the local gradient? Basically dE by dA. We have to look at that. So I'm a little times node inside a massive graph.

24:40

And I only know that I did a times b and I produced an e. So now what is dE by dA and dE by dB? That's the only thing that I know about. That's my local gradient. So because we have that e is a times b, we're asking what is dE by dA? And of course, we just did that here. We had a times, so I'm not going to re-derive it. But if you want to differentiate this with respect to a, you'll just get b, right? The value of b, which in this case is negative 3.0. So we have that dl by dA. Well, let me just do it right here. We have that a.grad, and we are applying chain rule here, is dl by dE, which we see here is negative 2, times

25:17

what is dE by dA. It's the value of b, which is negative 3. That's it. And then we have b.grad is again dl by dE, which is negative 2, in the same way, times what is dE by dB? dE by dB is the value of a, which is 2.0. That's the value of a. So these are our claimed derivatives. Let's redraw. And we see here that a.grad turns out to be 6, because that is negative 2 times negative 3. And b.grad is negative 2 times 2, which is negative 4. So those are our claims. Let's delete this and let's verify them. We have a here, a.data plus equals h. So the claim is that a.grad is 6. Let's verify. 6. And we have b.data plus equals h.

26:08

So nudging b by h and looking at what happens, we claim it's negative 4. And indeed, it's negative 4, plus minus, again, float oddness. And that's it. That was the manual back propagation all the way from here to all the leaf nodes. And we've done it piece by piece. And really, all we've done is, as you saw, we iterated through all the nodes one by one and locally applied the chain rule. We always know what is the derivative of l with respect to this little output. And then we look at how this output was produced. This output was produced through some operation. And we have the pointers to the children

26:51

nodes of this operation. And so in this little operation, we know what the local derivatives are, and we just multiply them onto the derivative always. So we just go through and recursively multiply on the local derivatives. And that's what back propagation is, just a recursive application of chain rule backwards through the computation graph. Let's see this power in action very briefly. What we're going to do is we're going to nudge our inputs to try to make l go up. So in particular, what we're doing is we want a.data, we're going to change it. And if we want l to go up,

27:45

that means we just have to go in the direction of the gradient. So a should increase in the direction of the gradient by some small step amount. This is the step size. And we don't just want this for a, but also for b, also for c, also for f. Those are leaf nodes, which we usually have control over. And if we nudge in the direction of the gradient, we expect a positive influence on l. So we expect l to go up positively. So it should become less negative. It should go up to say negative six or something like that. It's hard to tell exactly. And we'd have to rerun the forward pass. So let me just do that here.

28:36

This would be the forward pass, f would be unchanged. This is effectively the forward pass. And now if we print l.data, we expect, because we nudged all the values, all the inputs, in the direction of the gradient, we expect a less negative l. We expect it to go up. So maybe it's negative six or so. Let's see what happens. Okay, negative seven. And this is basically one step of an optimization that we'll end up running. And really this gradient just gives us some power because we know how to influence the final outcome. And this will be extremely useful for training and all that as well, as you can see.

29:21

So now I would like to do one more example of manual backpropagation using a bit more complex and useful example. We are going to backpropagate through a neuron. So we want to eventually build out neural networks. And in the simplest case, these are multilayer perceptrons, as they're called. So this is a two-layer neural net. And it's got these hidden layers made up of neurons. And these neurons

29:46

we expected less negative l. We expect it to go up. So maybe it's negative six or so. Let's see what happens. Okay, negative seven. And this is one step of an optimization that we'll end up running. And really this gradient just gives us some power because we know how to influence the final outcome. And this will be extremely useful for training and all that as well, as you can see.

29:55

So now I would like to do one more example of manual backpropagation using a bit more complex and useful example. We are going to backpropagate through a neuron. So we want to eventually build out neural networks. And in the simplest case, these are multilayer perceptrons, as they're called. So this is a two-layer neural net. And it's got these hidden layers made up of neurons. And these neurons are fully connected to each other. Now, biologically, neurons are very complicated devices, but we have very simple mathematical models of them. And so this is a very simple mathematical model of a neuron. You have some inputs, x's, and then you have these synapses that have weights on them. So the w's are weights. And then the synapse interacts with the input to this neuron multiplicatively. So what flows to this cell body of this neuron is w times x. But there's multiple inputs. So there's many w times x's flowing into the cell body. The cell body then also has some bias. So this is the innate trigger-happiness of this neuron. So this bias can make it a bit more trigger-happy or a bit less trigger-happy, regardless of the input. But we're taking all the w times x of all the inputs, adding the bias, and then we take it through an activation function. And this activation function is usually some kind of squashing function, like a sigmoid or tanh or something like that.

30:01

So as an example, we're going to use the tanh in this example. NumPy has an np.tanh. So we can call it on a range, and we can plot it. This is the tanh function. And you see that the inputs, as they come in, get squashed on the y coordinate here. So right at zero, we're going to get exactly zero. And then as you go more positive in the input, then you'll see that the function will only go up to one and then plateau out. And so if you pass in very positive inputs, we're going to cap it smoothly at one. And on the negative side, we're going to cap it smoothly to negative one. So that's tanh. And that's the squashing function, or an activation function. And what comes out of this neuron is just the activation function applied to the dot product of the weights and the inputs. So let's write one out.

30:07

I'm going to copy-paste because I don't want to type too much. But okay, so here we have the inputs x1, x2. So this is a two-dimensional neuron. So two inputs are going to come in. These are thought of as the weights of this neuron, weights w1, w2. And these weights, again, are the synaptic strengths for each input. And this is the bias of the neuron, b. And now what we want to do is, according to this model, we need to multiply x1 times w2. And then we need to add bias on top of it. And it gets a little messy here. But all we are trying to do is x1 w1 plus x2 w2 plus b. And these are multiplied here, except I'm doing it in small steps so that we actually have pointers to all these intermediate nodes. So we have x1 w1 variable, x times x2 w2 variable, and I'm also labeling them. So n is now the cell body raw activation without the activation function for now. And this should be enough to plot it. So draw dot of n gives us x1 times w1, x2 times w2 being added. Then the bias gets added on top of this. And this n is this sum.

30:13

So we're now going to take it through an activation function. And let's say we use the tanh so that we produce the output. So what we'd like to do here is we'd like to do the output, and I'll call it o, is n dot tanh. Okay, but we haven't yet written the tanh. Now, the reason that we need to implement another tanh function here is that tanh is a hyperbolic function, and we've only so far implemented a plus and a times. And you can't make a tanh out of just pluses and times. You also need exponentiation. So tanh is this kind of formula here. You can use either one of these. And you see that there's exponentiation involved, which we have not implemented yet for our little value node here. So we're not going to be able to produce tanh yet, and we have to go back up and implement something like it.

30:19

Now, one option here is we could actually implement exponentiation, right? And we could return the exp of a value instead of a tanh of a value. Because if we had exp, then we have everything else that we need, because we know how to add and we know how to multiply. So we'd be able to create tanh if we knew how to exp. But for the purposes of this example, I specifically wanted to show you that we don't necessarily need to have the most atomic pieces in this value object. We can actually create functions at arbitrary points of abstraction. They can be complicated functions, but they can also be very simple functions like a plus. And it's totally up to us. The only thing that matters is that we know how to differentiate through any one function. So we take some inputs and we make an output. The only thing that matters, it can be an arbitrarily complex function, as long as you know how to create the local derivative. If you know the local derivative of how the inputs impact the output, then that's all you need.

30:23

So we're going to cluster up all of this expression, and we're not going to break it down to its atomic pieces. We're just going to directly implement tanh. So let's do that.

30:28

So let's do that. And then out will be a value of, and we need this expression here. So let me actually copy-paste. Let's grab n, which is a self.data. And then this, I believe, is the tanh. Math.exp of 2, no, n minus 1 over 2n plus 1. Maybe I can call this x, just so that it matches exactly. Okay. And now this will be t and children of this node. There's just one child, and I'm wrapping it in a tuple. So this is a tuple of one object, just self. And here, the name of this operation will be tanh. And we're going to return that.

30:32

Okay. So now values should be implementing tanh. And now we can scroll all the way down here. And we can actually do n.dot.tanh. And that's going to return the tanh output of n. And now we should be able to draw a dot of o, not of n. So let's see how that worked.

30:37

There we go. And it went through tanh to produce this output. So now tanh is our little micrograd-supported node here as an operation. And as long as we know the derivative of tanh, then we'll be able to backpropagate through it. Now let's see this tanh in action. Currently, it's not squashing too much because the input to it is pretty low. So if the bias was increased to, say, 8, then we'll see that what's flowing into the tanh now is 2. And tanh is squashing it to 0.96. So we're already hitting the tail of this tanh. And it will smoothly go up to 1 and then plateau out over there.

30:43

Okay. So now I'm going to do something slightly strange. I'm going to change this bias from 8 to this number, 6.88, et cetera. And I'm going to do this for specific reasons because we're about to start backpropagation. And I want to make sure that our numbers come out nice. They're not very crazy numbers. They're nice numbers that we can understand in our head. Let me also add O's label. O is short for output here. So that's the r. Okay. So 0.88 flows into tanh, which comes out 0.7. So now we're going to do backpropagation and we're going to fill in all the gradients. So what is the derivative O with respect to all the inputs here? And of course, in a typical neural network setting, what we really care about the most is the derivative of these neurons on the weights specifically, the w2 and w1, because those are the weights that we're going to be changing as part of the optimization. And the other thing that we have to remember is here, we have only a single neuron, but in the neural net, you typically have many neurons and they're

30:47

not very crazy numbers. They're nice numbers that we can understand in our head. Let me also add O's label. O is short for output here. So that's the R. Okay. So 0.88 flows into 10h, which comes out 0.7. So now we're going to do backpropagation and we're going to fill in all the gradients. So what is the derivative O with respect to all the inputs here? And of course, in a typical neural network setting, what we really care about the most is the derivative of these neurons on the weights specifically, the w2 and w1, because those are the weights that we're going to be changing as part of the optimization. And the other thing that we have to remember is here, we have only a single neuron, but in the neural net, you typically have many neurons and they're connected. So this is only one small neuron, a piece of a much bigger puzzle. And eventually, there's a loss function that measures the accuracy of the neural net. And we're backpropagating with respect to that accuracy and trying to increase it. So let's start off backpropagation here in the end. What is the derivative of O with respect to O? The base case, we know always, is that the gradient is just 1.0. So let me fill it in. And then let me split out the drawing function here. And then here, cell, clear this output here. Okay. So now when we draw O, we'll see that O, that grad, is 1. So now we're going to backpropagate through the 10h. So to backpropagate through 10h, we need to know the local derivative of 10h. So if we have that O is 10h of n, then what is dO by dN? Now, what you could do is you could come here and you could take this expression and you could do your calculus derivative taking. And that would work. But we can also just scroll down on Wikipedia here into a section that hopefully tells us that derivative d by dx of 10h of x is any of these. I like this one, 1 minus 10h squared of x. So this is 1 minus 10h of x squared. So what this is saying is that dO by dN is 1 minus 10h of n squared. And we already have 10h of n. It's just O. So it's 1 minus O squared. So O is the output here. So the output is this number. O dot data is this number. And then what this is saying is that dO by dN is 1 minus this squared. So 1 minus O dot data squared is 0.5 conveniently. So the local derivative of this 10h operation here is 0.5. And so that would be dO by dN. So we can fill in that n.grad is 0.5. We'll just fill it in.

30:53

So this is exactly 0.5, 1.5. So now we're going to continue the backpropagation. This is 0.5 and this is a plus node. So how is backprop going to, what is backprop going to do here? And if you remember our previous example, a plus is just a distributor of gradient. So this gradient will simply flow to both of these equally. And that's because the local derivative of this operation is one for every one of its nodes. So 1 times 0.5 is 0.5. So therefore, we know that this node here, which we called this, its grad, is just 0.5. And we know that b.grad is also 0.5.

30:59

So let's set those and let's draw. So those are 0.5. Continuing, we have another plus. 0.5, again, we'll just distribute. So 0.5 will flow to both of these. So we can set theirs. X2w2 as well. That grad is 0.5. And let's redraw. Pluses are my favorite operations to backpropagate through because it's very simple. So now what's flowing into these expressions is 0.5. And so really, again, keep in mind what the derivative is telling us at every point in time along here. This is saying that if we want the output of this neuron to increase, then the influence on these expressions is positive on the output. Both of them are positive contribution to the output.

31:07

So now backpropagating to x2 and w2 first. This is a times node. So we know that the local derivative is the other term. So if we want to calculate x2.grad, then can you think through what it's going to be? So x2.grad will be w2.data times this x2w2.grad, right? And w2.grad will be x2.data times x2w2.grad, right? So that's the little local piece of chain rule.

31:12

Let's set them and let's redraw. So here we see that the gradient on our weight 2 is 0 because x2's data was 0, right? But x2 will have the gradient 0.5 because data here was 1. And so what's interesting here is because the input x2 was 0, then because of the way the times works, this gradient will be 0. And think about intuitively why that is. Derivative always tells us the influence of this on the final output. If I wiggle w2, how is the output changing? It's not changing because we're multiplying by 0. So because it's not changing, there is no derivative. And 0 is the correct answer because we're squashing that 0. And let's do it here. 0.5 should come here and flow through this times. And so we'll have that x1.grad is... Can you think through a little bit what this should be?

31:17

The local derivative of times with respect to x1 is going to be w1. So w1's data times x1w1.grad, and w1.grad will be x1.data times x1w1.grad. Let's see what those came out to be. So this is 0.5. So this would be negative 1.5. And this would be 1. And we've backpropagated through this expression. These are the actual final derivatives. So if we want this neuron's output to increase, we know that what's necessary is that w2, we have no gradient. w2 doesn't actually matter to this neuron right now. But this neuron, this weight should go up. So if this weight goes up, then this neuron's output would have gone up, and proportionally, because the gradient is 1.

31:24

Okay, so doing the backpropagation manually is obviously ridiculous. So we are now going to put an end to this suffering. And we're going to see how we can implement the backward pass a bit more automatically. We're not going to be doing all of it manually out here. It's now pretty obvious to us, by example, how these pluses and times are backpropagating gradients. So let's go up to the value object. And we're going to start codifying what we've seen in the examples below.

31:31

So we're going to do this by storing a special self.backward and underscore backward. And this will be a function which is going to do that little piece of chain rule. At each little node that took inputs and produced output, we're going to store how we are going to chain the output's gradient into the input's gradients. So by default, this will be a function that doesn't do anything. And you can also see that here in the value in micrograd. So we have this backward function, by default, doesn't do anything. This is an empty function. And that would be the case, for example, for a leaf node. For a leaf node, there's nothing to do.

31:40

But now when we're creating these out values, these out values are an addition of self and other. And so we will want to set out's backward to be the function that propagates the gradient. So let's define what should happen. And we're going to store it in a closure. Let's define what should happen when we call out's grad. For addition, our job is to take out's grad and propagate it into self's grad and other.grad. So we want to set self.grad to something and we want to set other's grad to something. Okay. And the way we saw below how chain rule works, we want to take the local derivative times the global derivative, I should call it, which is the derivative of the final output of the expression with respect to out's data. So the local derivative of self in an addition is 1.0. So it's just 1.0 times out's grad. That's the chain rule. And other's grad will be 1.0 times out grad. And what you're seeing here is that out's grad will simply be copied onto self's grad and other's grad, as we saw happens for an addition operation. So we're going to later call this function to propagate the gradient, having done an addition.

31:51

Let's now do multiplication. We're going to also define.backward. And we're going to set its backward to be backward. and we want to set others.grad to something. Okay. And the way we saw below how chain rule works, we want to take the local derivative times the global derivative, I should call it, which is the derivative of the final output of the expression with respect to outs data. So the local derivative of self in an addition is 1.0. So it's just 1.0 times outs grad. That's the chain rule.

32:07

And others.grad will be 1.0 times out grad. And what you're seeing here is that outs grad will simply be copied onto selfs grad and others.grad, as we saw happens for an addition operation. So we're going to later call this function to propagate the gradient, having done an addition. Let's now do multiplication. We're going to also define.backward. And we're going to set its backward to be backward. And we want to chain out grad into self.grad and others.grad. And this will be a little piece of chain rule for multiplication. So we'll have. So what should this be? Can you think through? So what is the local derivative here? The local derivative was others.data

32:55

and then times out.grad. That's chain rule. And here we have self.data times out.grad. That's what we've been doing. And finally here for 10h, def backward. And then we want to set outs backwards to be just backward. And here we need to backpropagate. We have out.grad And we want to chain it into self.grad. And self.grad will be the local derivative of this operation that we've done here, which is 10h. And so we saw that the local gradient is 1 minus the 10h of x squared, which here is t. That's the local derivative because t is the output of this 10h.

34:00

So 1 minus t squared is the local derivative. And then gradient has to be multiplied because of the chain rule. So out.grad is chained through the local gradient into salt.out.grad. And that should be it. So we're going to redefine our value node. We're going to swing all the way down here. And we're going to redefine our expression. Make sure that all the grads are zero. Okay. But now we don't have to do this manually anymore. We are going to be calling the dot backward in the right order.

35:12

So first we want to call os.backward. So o was the outcome of 10h. Right? So calling os.backward will be this function. This is what it will do. Now we have to be careful because there's a times out.grad. And out.grad, remember, is initialized to zero. So here we see grad zero.

36:24

So as a base case, we need to set os.grad to 1.0 to initialize this with 1. And then once this is 1, we can call o.backward. And what that should do is it should propagate this grad through 10h. So the local derivative times the global derivative, which is initialized at 1. So this should... don't. So I thought about redoing it, but I figured I should just leave the error in here because it's pretty funny. Why is nunti object not callable?

37:38

It's because I screwed up. We're trying to save these functions. So this is correct. This here, we don't want to call the function because that returns none. These functions return none. We just want to store the function. So let me redefine the value object. And then we're going to come back in. Redefine the expression. Draw dot. Everything is great.

38:49

O.grad is 1. O.grad is 1. And now... Now this should work, of course. Okay. So all that backward should have...

39:31

This grad should now be 0.5 if we redraw. And if everything went correctly. 0.5. Yay! Okay. So now we need to call ns.grad. ns.backward, sorry. ns.backward. So that seems to have worked.

40:42

ns.backward routed the gradient to both of these. So this is looking great. Now we can, of course, call b.grad. b.backward, sorry. What's going to happen? Well, b doesn't have a backward. b's backward, because b is a leaf node, b's backward is by initialization the empty function. So nothing would happen. But we can call it on it. But when we call this one, its backward. Then we expect this 0.5 to get further routed.

41:51

Right? So there we go, 0.5, 0.5. And then finally, we want to call it here on x2w2 and on x1w1. Let's do both of those. And there we go. So we get 0, 0.5, negative 1.5, and 1 exactly as we did before. But now we've done it through calling that backward manually. So we have one last piece to get rid of, which is us calling underscore backward manually. So let's think through what we are actually doing. We've laid out a mathematical expression, and now we're trying to go backwards through that expression.

43:03

So going backwards through the expression just means that we never want to call a dot backward for any node before we've done everything after it. So we have to do everything after it before we're ever going to call that backward on any one node. We have to get all of its full dependencies. Everything that it depends on has to propagate to it before we can continue backpropagation. So this ordering of graphs can be achieved using something called topological sort. So topological sort is a laying out of a graph such that all the edges go only from left to right. So here we have a graph. It's a directory as a click graph, a DAG.

44:13

And this is two different topological orders of it, I believe, where you'll see that it's a laying out of the nodes such that all the edges go only one way, from left to right. And implementing topological sort, you can look in Wikipedia and so on. I'm not going to go through it in detail. But this is what builds a topological graph. We maintain a set of visited nodes. And then we are going through starting at some root node, which for us is O.

45:23

That's where we want to start the topological sort. And starting at O, we go through all of its children and we need to lay them out from left to right. And this starts at O.

45:47

If it's not visited, then it marks it as visited. And then it iterates through all of its children and calls build topological on them. And then after it's gone through all the children, it adds itself. So this node that we're going to call it on, like say O, is only going to add itself to the topo list after all of the children have been processed. And that's how this function is guaranteeing that you're only going to be in the list once all your children are in the list. And that's the invariant that is being maintained. So if we build topo on O and then inspect this list, we're going to see that it ordered our value

46:50

objects. And the last one is the value of 0.707, which is the output. So this is O and then this is N and then all the other nodes get laid out before it.

47:07

So that builds the topological graph. And really what we're doing now is we're just calling dot underscore backward on all of the nodes in a topological order. So if we just reset the gradients, they're all zero. What did we do? We started by setting O.grad to be one. That's the base case. Then we built a topological order. And then we went for node in reversed optopo.

48:17

Now in the reverse order, because this list goes from, you know, we need to go through it in reversed order. So starting at O, node.backward. And that should be it. There we go. Those are the correct derivatives. Finally, we are going to hide this functionality. So I'm going to copy this and we're going to hide it inside the value class because we don't want to

49:30

have all that code lying around. So instead of an underscore backward, we're now going to define an actual backward. So that's backward without the underscore. And that's going to do all the stuff that we just derived. So let me just clean this up a little bit. So we're first going to build a topological graph starting at self. So build topo of self will populate the topological order into the topo list, which is a local variable. Then we set self.grad to be one. And then for each node in the reversed list, so starting at us and going to all the children,

50:41

underscore backward. And that should be it. So save. Come down here. Redefine. Okay, all the grads are zero. And now what we can do is O dot backward without the underscore. And there we go. And that's backpropagation. Place for one neuron. We shouldn't be too happy with ourselves, actually, because we have a bad bug.

51:51

And we have not surfaced the bug because of some specific conditions that we have to think about right now. So here's the simplest case that shows the bug. Say I create a single node A. And then I create a B that is A plus A.

52:16

And then I call backward. So build topo of self will populate the topological order into the topo list, which is a local variable. Then we set self.grad to be one. And then for each node in the reversed list, so starting at us and going to all the children, underscore backward. And that should be it. So save. Come down here. Redefine. Okay, all the grads are zero. And now what we can do is O dot backward without the underscore. And there we go.

53:29

And that's backpropagation. Place for one neuron. We shouldn't be too happy with ourselves, actually, because we have a bad bug. And we have not surfaced the bug because of some specific conditions that we have to think about right now. So here's the simplest case that shows the bug. Say I create a single node A. And then I create a B that is A plus A. And then I call backward. So what's going to happen is A is 3. And then B is A plus A.

54:40

So there's two arrows on top of each other here. Then we can see that B is, of course, the forward pass works. B is just A plus A, which is 6. But the gradient here is not actually correct that we calculated automatically. And that's because, of course, just doing calculus in your head, the derivative of B with respect to A should be 2. 1 plus 1.

55:27

It's not 1. Intuitively, what's happening here? So B is the result of A plus A. And then we call backward on it.

55:57

So let's go up and see what that does. B is the result of addition. So out is B. And then when we call backward, what happened is self.grad was set to 1. And then other.grad was set to 1. But because we're doing A plus A, self and other are actually the exact same object. So we are overriding the gradient. We are setting it to 1. And then we are setting it again to 1.

57:09

And that's why it stays at 1. So that's a problem. There's another way to see this in a little bit more complicated expression. So here we have A and B. And then D will be the multiplication of the two. And E will be the addition of the two. And then we multiply E times D to get F. And then we call it F.backward. And these gradients, if you check, will be incorrect. So fundamentally, what's happening here, again, is we're going to see an issue any time we use a variable more than once.

58:20

Until now, in these expressions above, every variable is used exactly once. So we didn't see the issue. But here, if a variable is used more than once, what's going to happen during backward pass? We're backpropagating from F to E to D. So far, so good. But now E calls its backward. And it deposits its gradients to A and B. But then we come back to D and call backward.

59:08

And it overwrites those gradients at A and B. So that's obviously a problem. And the solution here, if you look at the multivariate case of the chain rule and its generalization there, the solution there is that we have to accumulate these gradients. These gradients add.

59:57

And so instead of setting those gradients, we can simply do plus equals. We need to accumulate those gradients. Plus equals.

1:00:18

Plus equals. Plus equals. Plus equals. And this will be okay, remember, because we are initializing them at zero. So they start at zero. And then any contribution that flows backwards will simply add. So now if we redefine this one, because of the plus equals, this now works. Because A dot grad started at zero. And when we call B dot backward, we deposit one. And then we deposit one again.

1:01:25

And now this is two, which is correct. And here this will also work. And we'll get correct gradients. Because when we call E dot backward, we will deposit the gradients from this branch. And then we get to D dot backward. It will deposit its own gradients. And then those gradients simply add on top of each other. And so we just accumulate those gradients. And that fixes the issue. Okay, now before we move on, let me actually do a bit of cleanup here. And delete some of this intermediate work. So I'm not going to need any of this now that we've derived all of it.

1:02:36

We are going to keep this. Because I want to come back to it. Delete the tanh. Delete our mode of game example. Delete the step. Delete this. Delete this.

1:03:39

Keep the code that draws. And then delete this example. And leave behind only the definition of value. And now let's come back to this nonlinearity here that we implemented, the tanh. Now I told you that we could have broken down tanh into its explicit atoms in terms of other expressions if we had the exp function.

1:04:49

So if you remember, tanh is defined like this. And we chose to develop tanh as a single function. And we can do that because we know its derivative and we can back propagate through it. But we can also break down tanh and express it as a function of exp. And I would like to do that now because I want to prove to you that you get all the same results and all the same gradients. But also because it forces us to implement a few more expressions. It forces us to do exponentiation, addition, subtraction, division, and things like that.

1:05:44

And I think it's a good exercise to go through a few more of these. Okay, so let's scroll up to the definition of value. And here, one thing that we currently can't do is we can do a value of, say, 2.0. But we can't do, for example, here, we want to add a constant 1. And we can't do something like this. And we can't do it because it says int object has no attribute data. That's because A plus 1 comes right here to add. And then other is the integer 1.

1:06:55

And then here, Python is trying to access 1.data, and that's not a thing.

1:07:06

And that's because 1 is not a value object. And we only have addition for value objects. So as a matter of convenience, so that we can create expressions like this and make them make sense, we can simply do something like this. Basically, we let other alone if other is an instance of value. But if it's not an instance of value, we're going to assume that it's a number, like an integer or float. And we're going to simply wrap it in value. And then other will just become value of other. And then other will have a data attribute. And this should work.

1:08:16

So if I just save this, redefine value, then this should work. There we go. Okay, and now let's do the exact same thing for multiply. Because we can't do something like this, again, for the exact same reason. So we just have to go to mul. And if other is not a value, then let's wrap it in value. Let's redefine value. And now this works. Now here's a kind of unfortunate and not obvious part. A times 2 works. We saw that.

1:09:29

But 2 times A, is that going to work? You'd expect it to, right? But actually it will not. And the reason it won't is because Python doesn't know. When you do A times 2, Python will go and it will basically do something like A.mul of 2. That's basically what it will call. But to it, 2 times A is the same as 2.mul of A. And it doesn't, 2 can't multiply value.

1:10:34

And so it's really confused about that. So instead, what happens is, in Python, the way this works is you are free to define something called the rmul. And rmul is kind of like a fallback. So if Python can't do 2 times A, it will check if, by any chance, A knows how to multiply 2. And that will be called into rmul.

1:11:23

So because Python can't do 2 times A, it will check, is there an rmul in value? And because there is, it will now call that. And what we'll do here is we will swap the order of the operands. So basically, 2 times A will redirect to rmul. And rmul will basically call A times 2.

1:12:02

And that's how that will work. So redefining that with rmul, 2 times A becomes 4. Okay, now looking at the other elements that we still need, we need to know how to exponentiate and how to divide.

1:12:28

So let's first do the exponentiation part. We're going to introduce a single function exp here. And exp is going to mirror tanh in the sense that it's a single function that transforms a single scalar value and outputs a single scalar value. So we pop out the Python number. We use math.exp to exponentiate it, create a new value object, everything that we've seen before. The tricky part, of course, is how do you backpropagate through e to the x? And so here, you can potentially pause the video and think about what should go here. So 2 times a will redirect to rmol. And rmol will call a times 2. And that's how that will work.

1:13:34

So redefining that with rmol, 2 times a becomes 4. Okay, now looking at the other elements that we still need, we need to know how to exponentiate and how to divide. So let's first do the exponentiation part. We're going to introduce a single function exp here. And exp is going to mirror 10h in the sense that it's a single function that transforms a single scalar value and outputs a single scalar value. So we pop out the Python number. We use math.exp to exponentiate it, create a new value object, everything that we've seen before. The tricky part, of course, is how do you backpropagate through e to the x?

1:14:18

And so, you can potentially pause the video and think about what should go here. Okay, so we need to know what is the local derivative of e to the x. So d by dx of e to the x is, famously, just e to the x. And we've already just calculated e to the x.

1:14:38

And it's inside out.data. So we can do out.data times out.grad. That's the chain rule.

1:14:53

So we're just chaining on to the current running grad. And this is what the expression looks like. It looks a little confusing, but this is what it is. And that's the exponentiation.

1:15:14

So redefining, we should now be able to call a.exp. And hopefully, the backward pass works as well. Okay, and the last thing we'd like to do, of course, is if we'd like to be able to divide. Now, I actually will implement something slightly more powerful than division, because division is just a special case of something a bit more powerful. So in particular, just by rearranging, if we have some kind of a b equals value of 4.0 here, we'd like to be able to do a divide b, and we'd like this to be able to give us 0.5. Now, division actually can be reshuffled as follows. If we have a divide b, that's actually the same as a multiplying 1 over b,

1:15:41

and that's the same as a multiplying b to the power of negative 1. And so what I'd like to do instead is I'd like to implement the operation of x to the k for some constant k. So it's an integer or a float. And we would like to be able to differentiate this. And then as a special case, negative 1 will be division. And so I'm doing that just because it's more general, and you might as well do it that way. So what I'm saying is we can redefine division, which we will put here somewhere. Yeah, we can put it here somewhere. What I'm saying is that we can redefine division. So self-divide other. This can actually be rewritten as self times other to the power of negative 1.

1:16:04

And now, value raised to the power of negative 1, we have to now define that. So we need to implement the power function. Where am I going to put the power function? Maybe here somewhere. This is this color for it. So this function will be called when we try to raise a value to some power, and other will be that power. Now, I'd like to make sure that other is only an int or a float. Usually, other is some kind of a different value object. But here, other will be forced to be an int or a float. Otherwise, the math won't work for what we're trying to achieve in this specific case. That would be a different derivative expression if we wanted other to be a value.

1:16:41

So here, we create the other value, which is just this data raised to the power of other. And other here could be, for example, negative 1. That's what we are hoping to achieve. And then, this is the backward stub. And this is the fun part, which is what is the chain rule expression here for backpropagating through the power function, where the power is to the power of some kind of a constant?

1:17:09

So this is the exercise.

1:17:15

And maybe pause the video here and see if you can figure it out yourself as to what we should put here. Okay, so you can actually go here and look at derivative rules as an example. And we see lots of derivative rules that you can hopefully know from calculus. In particular, what we're looking for is the power rule. Because that's telling us that if we're trying to take d by dx of x to the n, which is what we're doing here, then that is just n times x to the n minus 1. Right? Okay. So that's telling us about the local derivative of this power operation. So all we want here, basically, n is now other and self.data is x.

1:18:05

And so this now becomes other, which is n, times self.data, which is now a Python int or a float. It's not a value object. We're accessing the data attribute raised to the power of other minus 1 or n minus 1. I can put brackets around this, but this doesn't matter because power takes precedence over multiply in pi helen. So that would have been okay. And that's the local derivative only. But now we have to chain it. And we chain it just simply by multiplying by our top grad. That's chain rule. And this should technically work. And we're going to find out soon. But now if we do this, this should now work. And we get 0.5.

1:19:12

So the forward pass works, but does the backward pass work? And I realized that we actually also have to know how to subtract. So right now, a minus b will not work. To make it work, we need one more piece of code here. And this is the subtraction. And the way we're going to implement subtraction is we're going to implement it by addition of a negation. And then to implement negation, we're going to multiply by negative one. So again, using the stuff we've already built and just expressing it in terms of what we have. And a minus b is not working. Okay, so now let's scroll again to this expression here for this neuron.

1:20:01

And let's just compute the backward pass here once we've defined o. And let's draw it. So here's the gradients for all of these leaf nodes for this two-dimensional neuron that has a 10h that we've seen before. So now what I'd like to do is I'd like to break up this 10h into this expression here. So let me copy-paste this here. And now, we'll preserve the label and we will change how we define o. So in particular, we're going to implement this formula here. So we need e to the 2x minus 1 over e to the x plus 1. So e to the 2x, we need to take 2 times n and we need to exponentiate it. That's e to the 2x.

1:20:46

And then because we're using it twice, let's create an intermediate variable e and then define o as e minus 1 over e plus 1. E minus 1 over e plus 1. And that should be it. And then we should be able to draw dot of o. So now before I run this, what do we expect to see? Number one, we're expecting to see a much longer graph here because we've broken up 10h into a bunch of other operations. But those operations are mathematically equivalent. And so what we're expecting to see is, number one, the same result here. So the forward pass works.

1:21:18

And number two, because of that mathematical equivalence, we expect to see the same backward pass and the same gradients on these leaf nodes. So these gradients should be identical. So let's run this. So number one, let's verify that instead of a single 10h node, we now have exp and we have plus, we have times negative 1. This is the division. And we end up with the same forward pass here. And then the gradients, we have to be careful because they're in slightly different order, potentially. The gradients for w2 x2 should be 0 and 0.5. w2 and x2 are 0 and 0.5. And w1 x1 are 1 and negative 1.5. 1 and negative 1.5.

1:22:07

So that means that both our forward passes and backward passes were correct because this turned out to be equivalent to 10h before. And so the reason I wanted to go through this exercise is, number one, we got to practice a few more operations and writing more backward passes. And number two, I wanted to illustrate the point that the level at which you implement your operations is totally up to you. You can implement backward passes for tiny expressions like a single individual plus or a single times.

1:22:20

Or you can implement them for, say, 10h, which is potentially a composite operation because it's made up of all these more atomic operations. But really all of this is a fake concept. All that matters is we have some kind of inputs and some kind of an output. And this output is a function of the inputs in some way. 1 and negative 1.5. So that means that both our forward passes and backward passes were correct because this turned out to be equivalent to tanh before. And so the reason I wanted to go through this exercise is, number one, we got to practice a few more operations and writing more backward passes.

1:22:47

And number two, I wanted to illustrate the point that the level at which you implement your operations is totally up to you. You can implement backward passes for tiny expressions like a single individual plus or a single times. Or you can implement them for, say, tanh, which is potentially a composite operation because it's made up of all these more atomic operations. But really all of this is a fake concept. All that matters is we have some inputs and an output. And this output is a function of the inputs in some way.

1:23:06

And as long as you can do the forward pass and the backward pass of that little operation, it doesn't matter what that operation is and how composite it is. If you can write the local gradients, you can chain the gradient and continue backpropagation. So the design of what those functions are is completely up to you. So now I would like to show you how you can do the exact same thing, but using a modern deep neural network library, for example, PyTorch, which I've roughly modeled micrograd by. And so PyTorch is something you would use in production. And I'll show you how you can do the exact same thing, but in the PyTorch API.

1:23:30

So I'm just going to copy-paste it in and walk you through it a little bit. This is what it looks like. So we're going to import PyTorch, and then we need to define these value objects like we have here. Now, micrograd is a scalar-valued engine.

1:23:43

So we only have scalar values like 2.0. But in PyTorch, everything is based around tensors. And like I mentioned, tensors are just n-dimensional arrays of scalars. So that's why things get a little bit more complicated here. I just need a scalar-valued tensor, a tensor with just a single element. But by default, when you work with PyTorch, you would use more complicated tensors like this. So if I import PyTorch, then I can create tensors like this. And this tensor, for example, is a 2 by 3 array of scalars in a single compact representation. So we can check its shape. We see that it's a 2 by 3 array, and so on.

1:24:14

So this is usually what you would work with in the actual libraries.

1:24:21

So here I'm creating a tensor that has only a single element, 2.0. And then I'm casting it to be double because Python is by default using double precision for its floating-point numbers. So I'd like everything to be identical. By default, the data type of these tensors will be float32. So it's only using a single-precision float. So I'm casting it to double so that we have float64, just like in Python. So I'm casting it to double. And then we get something similar to a value of 2. The next thing I have to do is because these are leaf nodes, by default, PyTorch assumes that they do not require gradients.

1:24:53

So I need to explicitly say that all of these nodes require gradients. Okay. So this is going to construct scalar-valued, one-element tensors. Make sure that PyTorch knows that they require gradients. Now, by default, these are set to false, by the way, because of efficiency reasons, because usually you would not want gradients for leaf nodes, like the inputs to the network. And this is just trying to be efficient in the most common cases. So once we've defined all of our values in PyTorch land, we can perform arithmetic just like we can here in micrograd land. So this would just work. And then there's a torch.tanh also. And then we get back a tensor again.

1:25:23

And we can, just like in micrograd, it's got a data attribute, and it's got grad attributes. So these tensor objects, just like in micrograd, have a .data and a .grad. And the only difference here is that we need to call .item, because otherwise PyTorch .item basically takes a single tensor of one element, and it just returns that element, stripping out the tensor. So let me just run this, and hopefully we are going to get this is going to print the forward pass, which is 0.707. And this will be the gradients, which hopefully are 0.50, negative 1.5, and 1. So if we just run this, there we go. 0.7, so the forward pass agrees. And then 0.50, negative 1.5, and 1.

1:25:46

So PyTorch agrees with us. And just to show you here, O: here's a tensor with a single element, and it's a double, and we can call .item on it to just get the single number out. So that's what .item does. And O is a tensor object, like I mentioned, and it's got a backward function, just like we've implemented. And then all of these also have a .grad. So like x2, for example, has a grad, and it's a tensor. And we can pop out the individual number with .item. So Torch can do what we did in micrograd as a special case when your tensors are all single-element tensors.

1:26:14

But the big deal with PyTorch is that everything is significantly more efficient because we are working with these tensor objects, and we can do lots of operations in parallel on all of these tensors. But otherwise, what we've built very much agrees with the API of PyTorch. Okay, so now that we have some machinery to build out pretty complicated mathematical expressions, we can also start building out neural nets. And as I mentioned, neural nets are just a specific class of mathematical expressions. So we're going to start building out a neural net piece by piece, and eventually we'll build out a two-layer multilayer perceptron, as it's called.

1:26:26

And I'll show you exactly what that means. Let's start with a single individual neuron. We've implemented one here, but here I'm going to implement one that also subscribes to the PyTorch API in how it designs its neural network modules. So just like we saw that we can match the API of PyTorch on the autograd side, we're going to try to do that on the neural network modules. So here's class Neuron. And just for the sake of efficiency, I'm going to copy-paste some sections that are relatively straightforward. So the constructor will take the number of inputs to this neuron, which is how many inputs come to a neuron. So this one, for example, has three inputs.

1:26:47

And then it's going to create a weight that is some random number between negative one and one for every one of those inputs, and a bias that controls the overall trigger happiness of this neuron. And then we're going to implement a def __call__ of self and x, some input x. And really what we want to do here is w times x plus b, where w times x here is a dot product specifically. Now, if you haven't seen call, let me just return 0.0 here for now. The way this works now is we can have an x, which is, say, 2.0, 3.0. Then we can initialize a neuron that is two-dimensional because these are two numbers. And then we can feed those two numbers into that neuron to get an output.

1:26:59

And so when you use this notation, n of x, Python will use call. So currently call just returns 0.0. Now, we'd like to actually do the forward pass of this neuron instead. So what we're going to do here first is we need to basically multiply all of the elements of w with all of the elements of x pairwise. We need to multiply them. So the first thing we're going to do is we're going to zip up self.w and x. And in Python, zip takes two iterators and creates a new iterator that iterates over the tuples of their corresponding entries. So for example, just to show you, we can print this list and still return 0.0 here. Sorry.

1:27:30

So we see that these w's are paired up with the x's: w with x. And now what we want to do is, for wi, xi in, we want to multiply wi times xi. And then we want to sum all of that together to come up with an activation and add also self.b on top. So that's the raw activation. And then, of course, we need to pass that through a nonlinearity. So what we're going to be returning is act.tanh. And here's out. So now we see that we are getting some outputs. And we get a different output from a neuron each time because we are initializing different weights and biases.

1:28:08

And then, to be a bit more efficient here, actually, sum, by the way, takes a second optional parameter, which is the start. And by default, the start is 0. So these elements of this sum will be added on top of 0 to begin with. But actually, we can just start with self.b. w with x. And now what we want to do is, for w_i x_i in, we want to multiply w times... w_i times x_i, and then we want to sum all of that together to come up with an activation and add also b on top. So that's the raw activation. And then, of course, we need to pass that through a nonlinearity. So what we're going to be returning is act.tanh. And here's out.

1:29:01

So now we see that we are getting some outputs. And we get a different output from a neuron each time because we are initializing different weights and biases. And then to be a bit more efficient here, actually, sum, by the way, takes a second optional parameter, which is the start. And by default, the start is 0. So these elements of this sum will be added on top of 0 to begin with. But actually, we can just start with b. And then we just have an expression like this. And then the generator expression here must be parenthesized in Python. There we go. Yep. So now we can forward a single neuron. Next up, we're going to define a layer of neurons.

1:29:51

So here we have a schematic for an MLP. So we see that these MLPs, each layer, this is one layer, has actually a number of neurons. And they're not connected to each other, but all of them are fully connected to the input. So what is a layer of neurons? It's just a set of neurons evaluated independently. So in the interest of time, I'm going to do something fairly straightforward here. It's literally a layer. It's just a list of neurons. And then how many neurons do we have? We take that as an input argument here. How many neurons do you want in your layer? Number of outputs in this layer. And so we just initialize completely independent neurons

1:30:52

with this given dimensionality. And when we call on it, we just independently evaluate them. So now instead of a neuron, we can make a layer of neurons.

1:31:13

They are two-dimensional neurons, and let's have three of them. And now we see that we have three independent evaluations of three different neurons. Right? Okay. And finally, let's complete this picture and define an entire multi-layer perceptron, or MLP. And as we can see here, in an MLP, these layers just feed into each other sequentially. So let's come here, and I'm just going to copy the code here in the interest of time. So an MLP is very similar. We're taking the number of inputs as before, but now instead of taking a single nout, which is number of neurons in a single layer, we're going to take a list of nouts, and this list defines the sizes of all the layers

1:32:19

that we want in our MLP. So here we just put them all together, and then iterate over consecutive pairs of these sizes, and create layer objects for them. And then in the call function, we are just calling them sequentially. So that's an MLP, really. And let's actually re-implement this picture. So we want three input neurons, and then two layers of four, and an output unit. So we want a three-dimensional input. Say this is an example input. We want three inputs into two layers of four, and one output. And this, of course, is an MLP. And there we go. That's a forward pass of an MLP. To make this a little bit nicer, you see how we have just a single element,

1:33:22

but it's wrapped in a list, because layer always returns lists. So for convenience, return outs[0] if len(outs) is exactly a single element, else return outs. And this will allow us to just get a single value out at the last layer that only has a single neuron. And finally, we should be able to draw dot of n of x. And as you might imagine, these expressions are now getting relatively involved. So this is an entire MLP that we're defining now, all the way until a single output.

1:34:09

Okay? And so obviously, you would never differentiate on pen and paper, these expressions.

1:34:26

But with micrograd, we will be able to backpropagate all the way through this, and backpropagate into these weights of all these neurons. So let's see how that works. Okay, so let's create ourselves a very simple example data set here. So this data set has four examples. And so we have four possible inputs into the neural net. And we have four desired targets. So we'd like the neural net to assign or output 1.0 when it's fed this example, negative one when it's fed these examples, and one when it's fed this example. So it's a very simple binary classifier neural net that we would like here. Now let's think what the neural net currently thinks about these four examples.

1:35:31

We can just get their predictions. We can just call n of x for x in xs. And then we can print. So these are the outputs of the neural net on those four examples. So the first one is 0.91, but we'd like it to be 1. So we should push this one higher. This one, we want to be higher. This one says 0.88, and we want this to be negative one. This is 0.88, we want it to be negative one. And this one is 0.88, we want it to be one. So how do we make the neural net, and how do we tune the weights, to better predict the desired targets? And the trick used in deep learning to achieve this is to calculate a single number that somehow measures

1:36:44

the total performance of your neural net. And we call the single number the loss. So the loss first is a single number that we're going to define that basically measures how well the neural net is performing. Right now, we have the intuitive sense that it's not performing very well because we're not very close to this. So the loss will be high, and we'll want to minimize the loss. So in particular, in this case, what we're going to do is we're going to implement the mean squared error loss.

1:37:57

So what this is doing is we're going to iterate for y_ground_truth and y_output in zip of ys and ypred. So we're going to pair up the ground truths with the predictions, and the zip iterates over tuples of them. And for each y_ground_truth and y_output, we're going to subtract them

1:39:11

and square them. So let's first see what these losses are. These are individual loss components. And so for each one of the four, we are taking the prediction and the ground truth. We are subtracting them and squaring them. So because this one is so close to its target, 0.91 is almost one. Subtracting them gives a very small number.

1:40:22

So here we would get a negative 0.1, and then squaring it just makes sure that regardless of whether we are more negative or more positive, we always get a positive number. Instead of squaring, we could also take, for example, the absolute value. We need to discard the sign. And so you see

1:41:37

that the expression is arranged so that you only get zero exactly when y_out is equal to y_ground_truth. When those two are equal, so your prediction is exactly the target, you are going to get zero. And if your prediction is not the target, you are going to get some other number. So here, for example, we are way off. And so that's why

1:42:50

the loss is quite high. And the more off we are, the greater the loss will be. So we don't want high loss, we want low loss. And so the final loss here will be just the sum of all of these numbers. So you see that this should be zero roughly plus zero roughly, but plus seven. So loss should be about seven here. And now we want to minimize the loss.

1:44:01

We want the loss to be low because if loss is low, then every one of the predictions is equal to its target. So the loss, the lowest it can be, is zero, and the greater it is, the worse off the neural net is predicting. So now, of course, if we do loss.backward, something magical happened when I hit enter. And the magical thing, of course, that happened is that we can look at n.layers.neuron n.layers

1:45:17

at, say, the first layer dot neurons at zero, because remember that MLP has the layers, which is a list, and each layer has neurons, which is a list, and that gives us an individual neuron, and then it's got some weights, and so we can, for example, look at the weights at zero. Oops, it's not called weights, it's called w. And that's a value, but now this value also has a grad because of the backward pass. And so we see that because this gradient here on this particular

1:46:32

weight of this particular neuron of this particular layer is negative, we see that its influence on the loss is also negative. So slightly increasing this particular weight of this neuron of this layer would make the loss go down. And we actually have this information for every single one of our neurons and all their parameters. Actually, it's worth looking at also a draw dot of loss, by the way. So previously we looked at the draw dot of a single neuron forward pass, and that was already a large expression.

1:47:45

But what is this expression? We actually forwarded every one of those four examples, and then we have the loss on top of them with the mean squared error. And so this is a really massive graph because this graph that we've built up now, oh my gosh, this graph that we've built up now, which is kind of excessive at zero. Oops, it's not called weights. It's called w, and that's a value. But now this value also has a grad because of the backward pass. And so we see that because this gradient here on this particular weight of this particular neuron of this particular layer is negative, we see that its influence on the loss is also negative. So slightly increasing

1:48:58

this particular weight of this neuron of this layer would make the loss go down. And we actually have this information for every single one of our neurons and all their parameters. Actually, it's worth looking at also a draw dot of loss, by the way. So previously, we looked at the draw dot of a single neuron forward pass, and that was already a large expression. But what is this expression? We actually forwarded every one of those four examples, and then we have the loss on top of them with the mean squared error. And so this is a really massive graph because this graph that we've built up now, oh my gosh, this graph that we've built up now, which is excessive,

1:50:08

it's excessive because it has four forward passes of a neural net for every one of the examples, and then it has the loss on top, and it ends with the value of the loss, which was 7.12. And this loss will now backpropagate through all the forward forward passes all the way through every single intermediate value of the neural net, all the way back to, of course, the parameters of the weights, which are the input. So these weight parameters here are inputs to this neural net,

1:51:22

and these numbers here, these scalars, are inputs to the neural net. So if we went around here, we will probably find some of these examples, this 1.0, potentially maybe this 1.0, or some of the others. And you'll see that they all have gradients as well. The thing is, these gradients on the input data are not that useful to us, and that's because the input data is assumed to be not changeable. It's a given to the problem, and so it's a fixed input. We're not going to be changing it or messing with it, even though we do have gradients for it. But some of these gradients here will be for the neural network parameters, the w's and the b's, and those, of course,

1:52:36

we want to change. Okay, so now we're going to want some convenience code to gather up all of the parameters of the neural net so that we can operate on all of them simultaneously. And every one of them we will nudge a tiny amount based on the gradient information. So let's collect the parameters of the neural net all in one array. So let's create a parameters of self, self, that just returns self.w, which is a list concatenated with a list of self.b. So this will just return a list. List plus list gives you a list. So that's parameters of neuron, and I'm calling it this way because also PyTorch has parameters on every single NN module, and it does exactly

1:53:34

what we're doing here. It just returns the parameter tensors. For us, it's the parameter scalars. Now layer is also a module, so it will have parameters, self. And what we want to do here is something like this: params is here, and then for neuron in self.neurons, we want to get neuron.parameters, and we want to params.extend, right? So these are the parameters of this neuron, and then we want to put them on top of params, so params.extend of piece, and then we want to return params. So this, there's way too much code. So actually, there's a way to simplify this, which is return p for neuron in self.neurons for p in neuron.parameters. So it's a single list comprehension.

1:54:36

In Python, you can nest them like this, and you can then create the desired array. So these are identical. We can take this out, and then let's do the same here: def.parameters self, and return a parameter for layer in self.layers for p in layer.parameters. And that should be good. Now let me pop out this so we don't re-initialize our network, because we need to re-initialize our... Okay, so unfortunately, we will have to probably re-initialize the network because we just had functionality because this class, of course. I want to get all the n dot parameters, but that's not going to work because this is the old class. Okay, so unfortunately, we do have to re-initialize

1:55:34

the network, which will change some of the numbers. But let me do that so that we pick up the new API. We can now do n-dot parameters, and these are all the weights and biases inside the entire neural net. So in total, this MLP has 41 parameters, and now we'll be able to change them. If we recalculate the loss here, we see that unfortunately we have slightly different predictions and slightly different loss, loss, but that's okay. Okay, so we see that this neuron's gradient is slightly negative. We can also look at its data right now, which is 0.85. So this is the current value of this neuron, and this is its gradient on the loss. So what we want to do now

1:56:27

is we want to iterate for every P in n.dot.parameters. So for all the 41 parameters in this neural net, we actually want to change P.data slightly according to the gradient information. Okay, so dot dot dot to do here, but this will be basically a tiny update in this gradient descent scheme. In gradient descent, we are thinking of the gradient, gradient, as a vector pointing in the direction of increased loss. And so in gradient descent, we are modifying P.data by a small step size in the direction of the gradient. So the step size, as an example, could be a very small number like 0.01 is the step size times P.grad, right? But we have to think through

1:57:14

some of the signs here. So, in particular, working with this specific example here, we see that if we just left it like this, then this neuron's value would be currently increased by a tiny amount of the gradient. The gradient is negative, so this value of this neuron would go slightly down. It would become like 0.8, 4, or something like that. But if this neuron's value goes lower, that would actually increase the loss. That's because the derivative of this neuron is negative, so increasing this makes the loss go down. So increasing it is what we want to do instead of decreasing it. So basically, what we're missing here is we're actually missing a negative sign.

1:57:56

And again, this other interpretation, and that's because we want to minimize the loss. We don't want to maximize the loss. We want to decrease it. And the other interpretation, as I mentioned, is you can think of the gradient vector, so basically the vector of all the gradients, as pointing in the direction of increasing the loss. But then we want to decrease it, so we actually want to go in the opposite direction. And so you can convince yourself that this does the right thing here with the negative because we want to minimize the loss. So if we nudge all the parameters by a tiny amount, then we'll see that this data will have changed a little bit. So now this neuron

1:58:37

is a tiny amount greater value. So 0.854 went to 0.857, and that's a good thing because slightly increasing this neuron data makes the loss go down according to the gradient. And so the correcting has happened sign-wise. And so now what we would expect, of course, is that because we've changed all these parameters, we expect that the loss should have gone down a bit. So we want to re-evaluate the loss. Let me, this is just a data definition that hasn't changed, but the forward pass here of the network we can recalculate. And actually, let me do it outside here so that we can compare the two loss values. So here, if I recalculate the loss, we'd expect the new loss

1:59:21

now to be slightly lower than this number. So hopefully what we're getting now is a tiny bit lower than 4.84. 4.36. Okay, and remember, the way we've arranged this is that low loss means that our predictions are matching the targets. So our predictions now are probably slightly closer to the targets. And now all we have to do is we have to iterate this process. So again, we've done the forward pass, and this is the loss. Now we can loss.backward. Let me take these out, and we can do a step size, and now we should have a slightly lower loss. 4.36 goes to 3.9.

2:00:04

And okay, so we've done the forward pass, here's the backward pass, nudge, and now the loss is 3.66, 3.47, and you get the idea. We just continue doing this, and this is gradient descent.

2:00:20

We're just iteratively doing forward pass, backward pass, update, forward pass, backward pass, update, and the neural net is improving its predictions. So here, if we look at ypred now, ypred, we see that this value should be getting closer to 1, so this value should be getting more positive. These should be getting more negative, and this one should also be getting more positive. So if we just iterate this a few more times, actually we may be able to afford to go a bit faster. Let's try a slightly higher learning rate. Oops. Okay, there we go. So now we're at 0.31. If you go too fast, by the way, if you try to make it too big of a step, you may actually overstep.

2:01:16

It's overconfidence because, again, remember, we don't actually know exactly about the loss function. The loss function has all kinds of structure, and we only know about the very local dependence of all these parameters on the loss. But if we step too far, we may step into a part of the loss that is completely different, and that can destabilize training and make your loss actually blow up even. So the loss is now 0.04, so actually the predictions should be really quite close. Let's take a look. So you see how this is almost 1, almost negative 1, almost 1. We can continue going. So, yep, backward, update. Oops, there we go. So we went way too fast, and

2:02:08

we actually overstepped. So we got too, too eager. Where are we now? Oops. Okay. a bit faster. Let's try a slightly higher learning rate. Oops. Okay. There we go. Now we're at 0.31. If you go too fast, by the way, if you try to make it too big of a step, you may actually overstep. It's overconfidence because, again, remember, we don't actually know exactly about the loss function. The loss function has all kinds of structure, and we only know about the very local dependence of all these parameters on the loss. But if we step too far, we may step into a part of the loss that is completely different, and that can destabilize training and make your loss actually blow up

2:02:58

even. The loss is now 0.04, so actually the predictions should be really quite close. Let's take a look. You see how this is almost 1, almost negative 1, almost 1? We can continue going. Yep. Backward. Update. Oops. There we go. We went way too fast, and we actually overstepped, so we got too, too eager. Where are we now? Oops. Okay. 7 in negative 9. This is very, very low loss, and the predictions are basically perfect. Somehow, we were doing way too big updates, and we briefly exploded, but then somehow we ended up getting into a really good spot. Usually, this learning rate and the tuning of it is a subtle art. You want to set your learning rate. If it's too low,

2:03:52

you're going to take way too long to converge, but if it's too high, the whole thing gets unstable, and you might actually even explode the loss, depending on your loss function. Finding the step size to be just right is a pretty subtle art sometimes when you're using vanilla gradient descent, but we happen to get into a good spot. We can look at end dot parameters. This is the setting of weights and biases that makes our network predict the desired targets very, very close, and we've successfully trained a neural net. Okay, let's make this a tiny bit more respectable and implement an actual training loop and what that looks like. This is the data definition that stays.

2:04:45

This is the forward pass.

2:04:49

For k in range, we're going to take a bunch of steps. First, you do the forward pass. We evaluate the loss. Let's reinitialize the neural net from scratch, and here's the data. We first do forward pass, then we do the backward pass, and then we do an update. That's gradient descent, and then we should be able to iterate this, and we should be able to print the current step, the current loss. Let's just print the number of the loss, and that should be it. The learning rate, 0.01, is a little too small. 0.1, we saw, is a little bit dangerously too high. Let's go somewhere in between, and we'll optimize this for not 10 steps, but let's go for 20 steps. Let me erase

2:05:31

all of this junk, and let's run the optimization. You see how we've actually converged slower, in a more controlled manner, and got to a loss that is very low. I expect white bread to be quite good. There we go.

2:05:50

That's it. Okay, this is embarrassing, but we actually have a really terrible bug in here, and it's a subtle bug, and it's a very common bug, and I can't believe I've done it for the 20th time in my life, especially on camera. I could have reshot the whole thing, but I think it's pretty funny, and you get to appreciate a bit what working with neural nets is like sometimes. We are guilty of a common bug. I've actually tweeted the most common neural net mistakes a long time ago now, and I'm not really going to explain any of these, except for we are guilty of number three: you forgot to zero grad before dot backward. What is that? Basically, what's happening,

2:06:22

and it's a subtle bug, and I'm not sure if you saw it, is that all of these weights here have a dot data and a dot grad, and dot grad starts at zero. Then we do backward, and we fill in the gradients, and then we do an update on the data, but we don't flush the grad. It stays there. When we do the second forward pass and we do backward again, remember that all the backward operations do a plus equals on the grad, and so these gradients just add up, and they never get reset to zero. Basically, we didn't zero grad. Here's how we zero grad. Before backward, we need to iterate over all the parameters, and we need to make sure that p dot grad is set to zero.

2:07:03

We need to reset it to zero, just like it is in the constructor. Remember, all the way here, for all these value nodes, grad is reset to zero, and then all these backward passes do a plus equals from that grad, but we need to make sure that we reset these grads to zero so that when we do backward, all of them start at zero, and the actual backward pass accumulates the loss derivatives into the grads. This is zero grad in PyTorch, and we will get a slightly different optimization. Let's reset the neural net. The data is the same. This is now, I think, correct, and we get a much more, we get a much more slower descent. We still end up with pretty good results,

2:07:34

and we can continue this a bit more to get down lower and lower and lower. Yeah. The only reason that the previous thing worked, it's extremely buggy, the only reason that worked is that this is a very, very simple problem, and it's very easy for this neural net to fit this data. The grads ended up accumulating, and it effectively gave us a massive step size, and it made us converge extremely fast. Basically, now we have to do more steps to get to very low values of loss and get YPRED to be really good. We can try to step a bit greater. Yeah. We're going to get closer and closer to one, minus one, and one. Working with neural nets is sometimes tricky because

2:08:09

you may have lots of bugs in the code, and your network might actually work, just like ours worked. But chances are that if we had a more complex problem, then actually this bug would have made us not optimize the loss very well, and we were only able to get away with it because the problem is very simple. Let's now bring everything together and summarize what we learned. What are neural nets? Neural nets are these mathematical expressions, fairly simple mathematical expressions in the case of multi-layer perceptron, that take input as the data, and they take input the weights and the parameters of the neural net. Mathematical expression for the forward pass,

2:08:42

followed by a loss function. The loss function tries to measure the accuracy of the predictions, and usually the loss will be low when your predictions are matching your targets or where the neural network is basically behaving well. We manipulate the loss function so that when the loss is low, the network is doing what you want it to do on your problem. Then we backward the loss, use backpropagation to get the gradient, and then we know how to tune all the parameters to decrease the loss locally.

2:09:19

But then we have to iterate that process many times in what's called gradient descent. We simply follow the gradient information, and that minimizes the loss, and the loss is arranged so that when the loss is minimized, the network is doing what you want it to do. Yeah. We just have a blob of neural stuff, and we can make it do arbitrary things, and that's what gives neural nets their power. This is a very tiny network with 41 parameters, but you can build significantly more complicated neural nets with billions, at this point almost trillions, of parameters. It's a massive blob of neural tissue, simulated neural tissue, roughly speaking, and you can make it

2:10:07

do extremely complex problems. These neural nets then have all kinds of very fascinating emergent properties when you try to make them do significantly hard problems, as in the case of GPT, for example. We have massive amounts of text from the internet, and we're trying to get a neural net to predict, to take a few words and try to predict the next word in a sequence. That's the learning problem. It turns out that when you train this on all of internet, the neural net actually has really remarkable emergent properties. That neural net would have hundreds of billions of parameters, but it works on fundamentally the exact same principles. The neural net, of course, will be

2:10:57

a bit more complex, but otherwise the value and the gradient is there and will be identical, and the gradient descent would be there and will be basically identical. People usually use slightly different updates. This is a very simple stochastic gradient descent update, and the loss function would not be a mean squared error. They would be using something called the cross entropy loss for predicting the next token. There's a few more details, but fundamentally the neural network setup and neural network training is identical and pervasive, and now you understand intuitively how that works under the hood. In the beginning of this video, I told you that by the end of it

2:11:44

you would understand everything in micrograd, and then we'd slowly build it up. Let me briefly prove that to you. I'm going to step through all the code that is in micrograd as of today. Actually, potentially some of the code will change by the time you watch this video because I intend to continue developing micrograd, but let's look at what we have so far, at least. init.py is empty. When you go to engine.py, that has the value. Everything here you should mostly recognize. We have the .data.grad attributes. We have the backward function. We have the previous set of children and the operation that produced this value. We have addition, multiplication, and raising to a

2:12:45

scalar power. We have the relu nonlinearity, which is a slightly different type of nonlinearity than 10h that we used in this video. Both of them are nonlinearities, and notably, 10h is not actually present in micrograd as of right now, but I intend to add it later. We have the backward, which is identical you would understand everything in micrograd, and then we'd slowly build it up. Let me briefly prove that to you. I'm going to step through all the code that is in micrograd as of today. Actually, potentially, some of the code will change by the time you watch this video because I intend to continue developing micrograd, but let's look at what we have so far.

2:13:19

At least init.py is empty. When you go to engine.py, that has the value. Everything here you should mostly recognize. We have the .data.grad attributes, we have the backward function, we have the previous set of children, and the operation that produced this value. We have addition, multiplication, and raising to a scalar power. We have the relu nonlinearity, which is a slightly different type of nonlinearity than 10h that we used in this video. Both of them are nonlinearities, and notably, 10h is not actually present in micrograd as of right now, but I intend to add it later. We have the backward, which is identical, and then all of these other operations, which are built up on top of operations here. So values should be very recognizable, except for the nonlinearity used in this video.

2:13:21

There's no massive difference between relu and 10h and sigmoid and these other nonlinearities. They're all roughly equivalent and can be used in MLPs. I use 10h because it's a bit smoother and because it's a little bit more complicated than relu, and therefore it stresses a little bit more the local gradients and working with those derivatives, which I thought would be useful.

2:13:23

nn.py is the neural networks library, as I mentioned, so you should recognize the identical implementation of neuron, layer, and mlp. Notably, or not so much, we have a class module here. There's a parent class of all these modules. I did that because there's an nn.module class in PyTorch, and so this exactly matches that API. nn.module in PyTorch also has a 0 grad, which I refactored out here.

2:13:26

So that's the end of micrograd, really. Then there's a test, which you'll see basically creates two chunks of code, one in micrograd and one in PyTorch, and we'll make sure that the forward and the backward pass agree identically for a slightly less complicated expression and a slightly more complicated expression. Everything agrees, so we agree with PyTorch on all of these operations.

2:13:27

Finally, there's a demo.ipy and b here, and it's a bit more complicated binary classification demo than the one I covered in this lecture. We only had a tiny data set of four examples. Here we have a bit more complicated example with lots of blue points and lots of red points, and we're trying to, again, build a binary classifier to distinguish two-dimensional points as red or blue. It's a bit more complicated MLP here. It's a bigger MLP. The loss is a bit more complicated because it supports batches. Because our data set was so tiny, we always did a forward pass on the entire data set of four examples, but when your data set is a million examples, what we usually do in practice is pick out some random subset. We call that a batch, and then we only process the batch: forward, backward, and update. So we don't have to forward the entire training set. This supports batching because there are a lot more examples here.

2:13:34

We do a forward pass. The loss is slightly more different. This is a max margin loss that I implement here. The one that we used was the mean squared error loss because it's the simplest one. There's also the binary cross entropy loss. All of them can be used for binary classification and don't make too much of a difference in the simple examples that we looked at so far.

2:13:34

There's something called L2 regularization used here. This has to do with generalization of the neural net and controls the overfitting in a machine learning setting, but I did not cover these concepts in this video, potentially later. The training loop you should recognize: forward, backward, with zero grad and update, and so on. You'll notice that in the update here, the learning rate is scaled as a function of number of iterations, and it shrinks. This is something called learning rate decay. In the beginning, you have a high learning rate, and as the network stabilizes near the end, you bring down the learning rate to get some of the fine details in the end.

2:13:36

In the end, we see the decision surface of the neural net, and we see that it learned to separate out the red and the blue area based on the data points. So that's the slightly more complicated example in the demo.ipy that you're free to go over. As of today, that is micrograd.

2:13:38

I also wanted to show you a little bit of real stuff so that you get to see how this is actually implemented in a production-grade library like PyTorch. In particular, I wanted to find and show you the backward pass for 10h in PyTorch. Here in micrograd, we see that the backward pass for 10h is 1 minus t square, where t is the output of the 10h of x, times that grad, which is the chain rule. We're looking for something that looks like this.

2:13:39

I went to PyTorch, which has an open-source github codebase, and I looked through a lot of its code. Honestly, I spent about 15 minutes, and I couldn't find 10h, and that's because these libraries, unfortunately, grow in size and entropy. If you just search for 10h, you get apparently 2800 results and 406 files. I don't know what these files are doing, honestly, and why there are so many mentions of 10h, but unfortunately these libraries are quite complex. They're meant to be used, not really inspected.

2:13:40

Eventually, I did stumble on someone who tries to change the 10h backward code for some reason, and someone here pointed to the CPU kernel and the CUDA kernel for 10h backward. This basically depends on if you're using PyTorch on a CPU device or on a GPU, which are different devices, and I haven't covered this. This is the 10h backward kernel for CPU, and the reason it's so large is that, number one, this is if you're using a complex type, which we haven't even talked about. If you're using a specific data type of bfloat16, which we haven't talked about, and then if you're not, then this is the kernel.

2:13:41

Deep here, we see something that resembles our backward pass. They have a times 1 minus b square, so this b here must be the output of the 10h, and this is the out.grad. Here we found it deep inside PyTorch, at this location, for some reason inside binary ops kernel, when 10h is not actually a binary op. Then this is the GPU kernel. We're not complex. We're here, and here we go with one line of code. We did find it, but unfortunately these codebases are very large, and micrograd is very, very simple. If you actually want to use real stuff, finding the code for it, you'll actually find that difficult.

2:13:44

I also wanted to show you an example here where PyTorch is showing you how you can register a new type of function that you want to add to PyTorch as a Lego building block. Here, if you want to, for example, add a Legendre polynomial 3, here's how you can do it. You will register it as a class that subclasses torch.argrad function, and then you have to tell PyTorch how to forward your new function and how to backward through it. As long as you can do the forward pass of this little function piece that you want to add and the backward pass for it, then you can use this as a Lego block in a larger Lego castle of all the different Lego blocks that PyTorch already has. That's the only thing you have to tell PyTorch, and everything would just work, and you can register new types of functions in this way following this example.

2:13:46

That is everything that I wanted to cover in this lecture. I hope you enjoyed building out micrograd with me. I hope you find it interesting, insightful, and I will post a lot of the links that are related to this video in the video description below. I will also probably post a link to a discussion forum or discussion group where you can ask questions related to this video, and then I can answer, or someone else can answer, your questions. I may also do a follow-up. Please subscribe or share if you can, so that YouTube knows to feature this video to more people, and that's it for now. I'll see you later. Now here's the problem. We know dl by. Wait, what is the problem?

2:13:49

That's everything I wanted to cover in this lecture, so I hope you enjoyed us building micrograd. Micrograd. Now let's do the exact same thing for multiply because we can't do something like a times two. Oops, I know what happened there. is that if we had a more complex problem then actually this bug would have made us not optimize the loss very well and we were only able to get away with it because the problem is very simple so let's now bring everything together and summarize what we learned what are neural nets neural nets are these mathematical expressions fairly simple mathematical expressions in the case of multi-layer perceptron that take input as the data

2:14:17

and they take input the weights and the parameters of the neural net mathematical expression for the forward pass followed by a loss function and the loss function tries to measure the accuracy of the predictions and usually the loss will be low when your predictions are matching your targets or where the new network is basically behaving well so we manipulate the loss function so that when the loss is low the network is doing what you want it to do on your problem and then we backward the loss use backpropagation to get the gradient and then we know how to tune all the parameters to decrease the loss locally but then we have to iterate that process many times

2:14:54

in what's called the gradient descent so we simply follow the gradient information and that minimizes the loss and the loss is arranged so that when the loss is minimized the network is doing what you want it to do and yeah so we just have a blob of neural stuff and we can make it do arbitrary things and that's what gives neural nets their power it's you know this is a very tiny network with 41 parameters but you can build significantly more complicated neural nets with billions at this point almost trillions of parameters and it's a massive blob of neural tissue simulated neural tissue roughly speaking and you can make it do extremely complex problems

2:15:35

and these neural nets then have all kinds of very fascinating emergent properties in when you try to make them do significantly hard problems as in the case of GPT for example we have massive amounts of text from the internet and we're trying to get a neural net to predict to take like a few words and try to predict the next word in a sequence that's the learning problem and it turns out that when you train this on all of internet the neural net actually has like really remarkable emergent properties but that neural net would have hundreds of billions of parameters but it works on fundamentally the exact same principles the neural net of course will be a bit more complex

2:16:12

but otherwise the value in the gradient is there and will be identical and the gradient descent would be there and will be basically identical but people usually use slightly different updates this is a very simple stochastic gradient descent update and the loss function would not be a mean squared error they would be using something called the cross entropy loss for predicting the next token so there's a few more details but fundamentally the neural network setup and neural network training is identical and pervasive and now you understand intuitively how that works under the hood in the beginning of this video I told you that by the end of it you would understand

2:16:48

everything in micrograd and then we'd slowly build it up let me briefly prove that to you so I'm going to step through all the code that is in micrograd as of today actually potentially some of the code will change by the time you watch this video because I intend to continue developing micrograd but let's look at what we have so far at least init.py is empty when you go to engine.py that has the value everything here you should mostly recognize so we have the .data.grad attributes we have the backward function we have the previous set of children and the operation that produced this value we have addition multiplication and raising to a scalar power we have the

2:17:25

relu nonlinearity which is a slightly different type of nonlinearity than 10h that we used in this video both of them are nonlinearities and notably 10h is not actually present in micrograd as of right now but I intend to add it later we have the backward which is identical and then all of these other operations which are built up on top of operations here so values should be very recognizable except for the nonlinearity used in this video there's no massive difference between relu and 10h and sigmoid and these other nonlinearities they're all roughly equivalent and can be used in mlps so I use 10h because it's a bit smoother and because it's a little bit more complicated

2:18:01

than relu and therefore it's stressed a little bit more the local gradients and working with those derivatives which I thought would be useful nn.py is the neural networks library as I mentioned so you should recognize identical implementation of neuron layer and mlp notably or not so much we have a class module here there's a parent class of all these modules I did that because there's an nn.module class in PyTorch and so this exactly matches that API and nn.module in PyTorch has also a 0 grad which I refactored out here so that's the end of micrograd really then there's a test which you'll see basically creates two chunks of code one in micrograd and one in PyTorch

2:18:46

and we'll make sure that the forward and the backward pass agree identically for a slightly less complicated expression and slightly more complicated expression everything agrees so we agree with PyTorch on all of these operations and finally there's a demo that I, pi, y, and b here and it's a bit more complicated binary classification demo than the one I covered in this lecture so we only had a tiny data set of four examples here we have a bit more complicated example with lots of blue points and lots of red points and we're trying to again build a binary classifier to distinguish two-dimensional points as red or blue it's a bit more complicated MLP here with

2:19:22

it's a bigger MLP the loss is a bit more complicated because it supports batches so because our data set was so tiny we always did a forward pass on the entire data set of four examples but when your data set is like a million examples what we usually do in practice is we basically pick out some random subset we call that a batch and then we only process the batch forward, backward and update so we don't have to forward the entire training set so this supports batching because there's a lot more examples here we do a forward pass the loss is slightly more different this is a max margin loss that I implement here the one that we used was the mean squared error

2:20:02

loss because it's the simplest one there's also the binary cross entropy loss all of them can be used for binary classification and don't make too much of a difference in the simple examples that we looked at so far there's something called L2 regularization used here this has to do with generalization of the neural net and controls the overfitting in machine learning setting but I did not cover these concepts in this video potentially later and the training loop you should recognize so forward backward with zero grad and update and so on you'll notice that in the update here the learning rate is scaled as a function of number of iterations and it shrinks and this is

2:20:42

something called learning rate decay so in the beginning you have a high learning rate and as the network sort of stabilizes near the end you bring down the learning rate to get some of the fine details in the end and in the end we see the decision surface of the neural net and we see that it learned to separate out the red and the blue area based on the data points so that's the slightly more complicated example in the demo.ipy that you're free to go over but yeah as of today that is micrograd I also wanted to show you a little bit of real stuff so that you get to see how this is actually implemented in a production grade library like PyTorch so in particular I wanted

2:21:18

to find and show you the backward pass for 10h in PyTorch so here in micrograd we see that the backward pass for 10h is 1 minus t square where t is the output of the 10h of x times of that grad which is the chain rule so we're looking for something that looks like this now I went to PyTorch which has an open source github codebase and I looked through a lot of its code and honestly I spent about 15 minutes and I couldn't find 10h and that's because these libraries unfortunately they grow in size and entropy and if you just search for 10h you get apparently 2800 results and 406 files so I don't know what these files are doing honestly and why there are so many mentions

2:22:08

of 10h but unfortunately these libraries are quite complex they're meant to be used not really inspected eventually I did stumble on someone who tries to change the 10h backward code for some reason and someone here pointed to the CPU kernel and the CUDA kernel for 10h backward so this so basically depends on if you're using PyTorch on a CPU device or on a GPU which these are different devices and I haven't covered this but this is the 10h backward kernel for CPU and the reason it's so large is that number one this is like if you're using a complex type which we haven't even talked about if you're using a specific data type of bfloat16 which we haven't talked about

2:22:52

and then if you're not then this is the kernel and deep here we see something that resembles our backward pass so they have a times 1 minus b square so this b here must be the output of the 10h and this is the out dot grad so here we found it deep inside PyTorch on this location for some reason inside binary ops kernel when 10h is not actually a binary op and then this is the GPU kernel we're not complex we're here and here we go with one line of code so we did find it but basically unfortunately these codebases are very large and micrograd is very very simple but if you actually want to use real stuff finding the code for it you'll actually find that difficult I also

2:23:44

wanted to show you a example here where PyTorch is showing you how you can register a new type of function that you want to add to PyTorch as a Lego building block so here if you want to for example add a Legendre polynomial 3 here's how you can do it you will register it as a class that subclass says torch.argrad that function and then you have to tell PyTorch how to forward your new function and how to backward through it so as long as you can do the forward pass of this little function piece that you want to add and as you can use this as a Lego block in a larger Lego castle of all the different Lego blocks that PyTorch already has and so that's the only thing you have

2:24:32

to tell PyTorch and everything would just work and you can register new types of functions in this way following this example and that is everything that I wanted to cover in this lecture so I hope you enjoyed building out micro grad with me I hope you find interesting insightful and yeah I will post a lot of the links that are related to this video in the video description below I will also probably post a link to a discussion forum or discussion group where you can ask questions related to this video and then I can answer or someone else can answer your questions and I may also do a follow so that YouTube knows to feature this video to more people and that's it for now

2:25:16

I'll see you later now here's the problem we know dl by wait what is the problem and that's everything I wanted to cover in this lecture so I hope you enjoyed us building a micro grab micro grab okay now let's do the exact same thing for multiply because we can't do something like a times two oops I know what happened there

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note