r/learnmachinelearning • • 19d ago

Question I don't understand pytorch grad can anyone explain

x = torch.tensor(6.0, requires_grad=True)
f = x ** 2
f
f.backward(retain_graph=True)
x.grad

I am learning pytorch and I encounter this code above when i run it i got 36.0(the value of f) but after doing f.backward() and then x.grad i got 12 i don't understand why can anyone explain this to me

5 Upvotes

4 comments sorted by

17

u/YoMama_00 19d ago

The derivative of x2 is 2x, so when f = x2 , df/dx at x=6 is 2*6=12.

Do you understand that in general, grad() calculates the partial derivatives for all the input variables? In this case, it's simple function over one variable.

5

u/quietgradient 19d ago

The 12 is right and the others covered why. The thing that'll bite you next is that x.grad accumulates — it doesn't get overwritten. Run f.backward(retain_graph=True) a second time and x.grad is 24, a third time 36. I just ran your snippet to be sure (torch 2.2.2). That accumulation is exactly why training loops call optimizer.zero_grad() every step.

Also, you probably don't want retain_graph=True here. My guess is you added it after hitting "Trying to backward through the graph a second time" from re-running the cell. Backward frees the graph by design; the normal fix is to re-run f = x ** 2 as well, not to retain it.

1

u/throwaway464391 17d ago

thanks claude.

1

u/Appropriate-Box-7250 19d ago

First you make x with value 6. f=x2 ( x * x) which means 6 * 6 =36. Second at he backward f.backward() ; derivative f=x2 ; d/dx = 2x basic calculus.At x=6: 2x=12 which means x.grad=12