Sab kuch start hota hai ek sawaal se: kya ek single brain cell reasoning kar sakti hai?
Spanish neurologist jisne prove kiya ki brain separate cells se bana hai jo signals ek direction mein pass karte hain. "Father of Modern Neuroscience."
McCulloch & Pitts (1943) — inhone neuron ko exact math mein convert kiya aur poochha: yeh kya compute kar sakta hai? Answer: koi bhi logic gate!
Neuron ek committee ki tarah hai. Har member (input) vote karta hai with different weight. Agar total votes threshold cross kare — fire ho jaata hai! Weights determine karte hain ki kiska vote zyada matter karta hai.
Bohot exciting result tha: ek single neuron koi bhi logic gate ban sakta hai!
| Gate | Weights | Threshold | Kab Fire karta hai |
|---|---|---|---|
| AND | w₁=1, w₂=1 | ≥ 2 | Sirf jab DONO inputs = 1 |
| OR | w₁=1, w₂=1 | ≥ 1 | Jab koi bhi EK input = 1 |
| NOT | w = -1 | ≥ 0 | Jab input = 0 (negative weight!) |
AND + OR + NOT se koi bhi logic circuit bana sakte hain. Matlab: neurons ka network = any computer!
But ek problem: weights kisi ne manually set kiye the. Network khud kuch nahi seekh raha tha!
Perceptron = pehla model jisme weights automatically data se learn kiye ja sakte the!
x₁ aur x₂ sirf first power pe hain — koi x², koi x₁x₂ nahi. Isliye boundary hamesha ek straight line rahegi. Activation function boundary change nahi karta — sirf sign check karta hai.
Minsky & Papert (1969) ne prove kiya: single perceptron XOR solve NAHI kar sakta. Isse First AI Winter aaya!
XOR = "exactly one switch ON." Jaise ghar ki staircase light jo do switches se control hoti hai — ek on to light on, dono on to light off!
| x₁ | x₂ | XOR | Meaning |
|---|---|---|---|
| 0 | 0 | 0 | Dono off |
| 0 | 1 | 1 | Ek on |
| 1 | 0 | 1 | Ek on |
| 1 | 1 | 0 | Dono on?! off! |
Problem: ek neuron = ek straight line. Agar do neurons ko different kaam dein, phir combine karein?
Ghar ki staircase light do switches se control hoti hai. Ek switch on = light on. Dono on = light off (wiring cancel). Yahi XOR hai!
Electrician ek switch se nahi, two-switch combination se solve karta hai. Neural network same — hidden layer!
Do hidden neurons ko alag kaam dein:
| x₁ | x₂ | h₁ (OR) | h₂ (NAND) | y = h₁ AND h₂ | XOR? |
|---|---|---|---|---|---|
| 0 | 0 | 0 | 1 | 0 AND 1 = 0 | ✅ |
| 0 | 1 | 1 | 1 | 1 AND 1 = 1 | ✅ |
| 1 | 0 | 1 | 1 | 1 AND 1 = 1 | ✅ |
| 1 | 1 | 1 | 0 | 1 AND 0 = 0 | ✅ |
Hidden layer ek naya representation space create karta hai. (x₁,x₂) space mein XOR non-linearly separable tha. But (h₁,h₂) space mein yeh linearly separable ho jaata hai!
Yahi hai hidden layers ki asli superpower: inputs ko aisa transform karo ki output layer ki job easy ho jaaye.
Raw data. Koi computation nahi!
Neurons = features in data
Example: 10 features → 10 neurons
Non-linear features seekhta hai.
a = φ(Wx + b)
Yahi XOR solve karta hai!
Final prediction. Size = problem type.
Binary: 1 neuron (sigmoid)
k classes: k neurons (softmax)
Regression: 1 (linear)
Think like a postal address: Country → City → Street → House. Weight notation bhi aisa hai!
| Symbol | Matlab kya hai | Example |
|---|---|---|
| x𝓎 | i-th input feature | x₁ = first column value |
| W[l] | Weight matrix for layer l | W[1] = input→hidden weights |
| b[l] | Bias vector for layer l | b[1] = hidden biases |
| z[l] | Pre-activation (Wx+b, before φ) | z[1] = W[1]x + b[1] |
| a[l] | Post-activation (after φ) | a[1] = φ(z[1]) |
| φ or σ | Activation function | sigmoid, ReLU, etc. |
Boundary equation: w₁x₁ + w₂x₂ + b = 0
x₁ aur x₂ dono first power pe hain. x₁² nahi, x₁x₂ nahi, koi bhi curve nahi.
School algebra: ax + by + c = 0 hamesha ek straight line hoti hai. W linear combination → boundary straight.
Activation function (step ya sigmoid) sirf sign check karta hai — boundary nahi badalta!
Boundary: 2x₁ + 3x₂ - 12 = 0
At x₁=0: x₂ = 4 | At x₂=0: x₁ = 6
Above line (2x₁+3x₂-12 > 0): Class 1
Below line (2x₁+3x₂-12 < 0): Class 0
Forward propagation = input leke prediction nikalna. Data layer by layer aage badhta hai.
Network mein kitne learnable parameters hain? Weights + Biases = Total Parameters.
| Layer | Weights | Biases | Total |
|---|---|---|---|
| Input→Hidden (2→3) | 2×3 = 6 | 3 | 9 |
| Hidden→Output (3→1) | 3×1 = 3 | 1 | 4 |
| TOTAL | 9 | 4 | 13 |
Hidden: 3×3 = 9 weights + 3 biases = 12
Output: 3×1 = 3 weights + 1 bias = 4
Total = 16 parameters
Sigmoid aur Tanh — dono S-shaped curves hain. Jab z bahut bada ho jaata hai, curve flat ho jaati hai. Flat curve = near-zero derivative = almost no learning!
Given: x=1, w=6, b=0, y=0, α=1
This is Vanishing Gradient: |z| bada hone pe sigmoid flat → derivative ≈ 0 → no learning!
Tanh ka max derivative 1.0 hai (sigmoid ka 0.25). Near z=0 tanh better hai. But z=1.663 pe dono cross karte hain, aur uske baad tanh aur tezi se flat ho jaati hai!
S-shape hi problem hai — koi bhi rescaling fix nahi karega.
S-curve band karo. Ek function use karo jo positive side pe kabhi flat nahi hoti!
| Activation | f'(z=6) | Step Taken | Ratio |
|---|---|---|---|
| Sigmoid | 0.00247 | 0.00246 | 1× |
| ReLU | 1 | 6 | ≈2439× bigger! |
ReLU same problem 1 step mein zero error karta hai jahan sigmoid thousands of steps leta!
Agar z negative ho gaya (e.g., w=-3, x=1 → z=-3):
f(-3) = 0, f'(-3) = 0 → gradient = 0 → weight update = 0 → neuron PERMANENTLY DEAD!
No learning rate can rescue a zero gradient. Ek baar dead = hamesha dead.
ReLU negative side mein f'=0. Fix: negative side ko tiny slope do — 0 nahi, 0.01!
| Activation | ŷ | f'(-3) | Gradient | Dead? |
|---|---|---|---|---|
| ReLU | 0 | 0 | 0 | ☠️ YES |
| Leaky ReLU | -0.03 | 0.01 | -0.0103 | ✅ No, but crawling |
Leaky ReLU: negative side straight line (slope 0.01). ELU: negative side ko smooth exponential curve se replace karo!
| Activation | ŷ | f'(-3) | Gradient | Verdict |
|---|---|---|---|---|
| ReLU | 0 | 0 | 0 | ☠️ Dead |
| Leaky ReLU | -0.03 | 0.01 | -0.0103 | 🐌 Crawling |
| ELU | -0.950 | 0.0498 | -0.097 | ✅ Moving properly! |
ELU step ≈9.4× Leaky ReLU se bada! More meaningful signal on negative side.
Sigmoid = ek class (yes/no). Agar 3 classes: cat, dog, bird? Teen separate sigmoids ka sum 1 nahi hoga! Softmax = ek function jo sabko simultaneously handle kare.
n=2 leke karo: e𝑧/(e𝑧+e⁰) = e𝑧/(e𝑧+1) = 1/(1+e⁻𝑧) = sigmoid!
Binary classification ek special case hai — separate idea nahi!
| Function | Formula | Range | Max f' | Use Case | Problem |
|---|---|---|---|---|---|
| Sigmoid | 1/(1+e⁻𝑧) | (0,1) | 0.25 | Output binary, LSTM gates | Vanishing gradient |
| Tanh | (e𝑧-e⁻𝑧)/(e𝑧+e⁻𝑧) | (-1,1) | 1 | RNNs, old shallow nets | Saturates faster |
| ReLU | max(0,z) | [0,∞) | 1 | Default hidden (CNNs) | Dying neurons (z<0) |
| Leaky ReLU | z or 0.01z | (-∞,∞) | 1 | GANs (dying destabilizes) | Tiny gradient, hand-tuned |
| ELU | z or e𝑧-1 | (-1,∞) | 1 | When ReLU units dying | Slower (exponential) |
| Softmax | e𝑧𝓪/Σe𝑧ⱼ | (0,1), sum=1 | - | Output multi-class | Layer-wise only |
Step activation use karo: output 1 if z≥0, else 0.
| x₁ | x₂ | h₁ (OR) | h₂ (AND) | h₁-h₂-0.5 | ŷ | XOR? |
|---|---|---|---|---|---|---|
| 0 | 0 | 0 | 0 | -0.5 | 0 | ✅ |
| 0 | 1 | 1 | 0 | 0.5 | 1 | ✅ |
| 1 | 0 | 1 | 0 | 0.5 | 1 | ✅ |
| 1 | 1 | 1 | 1 | -0.5 | 0 | ✅ |
Har hidden neuron ek line draw karta hai. OR line aur AND line parallel hain. XOR = strip between them.
Ek line kabhi XOR separate nahi kar sakti. Do lines mili ke — kar sakti hain!
Forward pass = prediction nikalna (left se right). Backward pass = mistakes se seekhna (right se left, chain rule use karke).
Composite functions: L depends on ŷ, ŷ depends on z[2], z[2] depends on w₂.
∂L/∂w₂ = (∂L/∂ŷ) × (∂ŷ/∂z[2]) × (∂z[2]/∂w₂)
Har weight ke liye chain continue hoti hai jab tak us weight tak nahi pahuncho. Deep networks mein chain bahut lambi hoti hai!
Sigmoid ka max derivative 0.25. 10 layers mein: 0.25¹⁰ ≈ 0.000001! First layers ko practically zero gradient milta hai — they learn nothing.
Isliye deep networks mein ReLU zaroori hai — f'=1 hamesha (positive side pe), multiplication se shrink nahi hota!
Pahari pe blindfold lagake khadhe ho. Minimum nahi dekh sakte. But local slope feel kar sakte ho.
Strategy: slope ke opposite direction mein ek chota step lo. Baar baar karo — eventually bottom!
Yahi hai gradient descent. Loss function = pahari. Parameters = tumhari position.
Gradient negative hai (w < 3), toh subtraction w ko badhata hai → minimum ki taraf!
Sab N examples ek saath use karo
✅ Accurate gradient
✅ Smooth convergence
❌ N=1M → huge memory, slow
1 update per epoch
Ek random example, turant update
✅ Super fast
✅ Noise helps escape local minima
❌ Very noisy zigzag path
N updates per epoch
Small group: 32, 64, 128 examples
✅ GPU efficiently use karta
✅ Balanced noise vs speed
✅ Modern DL ka default!
N/B updates per epoch
| Observation | Diagnosis | α kaisa hai? |
|---|---|---|
| Loss bahut slowly gir rahi hai | Stagnation | Too LOW ⬆️ |
| Loss gir rahi but bounce karti hai | Oscillation | Slightly HIGH |
| Loss tezi se badh rahi (NaN!) | Divergence | Much too HIGH 💥 |
Fixed α = ek hi speed everywhere — highway aur parking dono pe! Better: schedule karo.
Milestones pe suddenly reduce
Graph: staircase
Simple
Abrupt drops
Smoothly reduce — same % every step
αₜ = α₀ × e⁻𝛋ₜ
Smooth
Cosine curve follow karo
Gentle "soft landing"
Modern default!
Pehle chota α, phir badhao!
Early instability prevent
Transformers mein essential
Val loss improve nahi → reduce
Reactive strategy
No total epoch needed
SGD = boat steering by only the newest ripple. One bad ripple → wrong direction!
Memory optimisers = boat keeping a compass heading (direction history) and a sea chart (scale history). One noisy ripple doesn't derail them.
Cricket coach player ki form judge karta hai. Kal ki performance sabse zyada important, pichle week ki thodi kam, pichle mahine ki bahut kam. Coach ek "running impression" rakhta hai — ek hi number (sₜ)! EWMA exactly yahi hai.
EWMA sirf ek state sₜ store karta hai — puri history nahi!
EWMA ko gradients pe apply karo → velocity banao jo direction yaad rakhti hai.
Bicycle ek nayi direction mein ek chhoti wobble se immediately face nahi karta. Usme inertia hai! Momentum gradient zigzag ko smooth karta hai — overall downhill direction yaad rehti hai.
Momentum: gradient at current point measure karo. NAG: pehle predict karo ki momentum tum kahan le jaayega, WAHAN gradient measure karo!
Cyclist sharp bend approach karte waqt sirf current speed nahi dekhta. Bend par look karta hai — agar road upar hai, pehle se brake lagao! NAG: anticipate karo gradient kahan hoga, wahan se correct karo.
Momentum: "Yahan slope kya hai?" → gradient at current θₜ
NAG: "Momentum mujhe kahan le jaayega? Wahan slope kya hai?" → gradient at θ_look
Problem: different features ke gradients bahut different sizes ke ho sakte hain. Ek hi α sabke liye? Inefficient!
rₜ = sum of ALL past g² — never resets! Deep training mein rₜ → ∞ → step → 0. Learning stops!
Old gradients never forget — 1000 steps purana gradient aaj bhi utna hi influential hai.
Replace cumulative sum with EWMA — taaki puraane gradients fade ho jaayein!
| Property | AdaGrad | RMSProp |
|---|---|---|
| State | Cumulative Σg² | Fading EWMA of g² |
| Old scale | Stays forever | Repeatedly discounted |
| Late training | Increasingly timid | Can recover! |
Adam = Momentum + RMSProp. Direction memory AND scale memory dono ek saath!
m₀ = v₀ = 0 se start karte hain. Pehle steps mein moving averages zero ke paas hain (too much zero in them)!
Analogy: Class test average mein imaginary zero marks add karo — first real mark bahut chota lagta hai. Bias correction woh imaginary zeros remove karta hai!
t=1, β₁=0.9: m₁=0.4 (raw) → m̂₁ = 0.4/0.1 = 4. 10× correction! Very first step mein zaroori.
Regular L2 regularization Adam ke adaptive scaling ke saath interfere karta hai. Fix: decay aur adaptive step alag rakho!
Steering = destination ki taraf jaana (Adam adaptive step)
Lane keeping = road centre ke paas rehna (weight decay)
AdamW in dono mix nahi karta — alag rakhta hai! Regular Adam mein adaptive scaling se decay ka actual effect change ho jaata tha.
| Optimiser | Memory State | Solves | Limitation |
|---|---|---|---|
| SGD | None | Baseline | Noisy, zigzag |
| EWMA | sₜ (fading avg) | Smooth signals | Building block only |
| Momentum | vₜ (velocity) | Noisy/zigzag gradients | Can overshoot |
| NAG | vₜ + look-ahead | Overshoot prevention | More complex |
| AdaGrad | rₜ (Σg²) | Uneven coordinate scales | Denominator grows forever |
| RMSProp | sₜ (EWMA of g²) | Uneven scales + recovery | No direction memory |
| Adam | mₜ + vₜ | Direction + scale both | L2 reg interfered |
| AdamW | mₜ + vₜ + decay | Adam + proper weight decay | Extra λ hyperparameter |