CurrentSky · Learning
Small tools.Big ideas.
Interactive mini-labs for AI, mathematics, physics, space and engineering. Drag a slider, change a number and watch the idea move. Free, right in your browser, nothing to install.
- 27 tools
- 5 topics
The periodic table of tools
Topic 01
Artificial Intelligence
How neural networks actually work, one idea at a time: from a single neuron to self-attention. Read the idea, then try it.
Basics
Training
Vision and language
Using models
In your product
Ne · AI
A single neuron
Change the weights and the bias, or drag the white point and follow it through the neuron below.
The idea
A neuron multiplies each input by a weight, adds a bias and passes the sum z through an activation that squashes it into an output between 0 and 1.
On a map of its two inputs, the points where z = 0 form a straight line. One side of the line gets an output near 1, the other near 0: a single neuron is a straight-line classifier.
Try this
- Move w1 and w2: the line turns. Move b: it slides without turning.
- Drag the white point across the line and watch the output flip.
- Press Let it learn: the neuron adjusts its own weights until the colours are split.
In real life
Spam filters, credit scoring and many medical risk scores are exactly this: a weighted sum pushed through a sigmoid (logistic regression). When one line separates your data, it is often all you need, and the weights tell you which input matters most.
How it works
z = w1·x1 + w2·x2 + b, output = σ(z) = 1 / (1 + e−z). The yellow arrow is the weight vector (w1, w2): it always points across the line, towards the side the neuron calls “1”.
Let it learn runs logistic regression: the same gradient descent that trains big networks, on just three numbers.
Af · AI
Activation functions
Pick a function. Top: its shape and slope. Bottom: what a small network built with it can draw.
- Slope at x = 0
- —
- Slope at x = 4
- —
The idea
After the weighted sum, every neuron applies an activation function. It has to bend: a stack of purely linear layers collapses into one straight line, however deep it is.
Its slope matters just as much, because training follows the slope. Where the slope is flat, learning stalls.
Try this
- Pick None (linear): the network below can only draw a straight line.
- Pick ReLU and press New random network: now it draws bends and corners.
- Compare the slope at x = 4 for Sigmoid and ReLU: that gap is the vanishing-gradient problem.
In real life
Choosing an activation is a real design decision. ReLU is the safe default for vision and tabular networks, GELU sits inside most transformers and LLMs, sigmoid is used for the final yes/no probability. If a deep model stops learning, saturated sigmoids are a classic suspect.
How it works
Sigmoid 1/(1+e−x) and tanh squash everything into a narrow range; their slopes shrink towards 0 for large inputs. ReLU, max(0, x), keeps a slope of 1 for positive inputs and is cheap to compute, which made very deep networks trainable.
GELU, a smooth cousin of ReLU, is used inside most modern transformers. The bottom chart is a real two-hidden-layer network with 8 neurons per layer and random weights.
Nn · AI
Neural network playground
Pick a dataset or click to add points. The coloured field is what the network believes.
Click the screen to add a point
What each hidden neuron learned (border: its vote, cyan for B, violet for A)
The idea
Put several neurons side by side and you get a hidden layer. Each hidden neuron draws its own straight line (the small tiles below the map); the output neuron then votes with them.
Combining many simple lines lets the network carve out curves, rings and islands that no single neuron could draw. Training finds all the lines at once.
Try this
- Pick XOR with 1 hidden neuron: it cannot solve it. Try 2, then 4.
- On Circle, watch the tiles: several lines together build a ring.
- Click to add points of the wrong colour inside a region and watch the boundary bend around them.
In real life
Every model you train hits the same trade-off: too few neurons underfit, too many overfit, and the learning rate decides whether training is slow, smooth or explosive. Engineers run small experiments like this one to pick a first network size before an expensive training run.
How it works
A neural network is a stack of simple functions. Here the two inputs (x, y) feed one hidden layer of tanh neurons, and a single sigmoid output gives the probability of class B.
Training repeats two steps: predict every point, then nudge every weight a little against the gradient of the cross-entropy loss (backpropagation, with momentum). With one neuron the boundary can only be a straight line; XOR needs at least two, and circles or moons need several. Try a learning rate that is too high and watch it struggle.
Gd · AI
Loss and gradient descent
Brighter means a higher loss. Click to drop the ball anywhere, then compare learning rates and optimizers.
The green rings mark the minima. The strip below the map plots the loss at every step.
The idea
Training a network means finding the weights with the lowest loss, the number that says how wrong the predictions are. Picture the loss as a landscape over the weights.
The gradient points uphill, so each step moves a little the other way. The learning rate sets the size of that step.
Try this
- Raise the learning rate step by step in the Narrow bowl: first it zigzags, then it explodes.
- In Two valleys, plain descent from the right stops in the shallower valley: a local minimum.
- In the Banana valley, compare Plain with Adam: Adam adapts its step to each direction.
In real life
Every neural network you use was trained by a loop like this. A practical default: start with Adam and a learning rate around 0.001, lower it if the loss jumps or turns into NaN, raise it if the loss barely moves. The loss curve is the first thing to check when a run misbehaves.
How it works
Plain gradient descent: w ← w − lr·∇L. Momentum keeps a running average of past gradients, so it rolls through narrow valleys and over small bumps. Adam also rescales every direction by its recent gradient size.
Real networks have millions of weights instead of two, and the gradient is estimated on small batches of data (stochastic gradient descent), but the picture is the same.
Bp · AI
Backpropagation, step by step
One neuron with one training example. Run the forward pass, then the backward pass, then update.
Press Forward to start.
The idea
The forward pass computes the prediction and the loss. The backward pass walks the same graph in reverse and asks of every node: if this value grew a little, how much would the loss grow? That number is its gradient.
Each node only needs its own local slope; the chain rule multiplies them together on the way back. That is backpropagation, and it scales to billions of weights.
Try this
- Press 1, 2, 3 in order and read the chain-rule lines under the graph.
- Press Update a few times: the loss on the right falls as w and b move.
- Set the target y to 0 and train again: the gradients flip sign.
In real life
You never code backpropagation by hand: PyTorch and TensorFlow do it automatically. Knowing it still explains vanishing and exploding gradients, why frozen layers do not learn, and why fine-tuning can update only a few layers.
How it works
Forward: m = w·x, z = m + b, a = σ(z), L = (a − y)². Backward: ∂L/∂a = 2(a − y), ∂L/∂z = ∂L/∂a · a(1 − a), ∂L/∂b = ∂L/∂z, ∂L/∂w = ∂L/∂z · x. Update: w ← w − lr·∂L/∂w, and the same for b.
Frameworks such as PyTorch build this graph automatically while you compute (automatic differentiation) and run exactly these steps for every weight.
Ov · AI
Overfitting and generalisation
Fit curves of rising complexity to noisy points. Filled points train the model; hollow ones only test it.
- Verdict
- —
- Training error
- —
- Validation error
- —
- Best complexity here
- —
Filled cyan: training points. Hollow amber: validation points. Dashed: the true curve.
The idea
A model that is too simple misses the pattern (underfitting). One that is too flexible bends to every noisy point and memorises it (overfitting): perfect on its training data, poor on anything new.
So we always keep some data the model never trains on. The validation error tells us how it will do in the real world; its lowest point is the sweet spot.
Try this
- Slide the degree from 0 to 14 and watch the two error curves split.
- At a high degree, raise λ: regularization calms the wiggles.
- Press More data: with enough examples, even a complex model behaves.
In real life
This is why every serious ML project splits data into training, validation and test sets. A model with 99% accuracy on its own training data can still fail on real customers. Early stopping, regularization, dropout and more data are the standard cures.
How it works
The model is a polynomial of the chosen degree, fitted by least squares (in a numerically stable Chebyshev basis); λ adds a penalty on large coefficients (ridge regression). The error is the mean squared difference between prediction and data.
Neural networks overfit in the same way. The usual cures are more data, regularization (weight decay, dropout) and stopping training when the validation error starts to rise.
Cv · AI
Convolution: how CNNs see
Draw on the left, pick a filter, and hover the result to see the exact multiply-and-add.
Amber: positive response. Violet: negative.
The idea
A convolution slides a small grid of weights, the kernel, across the image. At every position it multiplies the pixels under it by the weights and adds them up: one number of the output, the feature map.
Hand-made kernels find edges or blur. A convolutional network (CNN) learns its kernels from data and stacks them: early layers see edges, deeper layers see corners, textures, then eyes, wheels and letters.
Try this
- Pick Vertical edges, then Horizontal edges: each one only sees its own direction.
- Draw a shape on the input and watch the edges trace it.
- Press Watch it slide to see the window move across the image.
In real life
Convolutions power OCR, document and ID checks, face detection, defect inspection on factory lines, medical imaging and phone cameras. Edge filters like these were the first step of classic computer vision; a CNN simply learns better filters than anyone can hand-write.
How it works
output[y, x] = Σi,j input[y+j, x+i] · kernel[j, i], with zeros outside the image. The same nine weights are reused at every position, which is why CNNs need far fewer parameters than fully connected networks and why they recognise a pattern wherever it appears.
Real CNNs apply dozens of kernels per layer to colour images and add pooling to shrink the maps between layers.
Tk · AI
Tokens: how text becomes numbers
Type or paste any text. The tokenizer learns its pieces from it, merge by merge.
- Tokens per word
- —
The idea
Language models never see letters or words; they see tokens, chunks of text that each have an ID number. Byte-pair encoding (BPE) builds them: start from single characters, find the most frequent neighbouring pair, glue it into a new token, and repeat.
Frequent words end up as one token, rare words split into pieces. The dot marks the start of a word. This is why prices are per token. Real tokenizers are trained once on huge piles of text, mostly English, so other languages often need more tokens for the same meaning; here the pieces are learned from your own text, so you will not see that effect.
Try this
- Drag Merges learned from 0 upwards and watch characters fuse into syllables and words.
- Paste a longer text where words repeat: the more it repeats, the fewer tokens per word it needs.
- Type a made-up word: it stays split into small known pieces.
In real life
Tokens are the unit of every LLM bill and every context limit. A rough rule for English: one token is about 4 characters, or three quarters of a word. It helps you estimate API cost, trim prompts and see why a long document does not fit the context window.
How it works
Each step counts every adjacent pair of tokens across the text (weighted by how often each word appears) and merges the most frequent pair everywhere. Real tokenizers run tens of thousands of merges on huge corpora and work on raw bytes, so they can represent any text.
After tokenizing, each token ID is mapped to an embedding vector (see the Embeddings lab), and that is what the network actually computes with.
Em · AI
Embeddings: meaning as vectors
Click a word to see its nearest neighbours, or build an analogy with vector arithmetic.
- —
- —
- Nearest to king
- —
A toy space with 14 dimensions, drawn as a 2D sketch. Real models learn hundreds of dimensions from text.
The idea
An embedding gives every token a list of numbers, a point in a space with many dimensions. Training pushes words that appear in similar contexts close together, so closeness means similar meaning.
Directions mean something too: the step from man to woman is roughly the same as from king to queen, so king − man + woman lands next to queen.
Try this
- Click Paris: its neighbours are other capitals, then France.
- Run each example and notice the two arrows on the map are parallel.
- Build your own analogy, for example kitten − cat + dog.
In real life
Embeddings sit behind semantic search, recommendations, duplicate detection and RAG chatbots that answer from your own documents: turn each text into a vector once, then find the vectors closest to the question.
How it works
Similarity is the cosine of the angle between two vectors: 100% means the same direction. The analogy result is the word whose vector is closest to A − B + C, excluding the three inputs.
Search engines, recommendation systems and retrieval-augmented generation (RAG) all compare embeddings this way to find related text.
At · AI
Self-attention
Click any word to see where it looks. Then change the last word and watch “it” change its mind.
Each row of the grid is one word asking a question; brighter cells are the words it listens to. The weights in every row add up to 100%.
Weights here are illustrative, set by hand to show the classic example.
The idea
In a transformer, every word builds three vectors: a query (what am I looking for?), a key (what do I contain?) and a value (what do I pass on?). Each query is compared with every key; softmax turns the scores into weights that add up to 1.
The word’s new representation is the weighted mix of the values. That is how “it” can take its meaning from “animal” in one sentence and from “street” in another.
Try this
- Click it, then switch the ending between tired and wide.
- Click The and the: each one looks at its own noun.
- Lower the Sharpness and watch attention spread out evenly.
In real life
Attention is the core of every modern LLM. It explains why a model can use a fact from early in the prompt, why cost grows quickly with long inputs, and why the placement of key instructions in a prompt matters.
How it works
Attention(Q, K, V) = softmax(Q·KT / √d)·V. Dividing by √d keeps the scores in a range where softmax still spreads its weight; the Sharpness slider multiplies the scores to show what happens without it.
Real models run many attention heads in parallel, each learning its own kind of relation (grammar, references, position), and stack dozens of such layers.
Tp · AI
LLM temperature and top-p
A model scores every possible next word. Temperature and top-p decide how it picks one.
The sky over the city is ___
- Entropy
- —
- Words left
- —
The idea
A language model does not write sentences directly. At each step it scores every possible next token, turns the scores into probabilities with softmax, and samples one. Then it repeats with the new token added.
Temperature and top-p change only that last step, which is why the same model can sound careful or creative.
Try this
- Set temperature to 0.05 and sample: always “blue”, as predictable as it gets.
- Raise it to 2 and press Sample ×100: rare words like “falling” start to appear.
- Lower top-p to 0.5: the unlikely words are cut off entirely.
In real life
Rules of thumb for API settings: temperature 0 to 0.3 for extraction, classification, code and anything that must be repeatable; 0.7 to 1 for writing and brainstorming. Change temperature or top-p, not both at once.
How it works
The model gives each candidate a score z (a logit). Softmax turns scores into probabilities: pi = ezi/T / Σ ezj/T. A low temperature T sharpens the distribution toward the favourite; a high one flattens it, so rare words appear more often.
Top-p (nucleus sampling) keeps only the smallest set of words whose probabilities add up to p, then renormalizes. Entropy, in bits, measures how unpredictable the next word is. The scores here are an illustrative example, not taken from a real model.
F1 · AI
Precision, recall and F1
Type the four counts from your test set. The metrics and a plain-language reading update instantly.
- F1 score
- —
- Precision
- —
- Recall
- —
- Accuracy
- —
- Specificity
- —
- Balanced accuracy
- —
- MCC
- —
The idea
After training we have to measure how good a classifier really is. Every prediction lands in one of four boxes of the confusion matrix: caught, missed, false alarm or correct no.
Different jobs care about different boxes: a fraud filter must not miss fraud (recall), a spam filter must not bin real mail (precision). F1 balances the two.
Try this
- Open the Rare-disease test: 90% accurate, yet most positive results are wrong.
- Move counts from FN to TP and watch recall climb.
- Make FP large on purpose and see precision and F1 drop while accuracy barely moves.
In real life
Use these numbers whenever you judge a detector: fraud, spam, defects, forged documents. Decide which mistake costs more, a missed case or a false alarm, then tune the decision threshold for precision or recall instead of chasing raw accuracy.
How it works
Precision = TP / (TP + FP): of everything flagged, how much was right. Recall = TP / (TP + FN): of everything real, how much was caught. F1 = 2·P·R / (P + R) is their harmonic mean, so it is only high when both are.
Accuracy can mislead when one class is rare: the rare-disease example is 90% accurate, yet most positive results are false alarms. The Matthews correlation coefficient (MCC, from −1 to 1) uses all four cells and stays honest on imbalanced data.
Ct · AI
LLM API cost calculator
Enter your traffic and token counts. Prices are per million tokens, so you can paste any provider’s numbers.
Prefixes work: 1.5k, 2M. Price tiers below are typical, not a quote: check your provider’s price page.
- Per month
- —
- Per request
- —
- Per day
- —
- Per year
- —
- Output share
- —
- Small tier
- —
- Mid tier
- —
- Large tier
- —
The three boxes show the monthly cost of the same traffic on each tier.
The idea
Language-model APIs charge per token, and you pay twice: for the tokens you send (input) and for the tokens the model writes back (output). Output tokens usually cost three to five times more.
Multiply tokens by price, then by how many requests you serve, and you get a real budget before you build.
Try this
- Raise Output tokens from 400 to 2,000: the output share of the bill jumps.
- Pick the Small tier: the same traffic costs a fraction of the Large tier.
- Double Requests per day: the cost doubles, nothing else changes.
In real life
Use it before you launch a chatbot, an agent or a document pipeline: it shows whether a feature costs dollars or thousands of dollars a month, and which lever (a smaller model, a shorter prompt, a cap on output length) saves the most.
How it works
Cost per request = (input tokens × price in + output tokens × price out) / 1,000,000. Per day = per request × requests, per month = per day × days.
Real bills can differ: many providers discount cached prompt tokens and batch jobs, and charge extra for images or long contexts. Treat the result as an estimate. A rough way to count tokens in English: words × 1.3.
Topic 02
Mathematics
Three everyday maths tools: move the plane with a matrix, fit a trend line to your data, and see why a positive test is not always true.
Mx · Math
Linear map visualizer
Edit the matrix or pick a preset. The grid, the unit square and the letter F move with it.
- Determinant
- —
- Trace
- —
- Eigenvalues
- —
- Inverse
- —
The idea
A 2×2 matrix is a rule that moves every point of the plane. Its two columns say where the two unit arrows land, and everything else follows. Rotating, stretching, shearing and mirroring are all just matrices.
The determinant tells you how areas change, and the eigenvectors are the directions that only get stretched and never turned.
Try this
- Press Rotate 45°: the F turns and its area does not change (determinant 1).
- Press Mirror: the determinant becomes −1 and the F flips over.
- Press Squash: the determinant is 0, the plane collapses onto a line and there is no inverse.
In real life
Matrices are how computers move things: rotation and scaling in graphics and games, camera and robot transforms, and every layer of a neural network. Eigenvectors drive PCA, which compresses data and finds its main directions.
How it works
A matrix [[a, b], [c, d]] sends every point (x, y) to (a·x + b·y, c·x + d·y). Its columns are simply where the unit vectors î and ĵ land.
The determinant ad − bc is the factor by which every area grows. A negative determinant flips the plane over (the F turns into its mirror image); zero squashes everything onto a line, so there is no inverse. Eigenvectors (dashed) are the directions that are only stretched, by their eigenvalue λ.
Lr · Math
Linear regression: fit a trend line
Click the screen to add points. The line is the one that makes the squared vertical gaps as small as possible.
- Prediction y at that x
- —
- Slope
- —
- Intercept
- —
- R²
- —
- Verdict
- —
The idea
Linear regression draws the straight line that best follows your points. “Best” means the smallest total of the squared vertical gaps (the residuals, shown as thin lines).
R² says how much of the ups and downs of y the line explains: 1 means a perfect fit, 0 means the line is no better than the average.
Try this
- Press Add an outlier: one far-off point drags the whole line.
- Pick No link: R² falls close to 0, the line explains nothing.
- Pick Curve: R² can still look decent, yet the line misses the shape. Always look at the plot.
In real life
It is the first model to try on almost any numbers: forecasting sales or demand from a trend, calibrating a sensor against a reference, estimating a price from size. It is also the simplest neural network: one neuron with no activation, trained with gradient descent.
How it works
For points (xi, yi) the best line is slope = Σ(x−x̄)(y−ȳ) / Σ(x−x̄)² and intercept = ȳ − slope·x̄. R² = 1 − (sum of squared residuals) / (total variation of y).
Correlation is not causation, and a line fitted to a curve extrapolates badly. Beyond the data you have, predictions are guesses.
By · Math
Bayes: when a positive result is not proof
A test can be 95% accurate and still be wrong most of the time it says “positive”. See why in 1,000 people.
- A positive result is truly positive
- —
- Negative is truly negative
- —
- False alarms per real case
- —
The idea
A test has two kinds of mistakes: it can miss a real case (sensitivity below 100%) and it can raise a false alarm (specificity below 100%). What you care about is different: when the test says positive, how likely is it really true?
That depends on how common the thing is. When it is rare, the few false alarms among the many healthy cases outnumber the real cases.
Try this
- Start with Rare disease: 90% sensitive, 95% specific, and still only about 15% of positives are real.
- Drag How common it is up to 20%: the same test suddenly becomes trustworthy.
- Raise specificity from 95% to 99%: far more helpful than raising sensitivity.
In real life
It is the reasoning behind medical screening, fraud and security alerts, spam filters and any AI detector: a model that looks excellent on paper can still bury your team in false alarms when the thing it hunts is rare. Check this before you trust (or buy) a detector.
How it works
Out of 1,000 people, prevalence × 1,000 are truly positive. Of those, a share equal to the sensitivity test positive (true positives). Of the healthy ones, one minus the specificity test positive anyway (false positives).
Positive predictive value = true positives / (true positives + false positives). This is Bayes’ theorem written with counts instead of probabilities. It is the same ratio as precision in the F1 lab.
Topic 03
Physics
Throw it, swing it, build a wave out of sines. Classical mechanics you can poke at.
Pj · Physics
Projectile motion
Set speed, angle and height, then launch. Earlier shots stay on screen for comparison.
- Range
- —
- Max height
- —
- Flight time
- —
- Impact speed
- —
The idea
After launch, gravity pulls the ball down at a steady rate while it keeps its sideways speed. Together they draw a parabola. Air drag slows the ball and shortens the flight.
That is why the same throw goes about six times farther on the Moon, and why a light ping-pong ball barely follows the textbook curve.
Try this
- With No air from the ground, 45° gives the longest range.
- Switch the world to Moon and launch again: same speed, far longer flight.
- Choose Ping-pong on Earth: air cuts the range sharply. Raise the launch height and a flatter angle wins.
In real life
Sports analytics, game physics, drone payload drops and ballistics all start here. Comparing flights with and without air shows why simple formulas are only a first estimate and why engineers simulate step by step.
How it works
Without air the only force is gravity: x = v·cosθ·t, y = h + v·sinθ·t − ½·g·t². From the ground the range is longest at 45°; a raised launch prefers a flatter angle.
With air, drag adds an acceleration −k·|v|·v against the motion, where k = ρ·Cd·A / 2m for each ball, scaled by the world’s air density: none on the Moon, about 1.6% of Earth’s on Mars. The path is integrated numerically with the Runge–Kutta (RK4) method.
Pe · Physics
Pendulum lab
Change the length, the starting angle and the world. The trace below records the angle over time.
- Period, small-angle
- —
- Period, exact
- —
- Difference
- —
The idea
A pendulum swings back and forth in a time that depends on its length and on gravity, and not on how heavy the bob is. For small swings the period is T = 2π√(L/g).
For wide swings the real period is longer, and a gap opens between the textbook number and the exact one.
Try this
- Compare 0.25 m with 1 m: four times the length, double the period.
- Switch to Jupiter: stronger gravity, faster swing.
- Start at 170° and compare the two periods: the textbook formula is far off.
In real life
Pendulums are in clocks, metronomes, seismometers and cranes swinging a load. With a string and a stopwatch you can even measure local gravity. The gap between exact and small-angle periods shows when a textbook formula stops being good enough.
How it works
The textbook formula T = 2π√(L/g) assumes small swings. The exact period of an undamped pendulum is T = 2π√(L/g) / AGM(1, cos(θ0/2)), where AGM is the arithmetic–geometric mean. At 90° the swing is already 18% slower.
The animation solves θ″ = −(g/L)·sinθ − b·θ′ in small time steps. Notice that the mass of the bob appears nowhere.
Fo · Physics
Fourier synthesis
Add harmonics one by one and watch the sum close in on the target, overshoot and all.
The idea
Every repeating wave is a stack of simple sine waves: a base frequency plus higher harmonics of the right sizes. Add more sines and the sum looks more and more like the target wave.
This idea turns sound and images into numbers you can filter, compress and analyse.
Try this
- Start with 1 harmonic, then drag up to 20 and watch the square wave appear.
- Compare Square with Triangle: the triangle needs far fewer harmonics.
- Watch the overshoot beside each jump: it never goes away (the Gibbs phenomenon).
In real life
Any signal can be split into frequencies. That is how audio equalizers, MP3 and JPEG compression, noise filters, radio and vibration analysis work. Finding a hidden frequency (the FFT) is a standard first step on sensor and audio data.
How it works
Every periodic signal can be written as a sum of sines: a Fourier series. A square wave uses only odd harmonics with amplitudes 1, 1/3, 1/5…; a triangle’s harmonics shrink as 1/k², so it converges much faster.
Next to a jump the partial sums always overshoot by about 9% of the jump, however many terms you add: the Gibbs phenomenon. The same idea powers audio and image compression.
Topic 04
Space
Einstein’s relativity, hands on: time dilation, spacetime, and the clock correction your phone’s GPS makes every day.
Relativity
Td · Space
Time dilation and the twin paradox
Pick a speed and a destination. The light clocks show why the moving clock falls behind.
- Time on Earth
- —
- Time on the ship
- —
- Lorentz factor γ
- —
- That is
- —
The idea
Light always moves at the same speed for everyone. Take a light clock: a flash bouncing between two mirrors. On a moving ship, an observer on Earth sees the flash follow a longer zigzag path, yet at the same speed, so each tick takes longer. Moving clocks run slow.
Every process slows the same way: atoms, hearts, ageing. A twin who flies to a star near light speed returns younger than the twin who stayed home.
Try this
- Pick 99% and Proxima Centauri: 4.3 years pass on Earth, about 7 months on the ship.
- Pick 99.99% and the Galactic centre: 26,700 years on Earth shrink to about 380 on the ship.
- Watch the tick counters: the ship clock ticks γ times slower.
In real life
Not a daily calculator, but the effect is real and measured: particle accelerators and cosmic-ray muons depend on it, and it is half of why GPS needs correcting (see the Gravity and GPS lab). It is the honest answer to how much younger a fast traveller would be.
How it works
γ = 1 / √(1 − v²/c²). Earth time for the trip is distance / v; ship time is that divided by γ. The numbers cover cruising one way and ignore the time spent speeding up and slowing down.
This is measured every day: muons from cosmic rays survive the trip down through the atmosphere only because their clocks run slow, and particle accelerators see the same effect.
Mk · Space
Spacetime and simultaneity
Events A and B happen at the same moment for us. Change the observer’s speed and see if they still agree.
- Time of A for them
- —
- Time of B for them
- —
- Verdict
- —
A and B are 4 light-years apart and happen at time 0 for us. C is a later event inside the light cone of the origin, so every observer agrees it comes after the origin.
The idea
Draw time upwards and space sideways and you get spacetime. Light travels along the 45° lines; the yellow cone holds everything the origin can ever affect or be affected by.
A moving observer slices spacetime at a tilt: their “now” is a slanted line. Two events that are simultaneous for us are not simultaneous for them. Only cause and effect, inside the light cone, keep their order for everyone.
Try this
- Set the speed to 0: A and B happen together.
- Move right: B now happens first. Move left: A happens first.
- Notice the observer’s time axis and their ‘now’ line close in on the light line like scissors.
In real life
The ordering of events is why relativity is consistent: observers can disagree about which of two distant events came first, but never about cause and effect. Physicists use this diagram to check that no signal can outrun light.
How it works
Lorentz transformation: t′ = γ(t − v·x/c²), x′ = γ(x − v·t). Units here are years and light-years, so c = 1 and light lines are at 45°.
Events outside each other’s light cones cannot influence one another, which is why disagreeing about their order never breaks cause and effect.
Gt · Space
Gravity slows time: GPS and black holes
Two cases of general relativity: clocks in orbit around Earth, and a clock hovering near a black hole.
- Net clock drift
- —
- From weaker gravity
- —
- From orbital speed
- —
- Uncorrected, GPS would drift
- —
- Your clock runs at
- —
- Time far away
- —
At the far left of the slider you hover just outside the horizon, as close as the planet in “Interstellar”.
The idea
Einstein’s general relativity says gravity is curved spacetime, and clocks deeper in a gravity well tick slower. A clock on a mountain runs a little faster than one at sea level.
GPS satellites feel two opposite effects: weaker gravity speeds their clocks up, their orbital speed slows them down. The net is about +38 microseconds a day; without correcting it, your phone’s position would drift by about 11 km every day.
Try this
- Slide the altitude: low orbits (ISS) actually run slow, high ones run fast.
- Switch to Black hole and move towards the horizon: the near clock’s hand barely moves.
- Find the distance where one hour near the hole equals a day far away.
In real life
A real engineering fix: GPS satellites carry clocks pre-adjusted for relativity, otherwise position errors would grow by about 11 km every day. The same effect is used to compare precision clocks at different heights.
How it works
Gravity: Δf/f = GM/c² · (1/REarth − 1/r). Speed: −v²/2c² with v = √(GM/r) for a circular orbit. Near a non-spinning black hole, a hovering clock runs at √(1 − rs/r), where rs is the horizon radius.
Real GPS clocks are built to tick slightly slow on the ground (by 4.465 parts in 10¹⁰) so that in orbit they keep time with Earth.
Topic 05
Engineering
Resistors, LEDs, voltage dividers, batteries and filters: the quick sums of building real hardware.
Rc · Engineering
Resistor color code
Pick colours band by band, or type a value like 4.7k, 220 or 4k7.
- Resistance
- —
- Tolerance
- —
- Range
- —
The idea
Resistors are too small for printed numbers, so their value is written as coloured bands: two or three digits, then a multiplier, then the tolerance.
Read from the end where the bands are bunched together, and the gap-separated band is the tolerance.
Try this
- Type 4.7k: the bands are yellow, violet, red, gold.
- Switch to 5 bands for precision resistors that have three digit bands.
- Type 0.47 and see the silver multiplier for values below 1 Ω.
In real life
Anyone repairing or building electronics reads resistor bands all the time. Type a value to see which bands to look for, or pick the colours from a real part when the markings are hard to read.
How it works
Read from the end where the bands are bunched together. The first two bands (three on precision resistors) are digits, the next is a power-of-ten multiplier and the last, set apart, is the tolerance.
Black 0, brown 1, red 2, orange 3, yellow 4, green 5, blue 6, violet 7, grey 8, white 9; as multipliers gold means ×0.1 and silver ×0.01. Example: yellow, violet, red, gold = 47 × 10² Ω = 4.7 kΩ ± 5%.
Ω · Engineering
Ohm’s law and LED resistor
Type any two values, prefixes included. The LED helper picks a standard resistor for you.
Prefixes work: 4.7k, 20m (milli), 1M, 330u. The two fields you typed last are the inputs.
LED resistor
- Exact R
- —
- Use (E12)
- —
- Actual current
- —
- Resistor power
- —
The idea
Voltage pushes current through a resistance: V = I·R. Power is how much heat a part turns that into: P = V·I.
Any two of the four values fix the other two, so the tool solves whichever pair you type.
Try this
- Type 12 V and 4.7k: the current is about 2.55 mA.
- Fill in V and P instead and read the resistance.
- In the LED helper choose 5 V, Red, 20 mA: you need a 150 Ω resistor.
In real life
The most used formulas in electronics: sizing the resistor so an Arduino or ESP32 LED does not burn out, checking a resistor’s power rating, estimating the current drawn from a supply.
How it works
Ohm’s law V = I·R links voltage, current and resistance, and power is P = V·I = I²·R = V²/R. Any two values fix the other two.
An LED needs a series resistor to drop the extra voltage: R = (Vsupply − VLED) / I. Round up to the next standard E12 value so the current stays at or below the target, and pick a resistor rated for at least twice the power it will dissipate. LED forward voltages here are typical values; check your datasheet.
Vd · Engineering
Voltage divider
Enter the supply and both resistors, or ask for a target voltage and get a pair of standard values.
Leave the load empty if nothing is attached. Prefixes work: 4.7k, 100k, 1M.
- Output voltage
- —
- Divider current
- —
- Power in R1 / R2
- —
Find resistors for a target
- Best standard pair (E12)
- —
- You get
- —
- Error
- —
The idea
Two resistors in series split a voltage in proportion to their sizes: the middle point gives Vout = Vin · R2 / (R1 + R2). It is the simplest way to scale a voltage down.
A divider only works if the thing connected to the middle point draws very little current. A heavy load in parallel with R2 pulls the voltage down.
Try this
- Keep 1k and 2k on 5 V: the output is 3.33 V.
- Type a load of 1k: the output falls a lot. The divider cannot power things.
- Press 12 V to 3.3 V to see which standard pair reads a car battery safely.
In real life
Dividers are everywhere in hardware work: shrinking a 5 V or 12 V signal to the 3.3 V an ESP32 or Raspberry Pi pin can survive, reading a battery’s voltage with an analog pin, setting thresholds in comparators. Use big resistors (10k and up) so the divider wastes little current.
How it works
The current through the divider is Vin / (R1 + R2) and the voltage across R2 is that current times R2. With a load RL, R2 is replaced by R2 in parallel with RL.
The finder tries every pair of E12 values between 100 Ω and 1 MΩ and keeps the one whose output is closest to your target, preferring a total of about 20 kΩ. Real resistors have a tolerance of 1% to 5%, so the real output differs a little.
Bt · Engineering
Battery runtime
Enter the battery and the current the device draws when awake and asleep. Sleep time is what decides battery life.
- Battery lasts about
- —
- Average current
- —
- If always awake
- —
The idea
Battery capacity is measured in milliamp-hours: a 3,000 mAh battery can supply 3,000 mA for one hour, or 3 mA for a thousand hours. Runtime is capacity divided by the average current.
Most gadgets are not always on. If a device sleeps nearly all the time, its average current is close to the tiny sleep current, and battery life is stretched enormously.
Try this
- Slide Awake share from 1% to 0.1%: the runtime multiplies.
- Set Current asleep to 5000 µA (5 mA): a leaky sleep mode ruins everything.
- Try the coin cell: it cannot deliver 120 mA, so it only suits low-power chips.
In real life
It is the first sum when you design a battery sensor, a tracker, a wearable or an IoT node on an ESP32 or Arduino: it tells you whether you need a bigger battery, deep sleep or a solar panel. Always keep margin: cold, ageing and peak currents all cut real capacity.
How it works
Average current = awake current × awake share + sleep current × (1 − awake share). Runtime = capacity × usable fraction / average current.
The usable fraction covers the part of the capacity you cannot draw in practice (cut-off voltage, temperature, ageing). The awake-share slider is logarithmic, from 0.01% to 100%.
Fc · Engineering
RC filter and time constant
Pick a resistor and a capacitor. The curve shows which frequencies pass and which are cut.
- Cutoff frequency
- —
- Time constant τ
- —
- Rise time 10–90%
- —
The idea
A resistor and a capacitor together make a filter. A capacitor needs time to charge through the resistor, so slow changes get through and fast wiggles are flattened: a low-pass. Swap their roles and you get a high-pass.
The cutoff frequency f = 1 / (2πRC) is where the signal has fallen to 71% (−3 dB). The time constant τ = RC is how long the capacitor takes to reach 63% of its final voltage.
Try this
- Pick Button debounce: τ of about 1 ms swallows the contact chatter of a switch.
- Make the capacitor 10 times bigger: the cutoff drops 10 times.
- Switch to High-pass and see low frequencies get cut instead.
In real life
RC filters appear in nearly every circuit: debouncing buttons, smoothing the output of a PWM pin into a steady voltage, taking noise off a sensor or microphone signal, shaping audio and setting delays. Pick R and C so the cutoff sits between the signal you want and the noise you do not.
How it works
Gain of a first-order low-pass: |H| = 1 / √(1 + (f/fc)²), in decibels 20·log10|H|. For a high-pass it is (f/fc) / √(1 + (f/fc)²). Past the cutoff the filter falls 20 dB for every ten times in frequency.
Resistor and capacitor sliders are logarithmic (10 Ω to 1 MΩ, 10 pF to 100 µF). The plot always shows three decades on each side of the cutoff.
Need a tool like this for your team?
CurrentSky builds AI-powered products for businesses and founders, from AI agents and computer vision to drones and custom hardware. Bring a problem or an idea; we turn it into a real, working product.