Your app graphs tonnage because it's easy, not because it predicts gains.
Every tracker defaults to sets times reps times load because it's arithmetic. The research counts hard sets. The metric your log stores decides the one it can ever show.
The chart that went up is the one you trust least
The volume chart in your tracker climbed twelve percent this block. Your working weight on the main lifts sat exactly where it was in May. Both numbers were pulled from the same log, on the same sets, and only one of them is telling you whether the block did anything. The chart that went up is the one you trust the least once you know what it is made of.
Tonnage is a receipt, not a training signal
Tonnage is a receipt. Sets times reps times load, summed over a session or a week, the single number nearly every tracker graphs as your progress line. It is honest arithmetic. It rises when you add a rep, add a set, add a kilo, or bolt a high-rep pump finisher onto the end of a session. That is exactly the problem. A number that climbs for four unrelated reasons cannot tell you which reason moved it, and only one of those reasons has much to do with growing muscle. A rising tonnage line is a record of iron that moved, not evidence that you delivered a stimulus.
The metric your app graphs is the one it can compute
Trackers default to volume-load for one reason: it is computable from data the app already holds. Reps, load, and set count are the fields every logger stores, so the product graphs the metric that falls out of them for free. The trouble is what that metric flattens. A back-off set of fifteen reps at sixty percent inflates tonnage far more than a grinding top single, even though the single is closer to the edge that drives adaptation. Two lifters can post an identical weekly tonnage while one took every set to RIR 1 and the other left four reps in the tank on all of them. The metric cannot separate them. It was chosen for the ease of adding it up. Ranking training quality was never in the job description.
What does training volume actually measure?
Look at where the volume research actually points and the unit is not the kilogram-total. Greg Nuckols' 'A New Approach to Training Volume' at Stronger by Science makes the case directly: count hard sets, not volume-load, because hard sets track the dose that drives hypertrophy while tonnage smears it across rep ranges. Brad Schoenfeld's dose-response meta-analysis on volume and hypertrophy counts weekly sets per muscle group, not tonnage. Renaissance Periodization's volume landmarks, MEV, MAV, and MRV, are all counted in sets. Examine's volume writeups land in the same place. The convergence across those sources is hard to miss: the dose unit is the hard set, roughly a set taken within a few reps of failure, and tonnage is the thing they are correcting away from.
Tonnage
Hard sets
What it sums
sets × reps × load
sets taken close to failure, per muscle per week
What the log must store
numbers it already has
set-level RPE or RIR
Moves when you
add reps, load, or light back-offs
add a set that is actually hard
Backed as a hypertrophy dose
confounded
SBS, Schoenfeld dose-response, RP landmarks
Why the two metrics need different things from your log. Sources: Nuckols (Stronger by Science), Schoenfeld et al.
What counts as a hard set
A hard set is one taken close enough to failure to register as a growth stimulus, commonly bracketed at RIR 0 to about 4. A set left at RIR 6 is a warm-up wearing a working set's clothes, and tonnage counts it identically to a set grinding at RIR 0.
Two sessions, and tonnage ranks the wrong one
Run the two sessions and watch tonnage rank them wrong. Session one: five sets of ten on bench at 100 kg, every set left at RIR 3. Tonnage, 5,000 kg. Session two: five sets of five at 140 kg pushed to RIR 1. Tonnage, 3,500 kg. On the chart, session one wins by a mile. In the muscle, session two delivered five genuinely hard sets and session one delivered a rehearsal that never got close to failure. Now add a two-by-fifteen pump finisher at 60 kg to session two and its tonnage climbs to 5,300 kg, nudging past session one for reasons that have nothing to do with why it was the better session. The number can be made to say anything. The hard-set count says the same thing either way: session two trained, session one rehearsed.
Tonnage can rank your hardest session below your easiest one.
You can't recover a metric your log never stored
Here is the part that decides what your tracker can ever tell you. Hard sets are computable only if the log stored the effort of each set while it happened. If the app kept a session tonnage and threw the per-set detail away, no query run later can reconstruct how close to failure any of those sets got, because that information was never written down. To count hard sets per muscle per week, or to track where you sit against your MRV, the database has to hold raw per-set rows: load, reps, RPE or RIR, the lift, the muscle it maps to, the date. A totals-only store has already lost the substrate. The better metric is not a feature you can add later; it is a data-model decision made the first time a set gets saved.
An app that stores totals can never show the better number
This is the line between a receipt and a substrate, and it is a build decision more than a marketing one. Platepusher stores each set as a raw row: the load, the reps, the RPE you logged, the lift, the date. Because the effort data is on disk and not summed away, hard-set counts and per-muscle weekly volume are computable from what you already logged, alongside the tonnage line if you still want it. The math runs on the substrate the log kept. An app that saved only the total can graph the arithmetic forever and never surface the number the research actually cares about, because it discarded the raw material on the way in.
What we're watching next
The open question is not which metric wins on paper; hard sets already have the evidence. It is whether set-level RPE, logged in the moment between gasps, is accurate enough to trust as the input. Effort ratings drift, lifters round to RPE 8 out of habit, and the last rep always feels harder than it was. The metric is only as honest as the substrate feeding it. That is the real frontier: not tonnage versus hard sets, which is settled, but whether the effort data the better metric runs on can be logged honestly enough to trust.
Log the raw set, not just the total. Get Platepusher and keep the substrate the better volume metric runs on.
Serious lifters have watched the volume-load chart climb through stalled blocks for years. Platepusher counts sets the way the research does: raw per-set rows first, tonnage as one view among several, hard-set and per-muscle volume computable because the effort data was never summed away.