Video summary
Pat Schloss from the University of Michigan presents a constructive critique of a complex figure published in *Nature* regarding new weight loss drugs, specifically focusing on panels I through T which initially appear identical. The study investigates how humanized mouse models can better replicate the effects of primate-specific GLP-1 receptor agonists compared to wild-type mice. Schloss outlines his four-step process for evaluation: description, analysis, interpretation, and judgment. He notes that while the paper involves numerous co-authors and a lengthy peer review period, the core issue lies in the organization of data visualization within Figure 1, which contains over a hundred lettered panels across five main figures.
During the analysis phase, Schloss observes that the twelve selected panels share consistent axes ranges and statistical markers but suffer from subtle alignment inconsistencies, such as varying panel widths and non-proportional time spacing on the x-axes. He points out that the authors use different colors for vehicle controls across different drug groups, which could confuse readers into thinking there are significant differences where none exist. Furthermore, he critiques the text descriptions that force readers to skip panels (e.g., comparing J, L, and N) rather than presenting them sequentially, arguing that this layout hinders direct visual comparison between control wild-type mice and humanized S33W mice under standard or high-fat diets.
In his judgment phase, Schloss concludes that while the individual plots are technically sound, the overall design misses a crucial opportunity to improve storytelling efficiency. He suggests that merging these twelve panels into a single, synthesized figure would allow authors to de-emphasize the similar vehicle lines by making them gray and highlight the distinct drug effects with color. This reorganization would not only save space but also use larger fonts for better readability. Ultimately, Schloss argues that by rearranging columns so that comparable groups are adjacent or merging redundant panels, authors can guide their audience's attention more effectively, ensuring the data story is told clearly without forcing readers to mentally bridge gaps between scattered visual elements.
Read the full video transcript
Hey folks, welcome back for another
episode of Code Club. In this episode, I
will be doing a constructive critique of
this set of panels, specifically panels
I through T of this figure that was
recently published in a paper in the
journal Nature. My name is Pat Schloss.
I'm a professor at the University of
Michigan and I am devoted to helping you
on your journey to make better
visualizations. This visualization, this
figure caught my eye because you betcha,
it has Z panels,
A through Z, 26 panels in this figure
alone. I certainly have thoughts on the
large number of panels that seem to be
populating modern biomedical research,
but what caught my eye was panels I
through T effectively look the same and
I wondered, do they need to be the same?
Do they need to be 12 different panels?
Could they be two or four? What could we
do? And so, I wanted to take on this set
of panels and put it through my process
of a constructive critique to see if
there's a way to improve the
visualization, to improve the story that
the authors are trying to say with these
panels. What is a constructive critique,
I hear you ask? Well, it is my four-step
process.
Starts with description, trying to get
the context that the figure and these
panels are being presented in. Second is
analysis where we take the panels and we
break them down into their constituent
parts and try to understand the story
that they are trying to tell, which then
gets us into interpretation, right? And
then we kind of think about whether our
interpretation of the plot matches what
the authors' interpretation of the plot
was. Finally, we come to the judgment
phase where I look at what I like, what
I didn't like, and perhaps make some
suggestions on things that could improve
the visualizations and improve the
authors' ability to tell the data story
that they are interested in telling. So,
as I already mentioned, this figure came
from a paper published in the journal
Nature. The article was published at the
beginning of May under an open access
model. So, if you want to get this
paper, you can for free. Go down below
in the description and you will see a
link there where you can find this
article.
The paper was titled a brain reward
circuit inhibited by next generation
weight loss drugs in mice. There are 35
co-authors. Uh that is a lot of people.
Um they're at the University of Virginia
in Charlottesville, Virginia and there
are three co-first authors. Elizabeth
Godshall, Taha Berga Gungor, Isabel
Sajonia. And I'm apologizing, of course,
for my mispronunciation. If we come down
to the bottom of the paper, we can get a
little bit more information about the
peer review history of this paper. It
was initially submitted December 12th of
2024 and was accepted March 24th of
2026. So, it's about a year and a half
that it spent in peer review, in
revision before it was finally
published, as I said, May 6th of 2026.
As the title suggests, this paper is
interested in the effect of these new
generation weight loss drugs, the GLP-1
agonists, which are all the craze these
days with things like Wegovy. They are
neurobiologists who are interested in
brain reward circuit inhibition, okay?
And so,
that's part of the story that we won't
get to uh as I'm looking at figure one
of this paper. In the abstract, we read
that glucagon-like peptide one receptor
agonists, GLP-1 RAs, effectively reduce
body weight and improve metabolic
outcomes, right? And so, this is part of
the craze for these drugs. It goes on to
say, however, established peptide-based
therapies require injections and are
complex to manufacture, right? So, I
think it's like a monthly injection or
something like that. So, that is a
problem. They go on to say, "Small
molecule GLP-1 RAs promise oral
bioavailability and scalable
manufacturing.
Right? So, this would be the advantage,
right? You can take a pill rather than
an injection. However, but their
selective binding to human versus rodent
receptors has limited mechanistic
studies. So, we want to understand the
role of these agonists better, we need
to put them into rodent models. But,
>> [laughter]
>> there's a difference between mice and
humans. And so, that is a problem that
is limiting the advancement of
developing of other small molecules that
would perhaps facilitate an oral route
of administration of these drugs. Going
on, they say, "Here we developed
humanized GLP-1R mouse models to
investigate how small molecule GLP-1RAs
influence feeding behavior." And so,
this is where I'm going to stop in the
abstract because this gets us to the
first figure where they are trying to
develop a mouse model that does a better
job of recapitulating what they see in
humans for these different drugs. This
difference in activity of the agonists
between humans and rodents, they mention
in the first sentences of the subsection
in the results section called humanized
GLP-1R
S33W mice.
>> [laughter]
>> Right? And so, basically what they found
or others have found is that there's a
single amino acid difference where
humans have a tryptophan and rodents
have a serine in position 33 of the
GLP-1R, the receptor. So, what they did
was that they inserted that mutation
using CRISPR-Cas mediated genome editing
to change that serine in mice to a
tryptophan in humans. They then go on to
tell us that looking at humanized mice
versus wild type littermates, they don't
see any physiological differences that
are obvious,
things like respiratory, uh energy
expenditure, or body weight. That they
the mice appear the same when they're
raised on the same conditions. This then
is where they jump into thinking about
the rest of figure one, which is what
we're going to be talking about in
today's critique. Before I jump into the
rest of figure one, let's look at all of
the figures together to understand where
this single figure fits in with all the
other figures. And I say single figure,
of course, as single figure as in as
figure one, but as already mentioned,
there are 26 different panels in figure
one. Wait for it though, because figure
two has 25 panels going through Y.
Figure three only
>> [laughter]
>> has 14 lettered panels, although we can
see with these microscopy images, there
really are quite a few more panels.
Figure four also has 25 panels going
through Y. And then figure five has 22
panels going through V. All told, the
five main figures of this paper, there
are 112 lettered figures lettered panels
that they constructed. I'm not going to
go through all 112. I am going to look
at 12, which are panels I through T of
this figure one. One other thing I like
to do as part of my critique is to
understand how the authors made their
visualizations. In this paper, there's a
section called data and statistical
analysis. And so what they tell us is
that statistical tests, including a
whole bunch of different tests, were
performed using R Studio, Python,
Jupyter Lab, MATLAB, GraphPad Prism, or
Microsoft Excel. I suspect the plots
were made in GraphPad Prism. Looking at
the panels, like there's kind of
telltale signs in there that these plots
were made in Prism rather than in R or
Python. Um one thing I'll point out is
that these version numbers that they put
next to R Studio are actually R version
numbers. So R has the
fairly traditional approach of a major
release number, a minor release number,
and then the patch. So like 4.1.2,
4.3.0.
It's odd to me
that they use two different versions of
R in this project. RStudio versions
generally have the year, period, and
then like I think like a number for the
month. So, like it might be 2006.05
as a version for RStudio. Uh I think
this sentence also highlights a problem
that we see in these types of
statements. They did not do their
analysis using RStudio. Sure, they did
it in RStudio, but they used R to do the
analysis. The distinction is a lot like
the difference between Python and
Jupyter Lab, right? Jupyter Lab is
perhaps a notebook, a Jupyter notebook,
where they were running Python, which
they could also run R in Jupyter.
But, RStudio is a development
environment to program in R primarily,
although I think you could also do
Python and other languages, right? And
so, RStudio isn't telling you
what they used to do the analysis.
They're telling you what the what the
software environment was rather than the
specific language, which was R. And
again, these version numbers relate to
R. As I mentioned, I am going to be
focusing on panels I through T in this
critique. And I'm focusing on these
panels because they look fairly similar
to each other. So, I'll talk about the
similarities and dissimilarities between
these 12 panels as I continue on in this
analysis phase of the critique. So, if I
look at panel I here as an exemplar of
the other panels, it's clearly in a
Cartesian coordinate space with
quantitative data on the x-axis and on
the y-axis. The x-axis presents the data
for time in hours. One thing I noticed
as I was looking at this is that the
time spacing is not proportional to the
actual difference in the amount of time,
right? So, the space between 1 and 2
hours is the same as the amount of space
between 2 and 4 hours. On the y-axis
then, they have the amount of food
consumed by the mice. One thing I
noticed is that all of the y-axes have
the same range, starting at zero and
going up to about 2.25.
And this is in units of grams. Panels I
through N are looking at the active
phase SD consumed. So, that's the active
phase for the mice and the standard diet
that the mice are consuming. Whereas O
through T are an inactive phase where
they're eating a high-fat diet and
looking at how much of the high-fat diet
they are consuming. As they mentioned in
the data and statistical analyses
paragraph, all data are presented as
mean plus or minus the standard error of
the mean unless otherwise noted. Um I
guess one other thing to point out from
the caption of this figure is that each
point represents between eight and 16
different animals, right? So, like this
point here or this point here or any of
the points that we see in these panels
is a summary of the eight to 16
different mice.
And so, there's a three different
geometries that are used to display the
data.
So, the first and most obvious is the
line going through the means. They also
use a filled circle to indicate the
points to indicate where the mean is,
again drawing the reader's attention to
that larger symbol. They also then show
a confidence interval to show plus or
minus one standard error of the mean
around the mean. So, they're using color
to indicate different treatment groups.
So, like in panels I, J, O, and P, they
have this dark blue to indicate the
vehicle, the control, whereas the blue,
the light blue
uh to indicate the drug. And similarly
in panels K, L, Q, and R, black is the
vehicle, red is the DAN group. That's a
shorthand for the drug name. And then
for M, N, S, and T, they have gray to
indicate the vehicle and this like
magenta color to indicate ORPHO, which
is another abbreviation for a drug name.
So, I'd say they're picking a more
neutral color to indicate the vehicle
and a different color to indicate the
different drugs. They also add
statistical information where you can
see the different numbers of stars above
the points indicating that that
comparison was statistically
significant. They tell us down below in
the caption that they used a one-way
ANOVA with Bonferroni correction for
panels F through Z, so that includes I
through T. And if there's a single star
like they have here in O, that means the
P value is less than .05. If there's two
stars like they have at two hours in the
same panel, that the P value is less
than .01. And if there's three stars,
that means the P value was less than
0.001.
In terms of an annotation layer, they do
have kind of a title for each of these
12 panels indicating if the mice were
wild type or whether they had that
CRISPR cast mutation changing the serine
for a tryptophan. So you can think of
the WT panels as being the mouse and the
S33W as
we can then see in terms of organization
that I, J, O, P go together with I and O
being the wild type, J and P being the
humanized version. And then as I
mentioned before, the top row I and J
being the standard diet and O and P
being the high-fat diet. And they take
kind of this tetrad of plots and repeat
it for the three different drugs. So now
moving into the interpretation phase of
the critique, I always like to interpret
the plot without relying on the authors
so that I'm perhaps not biased by what
their text says and see if we're drawing
out the same type of information. Think
the first thing I noticed is that across
the panels, the lines for the vehicle
data largely track with each other.
There's some subtle differences,
but by and large they're pretty similar
to each other, which tells me that there
wasn't really a difference in the amount
of food consumed across these different
drug regimens for the vehicle. And I
think they use slightly different
vehicles for different drugs. We also
don't see a difference in the vehicle
between the wild-type mice and the
humanized mice, which again is another
good control. Looking at the wild-type
mice,
again, that would be panels like I O K Q
M S, liraglutide is the only drug that
was administered where there was a
reduction in the amount of food consumed
by the mice. The Dan and the Orfo had no
statistical impact in wild-type mice on
food consumption. This result makes
sense because we're told up in panel A
that the liraglutide has specificity for
both rodents and primates, whereas the
Dan and the Orfo are specific to
primates and don't have activity in
mice. In the humanized mice, like again,
J P L R and N T, the mice receiving the
drug consumed less chow than the mice
receiving the vehicle. And that's really
regardless of whether it was the
standard diet or the high-fat diet. As I
already mentioned, I don't really see a
big difference between mice receiving
the standard diet here in this first row
versus those receiving the high-fat diet
here in the second row. So now let's see
if my interpretation aligns with the
authors' interpretation. The title of
the figure, as I already mentioned, is
descriptive, right? It's the validation
of small molecule GLP-1R agonist
responsive mouse model.
And so this makes sense, right? The data
are validating that they can take mice,
they can recreate the effect where these
primate-specific drugs have no effect in
mice, and when they then humanize the
mice, they then get the primate-like
effects. So I would say they have
validated this, right? If they had come
to me before they published this, I
would encourage them to come up with a
more declarative title, something like
humanized mouse model recapitulates the
primate-like activity in rodents,
something like that, right? These
results are being discussed in the
subsection humanized GLP-1R S33W mice.
These panels that I'm interested in are
described then in this second paragraph
that starts beyond their effects on
treating type 2 diabetes, GLP-1RAs
induce significant weight loss. Again,
there are 26 panels that they are now
going to summarize in this paragraph,
which is, I think, quite commendable.
So, the first description of the results
in these panels is in this highlighted
sentence. At these doses, all three
agonists markedly reduced active phase
SD consumption and rest phase HFD
feeding in GLP-1R S33W mice. In those
humanized mice, right? All three
agonists suppressed diet consumption
regardless of whether it's standard diet
or the high-fat diet. And again, that is
this second, fourth, and sixth column of
panels that they are describing here,
agreeing with my interpretation of the
results. Further down in this paragraph,
they go on to say, "As predicted by
species-specific receptor activation,
wild-type mice exhibited reduced SD and
HFD intake only with the liraglutide."
Again, that would be the first column of
plots in each of these tetrads, which
again agrees with what I had said about
their results. So, on the whole, my
interpretation agrees with their
interpretation. So, that is good. That
is the most important thing that we come
to the same conclusion about the data.
Now, let's turn to the judgment phase
where we can begin to think about how
efficient was it for me to come to the
same conclusion as them. When we think
about having 112 panels in the main body
of the paper and probably more than that
in the extended data, being able to
efficiently interpret plots is critical.
So, let's think about the positives. One
of the things I really appreciated was
that the axes on these 12 panels are the
same, right? The Y axis goes from 0 to
just over 2.
The X axis has the same three time
points, 1, 2, and 4. And so, that makes
it much easier to compare the lines or
the slopes of the lines across these
different panels. That seems like a
small thing, but that is something that
many authors don't do for their readers.
So, thank you, authors. So, let's now
talk about the negatives. Although the
axes look very similar to each other,
there are inconsistencies across these
12 panels that are distracting to the
audience and potentially make it harder
to interpret the data. So, what do I
mean by inconsistency? So,
let's start with the alignment of the
panels. So, here I have a horizontal red
line that I've created. So, let's move
this line up so it's right below the
tick for two in panel I.
And what we can see is that there are
other There are other panels here where
I believe the tick is above the red
line, as we see like in M and N. It's
not a huge difference, but again, it's
not entirely consistent. When I looked
at the alignment in the vertical
dimension, things appeared to line up
much better. Something I've noticed as I
look at this is that I believe panel O
and P might be different widths. And so,
if I
kind of go like this, that's 325 pixels.
And if I move it over to say here,
again, this blue dot was at the right
edge of the X axis in panel O, and you
can see it's further to the right of the
X axis in panel P. And if I kind of just
keep sliding like this,
we see kind of a similar type of effect
where perhaps O and R are the same
width, as is S and T, but P and Q are
more narrow than R, S, and T. I'll grant
that the alignment of the individual
panels and the length of the axes in
this case aren't huge, but I think it
really underscores the importance if you
are taking individual panels and
assembling them together, it's really
important to make sure that everything
is aligned, everything is the same width
and the same height. Otherwise, there is
a risk that people will misinterpret the
data, mainly because of how you have
scaled the individual plots. Some other
inconsistencies that I noticed was in
the color of the vehicle. And so, well,
within each of the tetrads, the vehicle
is the same color. I think it would have
been helpful to have the same color for
all of the vehicle lines. And so,
perhaps have black or blue or gray for
all 12 vehicle lines. But, having three
different makes you think that there's
really big differences in the vehicle or
there's perhaps differences that we
should be drawing between these sets of
tetrad plots. Perhaps an argument could
be made that different colors are
necessary because there were different
vehicles. If that's the case, then I
think it would be better to indicate
what the vehicle was. What was different
between the three different vehicles?
Another inconsistency I noticed was in
the legend for M that again, we have the
vehicle and we have the Orf O, but the
lines here, the gray line and this
magenta line, are a bit thicker than
what we saw in the other two legends for
this set of panels. Again, try to keep
everything the same width or people in
your audience are going to start
wondering like, is there something else
different about this set of panels than
the other eight panels? So, now let's
turn to the x-axis where I'm sure you
know what I'm going to say. The data on
the x-axis should be spaced
proportionate to the difference in the
values on the x-axis. So, the space
between two and four hours should be
twice as wide as the space between one
and two hours. I think if they did this,
that they would get these lines, like
let's look at these blue vehicle lines,
would actually look more straight and
wouldn't appear to be ramping up, kind
of like an exponential increase in diet
consumption. So again, because these
aren't on the correct spacing, it does
appear that there's an exponential
increase in food consumption unless you
go and you look at the x-axis and
realize that they're actually on a log
two scale, although I don't think the
authors intended to put things on a log
two scale. Some of these critiques you
might say, "Well, Pat, that seems a bit
petty. These aren't really fatal flaws
or big problems with interpretation of
the data. Sure, the spacing on the
x-axis is a problem, but eh, there's
only three points and
you know, that whether it's exponential
change or things like that, nobody
mentioned anything about that in the
paper, so who cares?"
Okay, let me tell you what I think is
the big problem with this set of 12
panels. It's that there's 12 panels.
>> [laughter]
>> Right, there's there's 26 panels. And so
when I look back at the results section,
we have this sentence
and we have this sentence, right? When
these panels are described,
you can see that they have Figure 1 J L
N. So if I want to compare those, I have
to jump over panels. So there's J, L,
and N. And it's actually visually
challenging to see if this line
is that different from this line or this
line. And similarly, for the colored
lines, right? Are those different from
each other? What would be far easier
would be to perhaps put those right next
to each other. If you want me to make
that comparison, put them next to each
other. And if you find yourself writing
a sentence and then putting a caption
like this, Figure 1 J L N, Figure 1 P R
T, again, that's
P R T, where you're skipping letters,
that's probably a sign that those panels
should be right next to each other.
Right? So, it should perhaps be IJK.
OPQ, right? Put those right next to each
other so that it's easier for your
audience to compare. In the second
sentence, they have the same type of
thing, right? IKM, OQS. They are
skipping over columns. There is no
sentence comparing I and J, K and L, M
and N, and so forth. There's also no
comparison of I and O or J and P. But,
the comparison they're calling out is
IKM,
JLN,
right? And similarly for the second row.
So, put those panels right next to each
other so that your audience can more
easily compare them. That might look
something more like this, where again,
we have IJK, LMN, OPQ, RST. It's the
same data, but where I've shuffled the
columns so that we're putting the
columns next to each other that we
actually want them to compare. The other
thing I would point out in the writing
of this paragraph is that kind of the
control, the wild type mouse, is
described as second. I would have
preferred to see this sentence before
this sentence, right? So, as predicted
by a species-specific receptor
activation that they showed in panel A,
wild type mice saw reduced SD and HFD
intake only with liraglutide. So, then
in the panel, in my refactored version
of the panel, I would then put the wild
type on the left and the humanized on
the right. Because again, those are the
comparisons you're making, and it is the
flow of the story. Why start with the
humanized data and then go to the wild
type as they did in the writing? In the
writing, if they flipped those
sentences, they could say, "Okay, this
is what we expected to see. This is like
a control. If we give mice the
rodent-specific drug, it's going to
work. But, if we give mice the
primate-specific drug, it's not going to
work, right? That's what we see here on
the left. That's like the control. Now,
the cool new thing we did was that we
humanized these mice. And those are the
panels on the right here, this S33W.
Organization of the panels really helps
to tell the story more easily. Now, I'm
going to take it up a notch and suggest
that these three panels, IJK, LMN, OPQ,
RST, that they could actually be pulled
together.
Because the vehicles
from this perspective don't look that
different from each other. And so,
perhaps what we could do is we could
make the vehicle lines gray and try to
put them into the background. And then
we'd have the three drugs, and those
could be colored. And then we could see
again, how the vehicle lines fall on top
of each other, and then how the drug
lines compare to each other. To make it
easier to compare the three different
drugs, again, within the same wild type
or humanized mice experiments. So,
that's what I've done in this version of
the plot. It's the same data as in these
12 panels, but again, synthesized
together. So, we have the wild type and
the humanized mice left and right, the
standard diet on top, the high-fat diet
on the bottom.
And when this is scaled to the same
height as those 12 panels, the two rows
that they take up, we can actually see
this font in like an eight or 10-point
font versus the font that's here, that's
probably more on the order of five or
six-point font. One thing I would love
to have been able to do is to perhaps
combine the wild type and the S33W
columns, perhaps giving the wild type a
different symbol than the humanized, but
when I've played around with that, it
gets a little bit jumbled and just too
busy of a visualization. But, we can
have these four facets, we can have the
legend off on the right, we can write
out the full name of the drugs, we can
still indicate the significance by
putting colored stars below those points
that were significantly different with
the color corresponding to the drug, and
it actually uses less space, and we can
use a larger font to tell the same
story. And again, what I just want to
highlight is that the authors knew this
was the right thing to do, but they
didn't do it. And they knew it was the
right thing to do because when they
wrote these citations to their data,
they were skipping panels, right? And
so, if you find yourself when needing to
skip panel letters when you're trying to
ask your audience to compare those
different panels,
think, should I be putting those panels
next to each other so my audience can
compare them, or should I even be
merging those panels together so that
they can more directly make the
comparisons that I want them to see? So
again, individually, these panels are
not horrible. They're fine just as they
are. Some subtle things that I would
change about them, but on the whole, and
thinking about the storytelling of the
data, I think they really missed an
opportunity to make it easier on their
audience to see the differences that the
authors wanted the audience to see. If
you're interested in seeing me refactor
that visualization into this version,
on Wednesday morning at 9:00 a.m., I
will be running a live stream, and if
you can't be there in time, there will
be a recording posted as well. But I
would really love to show you my thought
process of how I would take those 12
panels and convert them into this single
figure, this maybe one lettered panel,
to simplify the data and improve the
interpretation of the data. All right.
Well, I hope you have gotten a lot out
of this discussion of thinking about how
we can organize our data visualizations
to improve the ability to tell our data
stories.
Please tell your friends what we are
doing here, and I will see you next time
for another episode of Code Club.