Submind YouTube summaries
Thumbnail for How to align your text with your data visualizations (CC424)

How to align your text with your data visualizations (CC424)

Watch on YouTube

Video summary

Pat Schloss from the University of Michigan presents a constructive critique of a complex figure published in *Nature* regarding new weight loss drugs, specifically focusing on panels I through T which initially appear identical. The study investigates how humanized mouse models can better replicate the effects of primate-specific GLP-1 receptor agonists compared to wild-type mice. Schloss outlines his four-step process for evaluation: description, analysis, interpretation, and judgment. He notes that while the paper involves numerous co-authors and a lengthy peer review period, the core issue lies in the organization of data visualization within Figure 1, which contains over a hundred lettered panels across five main figures. During the analysis phase, Schloss observes that the twelve selected panels share consistent axes ranges and statistical markers but suffer from subtle alignment inconsistencies, such as varying panel widths and non-proportional time spacing on the x-axes. He points out that the authors use different colors for vehicle controls across different drug groups, which could confuse readers into thinking there are significant differences where none exist. Furthermore, he critiques the text descriptions that force readers to skip panels (e.g., comparing J, L, and N) rather than presenting them sequentially, arguing that this layout hinders direct visual comparison between control wild-type mice and humanized S33W mice under standard or high-fat diets. In his judgment phase, Schloss concludes that while the individual plots are technically sound, the overall design misses a crucial opportunity to improve storytelling efficiency. He suggests that merging these twelve panels into a single, synthesized figure would allow authors to de-emphasize the similar vehicle lines by making them gray and highlight the distinct drug effects with color. This reorganization would not only save space but also use larger fonts for better readability. Ultimately, Schloss argues that by rearranging columns so that comparable groups are adjacent or merging redundant panels, authors can guide their audience's attention more effectively, ensuring the data story is told clearly without forcing readers to mentally bridge gaps between scattered visual elements.
Read the full video transcript
Hey folks, welcome back for another episode of Code Club. In this episode, I will be doing a constructive critique of this set of panels, specifically panels I through T of this figure that was recently published in a paper in the journal Nature. My name is Pat Schloss. I'm a professor at the University of Michigan and I am devoted to helping you on your journey to make better visualizations. This visualization, this figure caught my eye because you betcha, it has Z panels, A through Z, 26 panels in this figure alone. I certainly have thoughts on the large number of panels that seem to be populating modern biomedical research, but what caught my eye was panels I through T effectively look the same and I wondered, do they need to be the same? Do they need to be 12 different panels? Could they be two or four? What could we do? And so, I wanted to take on this set of panels and put it through my process of a constructive critique to see if there's a way to improve the visualization, to improve the story that the authors are trying to say with these panels. What is a constructive critique, I hear you ask? Well, it is my four-step process. Starts with description, trying to get the context that the figure and these panels are being presented in. Second is analysis where we take the panels and we break them down into their constituent parts and try to understand the story that they are trying to tell, which then gets us into interpretation, right? And then we kind of think about whether our interpretation of the plot matches what the authors' interpretation of the plot was. Finally, we come to the judgment phase where I look at what I like, what I didn't like, and perhaps make some suggestions on things that could improve the visualizations and improve the authors' ability to tell the data story that they are interested in telling. So, as I already mentioned, this figure came from a paper published in the journal Nature. The article was published at the beginning of May under an open access model. So, if you want to get this paper, you can for free. Go down below in the description and you will see a link there where you can find this article. The paper was titled a brain reward circuit inhibited by next generation weight loss drugs in mice. There are 35 co-authors. Uh that is a lot of people. Um they're at the University of Virginia in Charlottesville, Virginia and there are three co-first authors. Elizabeth Godshall, Taha Berga Gungor, Isabel Sajonia. And I'm apologizing, of course, for my mispronunciation. If we come down to the bottom of the paper, we can get a little bit more information about the peer review history of this paper. It was initially submitted December 12th of 2024 and was accepted March 24th of 2026. So, it's about a year and a half that it spent in peer review, in revision before it was finally published, as I said, May 6th of 2026. As the title suggests, this paper is interested in the effect of these new generation weight loss drugs, the GLP-1 agonists, which are all the craze these days with things like Wegovy. They are neurobiologists who are interested in brain reward circuit inhibition, okay? And so, that's part of the story that we won't get to uh as I'm looking at figure one of this paper. In the abstract, we read that glucagon-like peptide one receptor agonists, GLP-1 RAs, effectively reduce body weight and improve metabolic outcomes, right? And so, this is part of the craze for these drugs. It goes on to say, however, established peptide-based therapies require injections and are complex to manufacture, right? So, I think it's like a monthly injection or something like that. So, that is a problem. They go on to say, "Small molecule GLP-1 RAs promise oral bioavailability and scalable manufacturing. Right? So, this would be the advantage, right? You can take a pill rather than an injection. However, but their selective binding to human versus rodent receptors has limited mechanistic studies. So, we want to understand the role of these agonists better, we need to put them into rodent models. But, >> [laughter] >> there's a difference between mice and humans. And so, that is a problem that is limiting the advancement of developing of other small molecules that would perhaps facilitate an oral route of administration of these drugs. Going on, they say, "Here we developed humanized GLP-1R mouse models to investigate how small molecule GLP-1RAs influence feeding behavior." And so, this is where I'm going to stop in the abstract because this gets us to the first figure where they are trying to develop a mouse model that does a better job of recapitulating what they see in humans for these different drugs. This difference in activity of the agonists between humans and rodents, they mention in the first sentences of the subsection in the results section called humanized GLP-1R S33W mice. >> [laughter] >> Right? And so, basically what they found or others have found is that there's a single amino acid difference where humans have a tryptophan and rodents have a serine in position 33 of the GLP-1R, the receptor. So, what they did was that they inserted that mutation using CRISPR-Cas mediated genome editing to change that serine in mice to a tryptophan in humans. They then go on to tell us that looking at humanized mice versus wild type littermates, they don't see any physiological differences that are obvious, things like respiratory, uh energy expenditure, or body weight. That they the mice appear the same when they're raised on the same conditions. This then is where they jump into thinking about the rest of figure one, which is what we're going to be talking about in today's critique. Before I jump into the rest of figure one, let's look at all of the figures together to understand where this single figure fits in with all the other figures. And I say single figure, of course, as single figure as in as figure one, but as already mentioned, there are 26 different panels in figure one. Wait for it though, because figure two has 25 panels going through Y. Figure three only >> [laughter] >> has 14 lettered panels, although we can see with these microscopy images, there really are quite a few more panels. Figure four also has 25 panels going through Y. And then figure five has 22 panels going through V. All told, the five main figures of this paper, there are 112 lettered figures lettered panels that they constructed. I'm not going to go through all 112. I am going to look at 12, which are panels I through T of this figure one. One other thing I like to do as part of my critique is to understand how the authors made their visualizations. In this paper, there's a section called data and statistical analysis. And so what they tell us is that statistical tests, including a whole bunch of different tests, were performed using R Studio, Python, Jupyter Lab, MATLAB, GraphPad Prism, or Microsoft Excel. I suspect the plots were made in GraphPad Prism. Looking at the panels, like there's kind of telltale signs in there that these plots were made in Prism rather than in R or Python. Um one thing I'll point out is that these version numbers that they put next to R Studio are actually R version numbers. So R has the fairly traditional approach of a major release number, a minor release number, and then the patch. So like 4.1.2, 4.3.0. It's odd to me that they use two different versions of R in this project. RStudio versions generally have the year, period, and then like I think like a number for the month. So, like it might be 2006.05 as a version for RStudio. Uh I think this sentence also highlights a problem that we see in these types of statements. They did not do their analysis using RStudio. Sure, they did it in RStudio, but they used R to do the analysis. The distinction is a lot like the difference between Python and Jupyter Lab, right? Jupyter Lab is perhaps a notebook, a Jupyter notebook, where they were running Python, which they could also run R in Jupyter. But, RStudio is a development environment to program in R primarily, although I think you could also do Python and other languages, right? And so, RStudio isn't telling you what they used to do the analysis. They're telling you what the what the software environment was rather than the specific language, which was R. And again, these version numbers relate to R. As I mentioned, I am going to be focusing on panels I through T in this critique. And I'm focusing on these panels because they look fairly similar to each other. So, I'll talk about the similarities and dissimilarities between these 12 panels as I continue on in this analysis phase of the critique. So, if I look at panel I here as an exemplar of the other panels, it's clearly in a Cartesian coordinate space with quantitative data on the x-axis and on the y-axis. The x-axis presents the data for time in hours. One thing I noticed as I was looking at this is that the time spacing is not proportional to the actual difference in the amount of time, right? So, the space between 1 and 2 hours is the same as the amount of space between 2 and 4 hours. On the y-axis then, they have the amount of food consumed by the mice. One thing I noticed is that all of the y-axes have the same range, starting at zero and going up to about 2.25. And this is in units of grams. Panels I through N are looking at the active phase SD consumed. So, that's the active phase for the mice and the standard diet that the mice are consuming. Whereas O through T are an inactive phase where they're eating a high-fat diet and looking at how much of the high-fat diet they are consuming. As they mentioned in the data and statistical analyses paragraph, all data are presented as mean plus or minus the standard error of the mean unless otherwise noted. Um I guess one other thing to point out from the caption of this figure is that each point represents between eight and 16 different animals, right? So, like this point here or this point here or any of the points that we see in these panels is a summary of the eight to 16 different mice. And so, there's a three different geometries that are used to display the data. So, the first and most obvious is the line going through the means. They also use a filled circle to indicate the points to indicate where the mean is, again drawing the reader's attention to that larger symbol. They also then show a confidence interval to show plus or minus one standard error of the mean around the mean. So, they're using color to indicate different treatment groups. So, like in panels I, J, O, and P, they have this dark blue to indicate the vehicle, the control, whereas the blue, the light blue uh to indicate the drug. And similarly in panels K, L, Q, and R, black is the vehicle, red is the DAN group. That's a shorthand for the drug name. And then for M, N, S, and T, they have gray to indicate the vehicle and this like magenta color to indicate ORPHO, which is another abbreviation for a drug name. So, I'd say they're picking a more neutral color to indicate the vehicle and a different color to indicate the different drugs. They also add statistical information where you can see the different numbers of stars above the points indicating that that comparison was statistically significant. They tell us down below in the caption that they used a one-way ANOVA with Bonferroni correction for panels F through Z, so that includes I through T. And if there's a single star like they have here in O, that means the P value is less than .05. If there's two stars like they have at two hours in the same panel, that the P value is less than .01. And if there's three stars, that means the P value was less than 0.001. In terms of an annotation layer, they do have kind of a title for each of these 12 panels indicating if the mice were wild type or whether they had that CRISPR cast mutation changing the serine for a tryptophan. So you can think of the WT panels as being the mouse and the S33W as we can then see in terms of organization that I, J, O, P go together with I and O being the wild type, J and P being the humanized version. And then as I mentioned before, the top row I and J being the standard diet and O and P being the high-fat diet. And they take kind of this tetrad of plots and repeat it for the three different drugs. So now moving into the interpretation phase of the critique, I always like to interpret the plot without relying on the authors so that I'm perhaps not biased by what their text says and see if we're drawing out the same type of information. Think the first thing I noticed is that across the panels, the lines for the vehicle data largely track with each other. There's some subtle differences, but by and large they're pretty similar to each other, which tells me that there wasn't really a difference in the amount of food consumed across these different drug regimens for the vehicle. And I think they use slightly different vehicles for different drugs. We also don't see a difference in the vehicle between the wild-type mice and the humanized mice, which again is another good control. Looking at the wild-type mice, again, that would be panels like I O K Q M S, liraglutide is the only drug that was administered where there was a reduction in the amount of food consumed by the mice. The Dan and the Orfo had no statistical impact in wild-type mice on food consumption. This result makes sense because we're told up in panel A that the liraglutide has specificity for both rodents and primates, whereas the Dan and the Orfo are specific to primates and don't have activity in mice. In the humanized mice, like again, J P L R and N T, the mice receiving the drug consumed less chow than the mice receiving the vehicle. And that's really regardless of whether it was the standard diet or the high-fat diet. As I already mentioned, I don't really see a big difference between mice receiving the standard diet here in this first row versus those receiving the high-fat diet here in the second row. So now let's see if my interpretation aligns with the authors' interpretation. The title of the figure, as I already mentioned, is descriptive, right? It's the validation of small molecule GLP-1R agonist responsive mouse model. And so this makes sense, right? The data are validating that they can take mice, they can recreate the effect where these primate-specific drugs have no effect in mice, and when they then humanize the mice, they then get the primate-like effects. So I would say they have validated this, right? If they had come to me before they published this, I would encourage them to come up with a more declarative title, something like humanized mouse model recapitulates the primate-like activity in rodents, something like that, right? These results are being discussed in the subsection humanized GLP-1R S33W mice. These panels that I'm interested in are described then in this second paragraph that starts beyond their effects on treating type 2 diabetes, GLP-1RAs induce significant weight loss. Again, there are 26 panels that they are now going to summarize in this paragraph, which is, I think, quite commendable. So, the first description of the results in these panels is in this highlighted sentence. At these doses, all three agonists markedly reduced active phase SD consumption and rest phase HFD feeding in GLP-1R S33W mice. In those humanized mice, right? All three agonists suppressed diet consumption regardless of whether it's standard diet or the high-fat diet. And again, that is this second, fourth, and sixth column of panels that they are describing here, agreeing with my interpretation of the results. Further down in this paragraph, they go on to say, "As predicted by species-specific receptor activation, wild-type mice exhibited reduced SD and HFD intake only with the liraglutide." Again, that would be the first column of plots in each of these tetrads, which again agrees with what I had said about their results. So, on the whole, my interpretation agrees with their interpretation. So, that is good. That is the most important thing that we come to the same conclusion about the data. Now, let's turn to the judgment phase where we can begin to think about how efficient was it for me to come to the same conclusion as them. When we think about having 112 panels in the main body of the paper and probably more than that in the extended data, being able to efficiently interpret plots is critical. So, let's think about the positives. One of the things I really appreciated was that the axes on these 12 panels are the same, right? The Y axis goes from 0 to just over 2. The X axis has the same three time points, 1, 2, and 4. And so, that makes it much easier to compare the lines or the slopes of the lines across these different panels. That seems like a small thing, but that is something that many authors don't do for their readers. So, thank you, authors. So, let's now talk about the negatives. Although the axes look very similar to each other, there are inconsistencies across these 12 panels that are distracting to the audience and potentially make it harder to interpret the data. So, what do I mean by inconsistency? So, let's start with the alignment of the panels. So, here I have a horizontal red line that I've created. So, let's move this line up so it's right below the tick for two in panel I. And what we can see is that there are other There are other panels here where I believe the tick is above the red line, as we see like in M and N. It's not a huge difference, but again, it's not entirely consistent. When I looked at the alignment in the vertical dimension, things appeared to line up much better. Something I've noticed as I look at this is that I believe panel O and P might be different widths. And so, if I kind of go like this, that's 325 pixels. And if I move it over to say here, again, this blue dot was at the right edge of the X axis in panel O, and you can see it's further to the right of the X axis in panel P. And if I kind of just keep sliding like this, we see kind of a similar type of effect where perhaps O and R are the same width, as is S and T, but P and Q are more narrow than R, S, and T. I'll grant that the alignment of the individual panels and the length of the axes in this case aren't huge, but I think it really underscores the importance if you are taking individual panels and assembling them together, it's really important to make sure that everything is aligned, everything is the same width and the same height. Otherwise, there is a risk that people will misinterpret the data, mainly because of how you have scaled the individual plots. Some other inconsistencies that I noticed was in the color of the vehicle. And so, well, within each of the tetrads, the vehicle is the same color. I think it would have been helpful to have the same color for all of the vehicle lines. And so, perhaps have black or blue or gray for all 12 vehicle lines. But, having three different makes you think that there's really big differences in the vehicle or there's perhaps differences that we should be drawing between these sets of tetrad plots. Perhaps an argument could be made that different colors are necessary because there were different vehicles. If that's the case, then I think it would be better to indicate what the vehicle was. What was different between the three different vehicles? Another inconsistency I noticed was in the legend for M that again, we have the vehicle and we have the Orf O, but the lines here, the gray line and this magenta line, are a bit thicker than what we saw in the other two legends for this set of panels. Again, try to keep everything the same width or people in your audience are going to start wondering like, is there something else different about this set of panels than the other eight panels? So, now let's turn to the x-axis where I'm sure you know what I'm going to say. The data on the x-axis should be spaced proportionate to the difference in the values on the x-axis. So, the space between two and four hours should be twice as wide as the space between one and two hours. I think if they did this, that they would get these lines, like let's look at these blue vehicle lines, would actually look more straight and wouldn't appear to be ramping up, kind of like an exponential increase in diet consumption. So again, because these aren't on the correct spacing, it does appear that there's an exponential increase in food consumption unless you go and you look at the x-axis and realize that they're actually on a log two scale, although I don't think the authors intended to put things on a log two scale. Some of these critiques you might say, "Well, Pat, that seems a bit petty. These aren't really fatal flaws or big problems with interpretation of the data. Sure, the spacing on the x-axis is a problem, but eh, there's only three points and you know, that whether it's exponential change or things like that, nobody mentioned anything about that in the paper, so who cares?" Okay, let me tell you what I think is the big problem with this set of 12 panels. It's that there's 12 panels. >> [laughter] >> Right, there's there's 26 panels. And so when I look back at the results section, we have this sentence and we have this sentence, right? When these panels are described, you can see that they have Figure 1 J L N. So if I want to compare those, I have to jump over panels. So there's J, L, and N. And it's actually visually challenging to see if this line is that different from this line or this line. And similarly, for the colored lines, right? Are those different from each other? What would be far easier would be to perhaps put those right next to each other. If you want me to make that comparison, put them next to each other. And if you find yourself writing a sentence and then putting a caption like this, Figure 1 J L N, Figure 1 P R T, again, that's P R T, where you're skipping letters, that's probably a sign that those panels should be right next to each other. Right? So, it should perhaps be IJK. OPQ, right? Put those right next to each other so that it's easier for your audience to compare. In the second sentence, they have the same type of thing, right? IKM, OQS. They are skipping over columns. There is no sentence comparing I and J, K and L, M and N, and so forth. There's also no comparison of I and O or J and P. But, the comparison they're calling out is IKM, JLN, right? And similarly for the second row. So, put those panels right next to each other so that your audience can more easily compare them. That might look something more like this, where again, we have IJK, LMN, OPQ, RST. It's the same data, but where I've shuffled the columns so that we're putting the columns next to each other that we actually want them to compare. The other thing I would point out in the writing of this paragraph is that kind of the control, the wild type mouse, is described as second. I would have preferred to see this sentence before this sentence, right? So, as predicted by a species-specific receptor activation that they showed in panel A, wild type mice saw reduced SD and HFD intake only with liraglutide. So, then in the panel, in my refactored version of the panel, I would then put the wild type on the left and the humanized on the right. Because again, those are the comparisons you're making, and it is the flow of the story. Why start with the humanized data and then go to the wild type as they did in the writing? In the writing, if they flipped those sentences, they could say, "Okay, this is what we expected to see. This is like a control. If we give mice the rodent-specific drug, it's going to work. But, if we give mice the primate-specific drug, it's not going to work, right? That's what we see here on the left. That's like the control. Now, the cool new thing we did was that we humanized these mice. And those are the panels on the right here, this S33W. Organization of the panels really helps to tell the story more easily. Now, I'm going to take it up a notch and suggest that these three panels, IJK, LMN, OPQ, RST, that they could actually be pulled together. Because the vehicles from this perspective don't look that different from each other. And so, perhaps what we could do is we could make the vehicle lines gray and try to put them into the background. And then we'd have the three drugs, and those could be colored. And then we could see again, how the vehicle lines fall on top of each other, and then how the drug lines compare to each other. To make it easier to compare the three different drugs, again, within the same wild type or humanized mice experiments. So, that's what I've done in this version of the plot. It's the same data as in these 12 panels, but again, synthesized together. So, we have the wild type and the humanized mice left and right, the standard diet on top, the high-fat diet on the bottom. And when this is scaled to the same height as those 12 panels, the two rows that they take up, we can actually see this font in like an eight or 10-point font versus the font that's here, that's probably more on the order of five or six-point font. One thing I would love to have been able to do is to perhaps combine the wild type and the S33W columns, perhaps giving the wild type a different symbol than the humanized, but when I've played around with that, it gets a little bit jumbled and just too busy of a visualization. But, we can have these four facets, we can have the legend off on the right, we can write out the full name of the drugs, we can still indicate the significance by putting colored stars below those points that were significantly different with the color corresponding to the drug, and it actually uses less space, and we can use a larger font to tell the same story. And again, what I just want to highlight is that the authors knew this was the right thing to do, but they didn't do it. And they knew it was the right thing to do because when they wrote these citations to their data, they were skipping panels, right? And so, if you find yourself when needing to skip panel letters when you're trying to ask your audience to compare those different panels, think, should I be putting those panels next to each other so my audience can compare them, or should I even be merging those panels together so that they can more directly make the comparisons that I want them to see? So again, individually, these panels are not horrible. They're fine just as they are. Some subtle things that I would change about them, but on the whole, and thinking about the storytelling of the data, I think they really missed an opportunity to make it easier on their audience to see the differences that the authors wanted the audience to see. If you're interested in seeing me refactor that visualization into this version, on Wednesday morning at 9:00 a.m., I will be running a live stream, and if you can't be there in time, there will be a recording posted as well. But I would really love to show you my thought process of how I would take those 12 panels and convert them into this single figure, this maybe one lettered panel, to simplify the data and improve the interpretation of the data. All right. Well, I hope you have gotten a lot out of this discussion of thinking about how we can organize our data visualizations to improve the ability to tell our data stories. Please tell your friends what we are doing here, and I will see you next time for another episode of Code Club.