Submind YouTube summaries
Thumbnail for Critique of a faceted heatmap ordered with a dendrogram (CC422)

Critique of a faceted heatmap ordered with a dendrogram (CC422)

Watch on YouTube

Video summary

The video presents a constructive critique of a complex figure from a high-impact paper published in *Nature Cancer*, which investigates how stromal cell reprogramming drives tissue disorganization in nodal B-cell lymphoma. The speaker outlines a four-step evaluation process: description, analysis, interpretation, and judgment. In the description phase, the presenter provides context on the study's focus on comparing healthy reactive lymph nodes against two types of lymphoma—follicular lymphoma (FL) and diffuse large B-cell lymphoma (DLBCL)—using single-cell RNA sequencing and multiplex immunofluorescence data. The critique notes that while the paper is open access and well-cited, the figure's title is overly descriptive rather than declarative, failing to clearly state the main takeaway for the reader. Moving into the analysis phase, the speaker examines how the visualizations were constructed, noting discrepancies between the published code on GitHub and the final rendered figures, suggesting that manual adjustments were made outside of the original script. The critique focuses heavily on panels E and F, which utilize faceted heatmaps and stacked bar plots to display cell enrichment and disease proportions. The presenter points out technical inconsistencies, such as mismatched color palettes in the source code versus the final output, and questions the statistical transparency regarding how dendrograms were generated and what criteria were used to cluster cell subpopulations into distinct facets. In the interpretation and judgment phases, the speaker argues that the visualizations suffer from issues related to sample size imbalance and poor design choices that hinder clarity. Specifically, the presence of five patients in the FL group versus only four in the other groups creates a potential bias that is not adequately addressed by scaling the bars to equal width. The presenter strongly advocates replacing the stacked bar plot with a jitter plot to better reveal individual patient variation and suggests using alternative visualization methods like biplots to explain the factors driving cluster separation. Additionally, the critique highlights minor but significant usability issues, such as perpendicular axis labels, inconsistent legend styles, and unnecessary abbreviations that confuse readers unfamiliar with specific immunological jargon. Ultimately, the video concludes by balancing praise for the authors' effort to simplify complex biological data with a call for greater transparency and improved design standards. The speaker commends the use of faceted heatmaps for organizing diverse cell populations but urges future researchers to provide clearer methodological details, such as distance metrics for clustering and baselines for statistical calculations. By offering specific suggestions like standardizing color schemes, removing axis expansion tails, and rewriting abbreviated terms for clarity, the presenter aims to help both the original authors and the broader scientific community create more interpretable and honest data visualizations that accurately reflect the underlying biological realities.
Read the full video transcript
Hey folks, welcome back for another episode of Code Club. In this episode, I will be doing a constructive critique of this set of panels that was recently published in the journal Nature Cancer. When I do my critiques, I try to be as constructive and positive as possible by following a four-step process. The first is description, where we try to get the context of the figure and trying to understand the point and the kind of utility of the figure in the overall work that is being presented. The second is analysis, to try to break down our figure or set of panels into its constituent elements to understand how this figure was assembled. The third is interpretation, trying to understand what the figure is saying and what the authors want me to take away from visualizing their data in the way that they've presented it. And finally, the fourth step is judgment, where we try to render judgment on what worked well, what didn't work so well, what I'd like to see improved or perhaps improve the interpretability, or just some kind of ideas that I have for things that I would like to try with their data. Let's head into description, where again we try to get the context of the figure in the overall paper. At a very broad level, the figure that we're talking about was published in a paper in the journal Nature Cancer. I haven't talked about Nature Cancer in these critique videos, but it's a very high-impact journal, at least from my perspective, with a journal impact factor of 28.5, which tells me that papers in this journal tend to get cited a lot. The title of the paper was reprogramming of stroma-derived chemokine networks drives the loss of tissue organization in nodal B cell lymphoma. I was drawn to this for a couple of reasons. First, I was intrigued by the figure and how I would make it. And also, about six or seven years ago, I had lymphoma. Not the type of lymphoma that they're talking about here. Don't worry, I'm doing fine. Um, but whenever I see lymphoma, uh, I jump at the opportunity to study it because it's relevant to me, right? And it's something that I care about and my family cares about, of course. Also, this paper, as you'll see here in the red, was published open access. And so, down below in the description to this critique, I will be presenting a link so that you can go get the article without having to pay anything. There is a long list of authors here, so many that they don't show them all to you. I believe I counted 39 co-authors. There are four senior corresponding authors. The first three authors, I'm going to butcher your names, I'm sorry. Felix uh, Srdan Novakovic, Anna Medaoki, Lia Jobava, are three co-corresponding authors. The work was done in Germany, primarily at the University of Heidelberg in Heidelberg, Germany. And then some of the other authors are at the University Hospital in Heidelberg, Germany, as well. Something I always love to look at is the peer review history of this paper. We can see that it was submitted, it was received June 28th of 2024, and it was accepted February 10th of 2026. So, again, that took about 20 months to go through peer review and be accepted before finally being published at the end of March 25th, 2026. So, again, from submission to publication, I would count that at 21 months. Almost 2 years to get published. Kudos to these authors for their dogged determination to see this work all the way through and have it published in Nature Cancer. Again, this work is talking about nodal B-cell lymphoma and the loss of tissue organization in lymph nodes. Reading through the abstract, we get a bit more context on the science here. Lymph node function requires the organization of cells into higher order spatial units. However, the principles governing lymph node architecture in healthy and disease remain poorly understood. Therefore, here we use single cell and spatial mapping to investigate the mechanisms directing immune cell organization in human lymph nodes and its disruption in architecturally distinct lymphoma entities, indolent follicular lymphoma and aggressive diffuse large B cell lymphoma, DLBCL. They go on to say that our data substantiate the central role of lymph node resident stromal cells and chemokine driven lymphocyte zonation and reveal an inflammatory feedback loop fueled by tumor reactive T cells that trigger stromal remodeling, progressive loss of homeostatic chemokine gradients, and tissue disorganization from non-malignant state to FL, follicular lymphoma, and DLBCL, which again is that diffuse large B cell lymphoma, the more aggressive form of lymphoma. That's getting a little bit beyond what we're going to be looking at with this figure 1 E and F. So, let's go ahead down to the results section where they introduce figure 1 to get more context about the figure that we'll be looking at. So, the very first section of the results section starts it out mapping the spatio-cellular principles of lymphoma induced lymph node remodeling. And so, what they tell us in figure 1A right here is the layout of their study. And that they were interested in looking at non-malignant reactive lymph nodes, so the RLNs, with a canonical organization into B cell rich follicles and T cell zones. Malignant lymph nodes from patients with FL, so the follicular lymphoma, characterized by a follicular growth pattern of malignant B cells, and then lymph nodes from patients with DLBCLs. Immunologists love their acronyms, love their jargon. So, they've got three populations, right? They've got non-malignant reactive lymph nodes that have the typical expected organization of B cells and T cells to make lymph nodes. And then they also looked at patients with FL, a form of lymphoma where you see this B cell remodeling, but not as aggressive as the DLBCL. So, what they did in figure 1A then was layout what the study looked like, the patients that they had. Don't get me started on the circular format. Um but then they did some single-cell RNA-seq work, and then they also did MIF or multiplex immunofluorescence. And they generated these ordination diagrams using UMAP on large numbers of cells, uh about 30,000 in the first case and about 3.7 million in the MIF, the multiplex immunofluorescence. And it's actually this multiplex immunofluorescence what's presented in E and F. To investigate spatial lymph node organization, is they obtained a 56-plex multiplex immunofluorescence data set comprising 3.7 million cells from four individuals with reactive lymph nodes. Again, that's the the more healthy state. Five with follicular lymphoma, and then four from people with diffuse large B cell lymphoma. There was some overlap between the single-cell group and this MIF group. Although, they don't really seem to be talking much about the overlap of those two different data sets. What they tell us then is that they identified 21 hematopoietic and structure-defining non-hematopoietic cell types or states, uh including all these. Um the latter included a distinct subset of B cell zone-residing FDCs defined by CD21 and CXCL13 positivity. I'm not going to go into that because I don't understand it myself. But I think basically what they were doing is they identified these different regions. And so then when we come down to E and F, we see those same regions down here. Okay? So basically they went from this ordination data from the multiplex immunofluorescence to look at the structure and they give you some examples of what that structure looks like down here where they and they have the six different neighborhoods right here in these three examples from the RLN, the reactive lymph node, the follicular lymphoma, and the DLBCL. And so then again, you can see like this creamish color I believe is the T zone, which again comes back from up here, which I believe is probably this purple. I It would be nice to have overlapping colors and overlapping jargon, but don't get me started. As I already mentioned, they talked about the BEC rich T zone, which is probably this BEC here, the T zone, the lymphatic sinus, plasma cells, follicles, and the B prol neighborhood. So again, that's a bit of the context, the description of where the data in E and F are situated in the overall study. Again, looking at the spatial organization and the factors, the cell types that give rise to this transition from kind of canonical, typical structure to disease within the follicular lymphoma and the large diffuse B cell lymphoma. Overall, this paper has five main figures. When I counted up the individual panels, they're about 50 individual panels. As we can see in the one we're looking at, like panel B here, it's kind of like two panels. The figure one, I would say this is a descriptive title, right? It's just saying lymphoma induced remodeling of the lymph node cellular ecosystem. I think that is descriptive. It's not telling us what they want us to take away. They're not telling us what the remodeling looked like, right? They're telling us that here's data related to the remodeling. Okay, but but what do you want me to see about the remodeling? I don't know, right? It's not in the title. The other figures though have more declarative titles like dysregulation of chemokine networks and FRCs is linked to a fused growth pattern in lymphoma, right? Or a reprogramming of lymph node stromal cells from a functionally tissue organized state to an inflammatory state underlies the loss of tissue organization in lymphoma, right? These are declarative. They are telling me what they want me to take away. Where again, in figure one, that's much more descriptive. It's describing the type of data in the figure rather than telling me what they want me to take away from these different panels. Something I always like to do before digging into the figure is trying to understand how they made the figures. So, if I do a search for Prism, nothing comes up. If I look for R, I see that some of the analysis was done in R um and that they also have some citations to various R packages and analysis that you can do in R and methods about how to do things in R. I notice there's also a code availability section where they have put their custom code on GitHub. If we go to that repository, we can see that they have R markdown documents for each of the figures. And figure one here has an R markdown document stepping us through how they made their code, right? And a lot of this is of course using R because it's R markdown. Something I see that they use was complex heat map, which is a heat map generating R package that does things more sophisticated than uh like a normal heat map like you might get out of ggplot2. When I was looking through this code, something I I that didn't quite align with the actual rendered figures. Down here they're using different shades of green, orange, and purple. And this is in figure F left, which again down here this is F left. They're using purple, blue, and green, not purple, orange, and green. So that's a little bit odd. Also, the number of colors, 8, 10, and 10, doesn't quite match what they had in the final paper. And so this is probably code that they developed along the way to generate the analysis and it was probably modified as they went along. So I don't think this is the actual code that they used to generate the figures that wound up in the paper. Also in here there's comments like for C saying it's Fiji output, no need for code. And so it appears that these panels were made in different parts and then assembled together probably in something like Illustrator or PowerPoint. So we get a little bit of an insight into how they made these panels with in this figure largely using R. I did see some mention of Python in some of their work, but it's largely made in R and then probably compiled somewhere else with some probable other modification along the way. So continuing on the analysis phase of this critique, let's now look at panels E and F and let's start with E. E is a heat map, right? And it's got dendrograms showing how the different rows and columns relate to each other. The coordinate system I would say is a discrete coordinate system. Sure we could say it's Cartesian, but it's got discrete values on the X axis, the different subpopulations of cells, and then on the Y axis it has the different categories being the different zones, neighborhoods, cell types, or whatever that they described up above in panel D or in panel B on the right, right? Where they used that UMAP ordination of the MIF data. The values on the X and Y axes are being sorted by the clustering of the dendrogram. And I would say that the dendrogram is a statistical layer. So, the fill color within each of the cells of the heat map here on the left is an indication of the enrichment, and we see its legend over here in the bottom right of panel E, where it says enrichment log 2 OR. That's the odds ratio. So, it's taking a log 2 of the odds ratio. So, what zero would mean here would be one, and what four here would mean would be two to the four, which is 16, or two to the negative four, which would be 1 over 16. So, again, the dark red are more enriched, and the dark blue are highly depleted. I don't know that I would necessarily call this an annotation, but it's perhaps a sense of organization in the data, is that they've taken these cellular subpopulations and broken them apart into five different groups, or five different facets, I would think of them as. And again, that is being defined by the clustering across the columns of data. I think that's interesting to have that type of organization within the visualization, to say that these five subpopulations go together, this one's on its own, these four go together, and so forth. All right, so that was 1E. Let's now turn to 1F, and here there are two parts to 1F. There's a left, and there's a right. For F on the left, we'll start there. On the X axis, they have the relative abundance, or the proportion of cells in this particular zone or neighborhood, and for this particular patient and disease entity, okay? And so, each rectangle, like this rectangle here, corresponds to one patient, who in this case had DLBCL, and the fraction of their cells that showed up in the B prol neighborhood, okay? So, it's maybe a lot to wrap your head around there, but that's kind of how it breaks down. The Y axis, again, is the same as what we saw in E, which are these six different clusters of overall zones or regions or cell types that their ordination analysis allowed them to break things down into. Again, the position on the Y axis is being dictated by the clustering in the dendrogram shown on the left side of 1 E. The fill color then, again, corresponds to the three different disease groups, whether it's a purple, a blue, or a green. And that different shades of purple or different shades of blue or different shades of green corresponds to the different patients that were representatives of those different disease entity groups. In terms of the geometry or the type of plot, this is a stacked bar plot laid on its side. And so, within, say the ggplot2 framework, you could generate this using something like geom_col. Of course, you would have to flip your traditional X and Y axis coordinates to get them to lay down on their side. I don't see any statistical layers or annotation layers like we did in panel 1 E. So, let's move on to the right side of panel F, where here they have another heat map. And again, heat maps can be generated in ggplot2 using the geom_tile function, where again, you give it discrete values on the X and Y axis and a continuous variable, or I guess you could do discrete, but in this case continuous variable as the fill color. And so, here on the X axis, again, is discrete, we have the three disease entities, right? So, the reactive lymph node, the follicular lymphoma, and the diffuse large B cell lymphoma. And the fill color is telling us what percentage of the cells within that neighborhood, which again is defined on the Y axis, as it was for the left side of 1 F and in 1 E, proportion of cells that could be assigned to that disease entity. And so, basically what we're seeing with this very pale color, uh where my cursor is, is that there weren't many cells at all from the reactive lymph node, right? And so we can see that over here in purple that very few of the cells are represented in these purple rectangles and perhaps a similar amount by the blue rectangles, but that most of them are in these green rectangles. And so that's why the B pearl neighborhood is dark red under DLBCL and quite faint for these others. That basically this heat map on the right as a summation of say the purple the blue and the green going to the three different columns in this right hand heat map. I would concede that in some ways this right hand heat map is a statistical summary of the data on the left side. They're basically summing the percentage of cells across the patients within a disease entity and then representing that as a different shade of red in this right hand heat map. So now we've broken things down into their constituent parts and try to put them back together again. What we'd like to do now is to interpret the data. And so I always like to interpret the data without reading what the paper wants me to take away from it. In E, my takeaway is that we have these six different general populations of cells or regions within a lymph node. And what they're trying to do is they're saying like these were defined by our MIF data, right? By their immunofluorescence data. And they want to better understand like what went into defining these clusters of MIF data, right? And so they're breaking it down then by these different cell populations that their assay, their multiplex immunofluorescence, was able to pull out. And so trying to say like what defines the plasma cell group? Oh, it's plasma cells, right? This PC. And it's also these to some degree, right? But it's certainly not these four cell types because those are quite blue. And also, again, the the B prol is pretty definitive of the pre prol neighborhood, right? And so, I think what they're trying to do here is to lend some information to understand the sub populations of cells that are going into making these six different categories. Then on the right side in F, what they're trying to do is then to the link that to the disease entities by saying, "What of these six different categories define the three different disease entities, right?" So, again, the DLBCL is quite abundant in the B prol neighborhood. And so, if you have a lot of cells in the B prol neighborhood, then it's likely that you've got something from a DLBCL patient. And that if you've got something in the plasma cells or the follicles, that that's going to be more indicative of the follicular lymphoma, right? And and so forth. And you know, in some places it gets a little bit more muddled uh as to what is being defined. I think some of the other things that I take away, especially in panel F on the left side stack bar plot, is that there's a lot of variation in the data, right? So, looking at like the plasma cells, you know, I had just said that that's kind of indicative of the FL state, but there's really only one patient that has a large number of plasma cells. Uh and so, is that indicative of FL or not? I don't know, right? And so, again, there's a fair amount of variation in the data among patients within a disease entity. So, in their results section, they go on to describe this MIF data. So, we're looking at, again, panels E and F, where they say, "As expected, follicular neighborhoods were expanded in both size and number in FL." And I believe that's this data down here, looking at the follicles, that it is more common in patients with FL. They go on to say that moreover, significantly enlarged plasma cell neighborhoods and enrichment of BECs in the T cell zone were observed in FL lymph nodes. They talk about those plasma cells being more pronounced in FL. Again, that's that's one subject. I don't know that I would go for that. They then go on to say that these BEC-rich T zone cells were more pronounced or expanded in the FL. And I don't know that I totally see that. Again, there are five patients in the blue where there's four in the green and the purple. And so, it might be that that looks expanded because there's five rather than four. And so, if we had four, maybe it wouldn't seem so expanded. Maybe it would just seem kind of consistent across all three disease entities. They go on to say that moreover DLBCL lymph nodes displayed a near complete depletion of lymphatic vessels and an expansion of non-FDC FRCs pointing toward a potential role for of stromal cell remodeling in the structural reorganization of diffusely growing lymphomas. And again, this I think points to the cases up here where like the T zone lymphatic sinus, plasma cells, and follicles are much more present among individuals with DLBCL relative to the other disease entities. So, I want to be a little bit guarded in saying that my interpretation doesn't or does agree with theirs just because again, this is way out of my wheelhouse, this type of analysis and this type of biology that they're interested in. But I do have some questions that I would ask the authors about their analysis and kind of already mentioned one of those being kind of this idea that there's five FL patients versus four for the others. And this large variation. Can we really say that plasma cells are expanded in people with FL relative to others when there's only one patient that had any appreciable level of FLs in them. Now, let's move on to the judgment phase of the critique. What did I like about this? Let's always start with the positives. So, I really was intrigued by this faceting of their heat map. Again, I was aside from kind of my own personal interest in lymphoma when I saw this heat map, I was really intrigued by how would I do this in R? And so that really got my attention that they have these fascinating of these different cell sub populations into the five different groups that are being defined by the dendrogram clustering those different columns. I thought that was pretty cool. Something else I really liked was in panel F that they use different shades of the same color to indicate replication of those different disease entities, right? So they use different shades of purple or blue or green. I thought it was pretty cool. Um and because I don't care necessarily about who the sample is from or which replicate number it is, I don't really need to worry a whole lot about was this the same shade of blue across these different um you know, the the the six different groupings or not. I just need to know it's blue or it's a different shade of blue. And and perhaps you could go ahead and be like, yeah, like this blue and this blue are the same and that and that's the same, right? But at the end of the day, it's blue. And so I know that comes from a patient with FL. And I think that's really what is most important in interpreting panel F. So those are some positive things that I could take away from these panels to help me think about my own future data visualization efforts and things that I might want to incorporate. As a general negative, I am left wondering why this is two panels instead of one panel or perhaps even three panels, right? That the left and the right really do go together and it's the very much the same data. I mean, they share a Y axis across all three of these different plots. So why isn't that one panel? I don't know. Or why isn't it three panels, right? That we've got this other heat map off here on the right. And perhaps it is, as I mentioned, that this is kind of a summarization of the data in the left side of F. So maybe that's why it's two. Anyway, this is something I've been thinking about lately is why do we label panels as different panels? When do we label them as different panels or when do we collect different plots together into the same panel letter? I don't know. It's If you've got thoughts on how to letter different panels or how to pull different panels apart, uh let me know because I would I would love to see different people's logic for why they make uh multiple figures under the same letter or different letters for different panels, right? And so this is something I wrestle with with my own stuff and as I've been looking through more and more figures with many panels, it's something that just strikes my curiosity. So let's look at E. And E, I feel like is an alternative approach to something that we often do in microbial ecology or ecology in general, which is called a biplot. Where you have an ordination and you then try to explain what are the factors that are pulling clusters into different parts of that ordination space. So if we look at this MIF ordination, which is again where the data that we're looking at in E come from, what are the cell types that are pulling NK up here or what's pulling PC out here or what's pulling BEC down here? That with a biplot, they generally would draw a vector coming down with effectively these different cell populations. That you'd have a different vector for each of these cell population and they would be centered in the middle here and they would be going out to the different cell populations describing how much that subpopulation is driving the separation of these different populations within this ordination. I don't I don't know that I totally buy this approach to doing it, but I think it's an interesting alternative to a biplot. Where in this case at least, you'd have like 21 different vectors. And perhaps some of those vectors would just get removed because they're not that important. But um again, it's an interesting alternative and I'll have to think more about what I think about this. One thing that I'm already questioning is with that biplot, the length of the vector tells you how important it is in pulling that population out of the center of the ordination. And so, if you have a very short arrow, it's not that important. And so, oftentimes the short arrows we might remove. But if it's a very long arrow, we know that that arrow is representing some factor that's really pulling that cluster out. In this case, I don't know how important these 20 or whatever different subpopulations are at an individual level in defining these six different categories. And so, I feel like that's something perhaps that's missing. There might be something I'm I am missing in the interpretation of this, but it would be nice to have some measure indicating how important these different factors are in discriminating between these six different groups. Something else I wonder about with panel E is is this odds ratio, and how was the odds ratio calculated? I suspect it was looking at all the cells within a group and then saying, "How much more enriched is this?" So, say for example, the granulo and the LEC are more enriched across all of the columns within this row. I think, right? And so, that you'll see that rows always have something red and something blue. And if you kind of average across all of them, it gets to zero, right? And so, you know, you've got these very dark red colors and you have these very dark blue colors balancing things out. And so, perhaps that goes back to my question of like, "How important are these columns?" Or you might look at something like T prol or T reg, but they're very muted colors and perhaps aren't really important in distinguishing between these six different groups. I don't know. So, it would have been nice to have some indication of how the odds ratio was calculated. So, if I'm right about how they're calculating the odds ratio, I wonder what this would look like if they calculated the odds ratio relative to what they were seeing in the RLN, the the reactive lymph nodes, which are not, as I understand it, disease. They're They're not cancerous, right? And so, what would it look like to take all the data and to scale it, again, in panel E, according to the RLN data as a baseline, instead of taking each row as its own baseline and looking at kind of the average, um what would it look like to use that reactive lymph node as the baseline? Also, there wasn't any description of how the dendrograms were drawn. I could see that they used the R package complex heatmap. My understanding is that complex heatmap is using Hclust under the hood to drive this clustering, and the default for Hclust is a complete linkage or furthest neighbor clustering algorithm. Again, it would be nice to say that. Um also, how are the distances calculated? That's not indicated, and again, I'll assume that they were Euclidean distances. Also, related to these dendrograms and the heatmap, is that it's not clear what decision criteria was used to pull apart the subpopulations into the different facets. And so, while I like the faceting, I think that's cool, I would like to know how did they pull that apart? Was there a certain level of dissimilarity between the different clusters that they used say, "Okay, this is one group, this is another group, this is another group." Uh that's missing, right? Or did they just say, "Biologically, these four go together. Look, they cluster together. We're going to make them a cluster." and so forth. That type of transparency, I think, would be helpful to understanding the visualization. So, I always like to read left to right, uh cuz that's my culture, and so I always kind of complain when labels, like we see on the x-axis here, are written perpendicular to the axis. I would love for these to be written horizontally, and so one strategy I always have is to pivot the whole thing 90°. But, of course, we have long names on the y-axis, and so that wouldn't work. And so, I think this worked out pretty well. One thing that kind of drives me nuts is that they are abbreviating things that I don't think need to be abbreviated. So, like PC, why not write plasma cells like you have here? If they wrote out plasma cell here, it wouldn't take up any extra space than any of the other sub-populations that they already have on the x-axis. And I suspect the same could be true for many of these, right? Uh I know immunologists love their jargon and their abbreviation, but I really do think it gets in the way. And it would be really nice to write some of these out so that when you're looking at something like PC, you know that's the same thing as plasma cells, and so it's not surprising that that is so red. Within the names of these six different groups, I know there is some kind of physiology going in here where they're looking at the spatial structure, but the fact that they have like zone, so like the BEC rich T zone, the T zone, and that they also have like the B prol neighborhood, why isn't this a zone? Is this like a different word for the sake of having a different word, or is there something else going on? Again, this might be something in the analysis or the biology, I don't know, but why not use zone here or use neighborhood up here? I think they use neighborhood a lot throughout the paper, so maybe use neighborhood, but neighborhood is a longer word, so maybe zone would be better. I don't know. I don't understand. Um it would be nice to be consistent though for people like me who get easily confused by these things. One final thing I'll mention about panel E is this legend enrichment log 2 OR. I really think could go over here um and be more closely connected with the data versus kind of sitting in the middle here where it's perhaps not immediately clear what side it goes with or if it's both, but saying oh yeah, that goes over here uh because the colors match, right? Well, why not just move the legend over here? You've got a nice big gap here. Let's put it there cuz that seems like a like a logical place for it. All right, so that's panel E. Now, let's turn to panel F. One thing I'll start out with panel F is that there's no title on the x-axis. And so, you have to go down through the caption which is rather lengthy and in the PDF version of the figure and there's no caption. It's the entire figure takes up the entire page. So is the caption down here on the next page? No, it's back up here. And then I'm looking through here. F left bar plot illustrating the proportions of identified neighborhoods across patient samples and disease entities. It'd be far preferable to give something to your audience to know what this x axis represents. I get that it's a proportion but proportion of what, right? Go ahead and say that here. You've already got it down in the caption. Maybe come up with a more pithy statement of saying that but put it right here. Again, there's plenty of room. Also, there's making more room would be taking these numbers on the x axis and making them parallel to the x axis. There's no need for these values to be perpendicular. It actually looks kind of weird for them to be perpendicular. Sure everything else is perpendicular but generally when you have a continuous x axis, those values are parallel to the x axis as a convention. So go ahead and flip those. That'll get you more room here to then put in your title. Also thinking about the axis, I see that again they probably used ggplot2 which naturally adds this expansion factor to the x axis and so I think we're seeing that here on the left and right where we've got the tails of the x axis sticking out beyond the data. I'd go ahead and remove that expansion, expand equals false, whatever to go ahead and remove those tails so that the axis starts at zero and ends at one. As I've already mentioned a few times, the FL category, the disease entity has five patients represented whereas the others only have four and so I think it is clouding the interpretation of the data. I don't know how much this extra patient is biasing my interpretation of the data. If all the data were pooled at a equal patient level, then I would expect the blue category to be wider than the other categories all other things being equal. But, if they pull based on disease category, then I would know that okay, there's five, but it's been scaled so that the three disease entities are equal. And so, maybe to that, it would be interesting to know what panel F on the left would look like across all six of the categories. Again, they're taking these six categories and pulling them apart. We don't totally get a sense of how many cells fall into each of these six different categories. And that would be that would be informative. But again, pulling all six together to give me a sense of is it a third, a third, a third, or is it like a 13th, a 13th, a 13th, a 13th, one per patient. And so, it would be good to see that and to see how evenly distributed the data were across the 13 patients and the three different disease groups. Otherwise, I think they are biasing the data towards overemphasizing the FL data. The last thing I'll say about panel F on the left here is that I do not like stacked bar plots. I'm pretty well on the record with that. I would I would prefer to see is what would this look like as a jitter plot, right? And so, we could take the 13 different patients and represent them each as a colored point, perhaps using the same colors, and then we could jitter the points within each of the six different categories using the same x-axis. And that way then it'd be easier to see how much the purple you know, how much variation there is in the purple relative to the blue, relative to the green, and how different the overall proportion is of the blue to the green to the purple, right? Uh as it is here, it's much more difficult to interpret. I also think that this will help a little bit with that problem with the FL having five versus four patients. Finally, on the right, we have this extra heat map. Not totally sure I see what the point of this is since it is summing up these other bars that we see on the left side of panel F. Regardless, it's here and if we're going to use it, let's make the best of it. One suggestion I would have would be to use a different color. They have white to red, which matches the white to red that they have with the enrichment going from zero to four, here going from zero to 80. Again, use a different color because you're representing different type of data. The other thing I notice about this legend for panel F on the right is that it has white tick marks and a white border or no border at all, whereas the legend for the heat map for E has black tick marks and a black border. It's a little subtle thing, but it'd be nice to have consistency between the two. So, either have white tick marks and border on both or a black tick marks and black border on the others. I kind of like having the black border and the black tick marks. It kind of matches having the black border within the panel F on the left legend as well. Well, that is my critique of panel E and F today. I hope you've gotten something out of this and thinking about again how we represent complex data using these types of visualizations. I know that these data sets are really complicated and so I commend the authors for their efforts in trying to simplify things. At the same time, hopefully my suggestions are helpful to them if this ever gets back to them or to you as you're thinking about analyzing your own data. I hope I haven't made a fool of myself because of my ignorance about this type of analysis or about the underlying biology, but I think again, even as an outsider, I can have opinions, I can have perspectives that will help to improve the interpretation of the data visualization. Well, that's enough for today. Thanks for watching. Please subscribe. Please tell your friends what we're doing here on Code Club and I will see you next time for the live stream where we will try to recreate this data visualization.