Video summary
The video presents a constructive critique of a complex figure from a high-impact paper published in *Nature Cancer*, which investigates how stromal cell reprogramming drives tissue disorganization in nodal B-cell lymphoma. The speaker outlines a four-step evaluation process: description, analysis, interpretation, and judgment. In the description phase, the presenter provides context on the study's focus on comparing healthy reactive lymph nodes against two types of lymphoma—follicular lymphoma (FL) and diffuse large B-cell lymphoma (DLBCL)—using single-cell RNA sequencing and multiplex immunofluorescence data. The critique notes that while the paper is open access and well-cited, the figure's title is overly descriptive rather than declarative, failing to clearly state the main takeaway for the reader.
Moving into the analysis phase, the speaker examines how the visualizations were constructed, noting discrepancies between the published code on GitHub and the final rendered figures, suggesting that manual adjustments were made outside of the original script. The critique focuses heavily on panels E and F, which utilize faceted heatmaps and stacked bar plots to display cell enrichment and disease proportions. The presenter points out technical inconsistencies, such as mismatched color palettes in the source code versus the final output, and questions the statistical transparency regarding how dendrograms were generated and what criteria were used to cluster cell subpopulations into distinct facets.
In the interpretation and judgment phases, the speaker argues that the visualizations suffer from issues related to sample size imbalance and poor design choices that hinder clarity. Specifically, the presence of five patients in the FL group versus only four in the other groups creates a potential bias that is not adequately addressed by scaling the bars to equal width. The presenter strongly advocates replacing the stacked bar plot with a jitter plot to better reveal individual patient variation and suggests using alternative visualization methods like biplots to explain the factors driving cluster separation. Additionally, the critique highlights minor but significant usability issues, such as perpendicular axis labels, inconsistent legend styles, and unnecessary abbreviations that confuse readers unfamiliar with specific immunological jargon.
Ultimately, the video concludes by balancing praise for the authors' effort to simplify complex biological data with a call for greater transparency and improved design standards. The speaker commends the use of faceted heatmaps for organizing diverse cell populations but urges future researchers to provide clearer methodological details, such as distance metrics for clustering and baselines for statistical calculations. By offering specific suggestions like standardizing color schemes, removing axis expansion tails, and rewriting abbreviated terms for clarity, the presenter aims to help both the original authors and the broader scientific community create more interpretable and honest data visualizations that accurately reflect the underlying biological realities.
Read the full video transcript
Hey folks, welcome back for another
episode of Code Club. In this episode, I
will be doing a constructive critique of
this set of panels that was recently
published in the journal Nature Cancer.
When I do my critiques, I try to be as
constructive and positive as possible by
following a four-step process. The first
is description, where we try to get the
context of the figure and trying to
understand the point and the kind of
utility of the figure in the overall
work that is being presented. The second
is analysis, to try to break down our
figure or set of panels into its
constituent elements to understand how
this figure was assembled. The third is
interpretation,
trying to understand what the figure is
saying and what the authors want me to
take away from visualizing their data in
the way that they've presented it. And
finally, the fourth step is judgment,
where we try to render judgment on what
worked well, what didn't work so well,
what I'd like to see improved or perhaps
improve the interpretability, or just
some kind of ideas that I have for
things that I would like to try with
their data. Let's head into description,
where again we try to get the context of
the figure in the overall paper. At a
very broad level, the figure that we're
talking about was published in a paper
in the journal Nature Cancer. I haven't
talked about Nature Cancer in these
critique videos, but it's a very
high-impact journal, at least from my
perspective, with a journal impact
factor of 28.5,
which tells me that papers in this
journal tend to get cited a lot. The
title of the paper was reprogramming of
stroma-derived chemokine networks drives
the loss of tissue organization in nodal
B cell lymphoma. I was drawn to this for
a couple of reasons. First, I was
intrigued by the figure and how I would
make it. And also, about six or seven
years ago, I had lymphoma. Not the type
of lymphoma that they're talking about
here. Don't worry, I'm doing fine. Um,
but whenever I see lymphoma,
uh, I jump at the opportunity to study
it because it's relevant to me, right?
And it's something that I care about and
my family cares about, of course. Also,
this paper, as you'll see here in the
red, was published open access. And so,
down below in the description to this
critique, I will be presenting a link so
that you can go get the article without
having to pay anything. There is a long
list of authors here, so many that they
don't show them all to you. I believe I
counted 39 co-authors. There are four
senior corresponding authors. The first
three authors, I'm going to butcher your
names, I'm sorry. Felix
uh, Srdan Novakovic,
Anna Medaoki, Lia Jobava, are three
co-corresponding authors. The work was
done in Germany, primarily at the
University of Heidelberg in Heidelberg,
Germany. And then some of the other
authors are at the University Hospital
in Heidelberg, Germany, as well.
Something I always love to look at is
the peer review history of this paper.
We can see that it was submitted, it was
received June 28th of 2024,
and it was accepted February 10th of
2026.
So, again, that took about 20 months to
go through peer review and be accepted
before finally being published at the
end of March 25th, 2026.
So, again, from submission to
publication, I would count that at 21
months.
Almost 2 years to get published. Kudos
to these authors for their dogged
determination to see this work all the
way through and have it published in
Nature Cancer. Again, this work is
talking about nodal B-cell lymphoma and
the loss of tissue organization in lymph
nodes. Reading through the abstract, we
get a bit more context on the science
here. Lymph node function requires the
organization of cells into higher order
spatial units. However, the principles
governing lymph node architecture in
healthy and disease remain poorly
understood. Therefore, here we use
single cell and spatial mapping to
investigate the mechanisms directing
immune cell organization in human lymph
nodes and its disruption in
architecturally distinct lymphoma
entities, indolent follicular lymphoma
and aggressive diffuse large B cell
lymphoma, DLBCL. They go on to say that
our data substantiate the central role
of lymph node resident stromal cells and
chemokine driven lymphocyte zonation and
reveal an inflammatory feedback loop
fueled by tumor reactive T cells that
trigger stromal remodeling, progressive
loss of homeostatic chemokine gradients,
and tissue disorganization from
non-malignant state to FL, follicular
lymphoma, and DLBCL,
which again is that diffuse large B cell
lymphoma, the more aggressive form of
lymphoma. That's getting a little bit
beyond what we're going to be looking at
with this figure 1 E and F. So, let's go
ahead down to the results section where
they introduce figure 1 to get more
context about the figure that we'll be
looking at. So, the very first section
of the results section starts it out
mapping the spatio-cellular principles
of lymphoma induced lymph node
remodeling. And so, what they tell us
in figure 1A right here is the layout of
their study. And that they were
interested in looking at non-malignant
reactive lymph nodes, so the RLNs, with
a canonical organization into B cell
rich follicles and T cell zones.
Malignant lymph nodes from patients with
FL, so the follicular lymphoma,
characterized by a follicular growth
pattern of malignant B cells, and then
lymph nodes from patients with DLBCLs.
Immunologists love their acronyms, love
their jargon. So, they've got three
populations, right? They've got
non-malignant reactive lymph nodes that
have the typical expected organization
of B cells and T cells to make lymph
nodes. And then they also looked at
patients with FL, a form of lymphoma
where you see this B cell remodeling,
but not as aggressive as the DLBCL.
So, what they did in figure 1A then was
layout what the study looked like, the
patients that they had.
Don't get me started on the circular
format.
Um but then they did some single-cell
RNA-seq work, and then they also did MIF
or multiplex immunofluorescence.
And they generated these ordination
diagrams using UMAP on large numbers of
cells, uh about 30,000 in the first case
and about 3.7 million in the MIF, the
multiplex immunofluorescence. And it's
actually this multiplex
immunofluorescence what's presented in E
and F. To investigate spatial lymph node
organization, is they obtained a 56-plex
multiplex immunofluorescence data set
comprising 3.7 million cells from four
individuals with reactive lymph nodes.
Again, that's the the more healthy
state. Five with follicular lymphoma,
and then four from people with diffuse
large B cell lymphoma. There was some
overlap between the single-cell group
and this MIF group. Although, they don't
really seem to be talking much about the
overlap of those two different data
sets. What they tell us then is that
they identified 21 hematopoietic and
structure-defining non-hematopoietic
cell types or states,
uh including all these.
Um the latter included a distinct subset
of B cell zone-residing FDCs defined by
CD21 and CXCL13
positivity. I'm not going to go into
that because I don't understand it
myself. But I think basically what they
were doing is they identified these
different regions. And so then when we
come down to E and F, we see those same
regions down here. Okay? So basically
they went from this ordination data from
the multiplex immunofluorescence to look
at the structure and they give you some
examples of what that structure looks
like down here where they and they have
the six different neighborhoods right
here in these three examples from the
RLN, the reactive lymph node, the
follicular lymphoma, and the DLBCL.
And so then again, you can see like this
creamish color I believe is the T zone,
which again comes back from up here,
which I believe is probably this purple.
I It would be nice to have overlapping
colors and overlapping jargon, but
don't get me started. As I already
mentioned, they talked about the BEC
rich T zone, which is probably this BEC
here, the T zone, the lymphatic sinus,
plasma cells, follicles, and the B prol
neighborhood. So again, that's a bit of
the context, the description of where
the data in E and F are situated in the
overall study. Again, looking at the
spatial organization and the factors,
the cell types that give rise to this
transition from kind of canonical,
typical structure to disease within the
follicular lymphoma and the large
diffuse B cell lymphoma. Overall, this
paper has five main figures. When I
counted up the individual panels,
they're about 50 individual panels. As
we can see in the one we're looking at,
like panel B here, it's kind of like two
panels. The figure one, I would say this
is a descriptive title, right? It's just
saying lymphoma induced remodeling of
the lymph node cellular ecosystem. I
think that is descriptive. It's not
telling us what they want us to take
away. They're not telling us
what the remodeling looked like, right?
They're telling us that here's data
related to the remodeling. Okay, but but
what do you want me to see about the
remodeling? I don't know, right? It's
not in the title. The other figures
though have more declarative titles like
dysregulation of chemokine networks and
FRCs is linked to a fused growth pattern
in lymphoma, right? Or a reprogramming
of lymph node stromal cells from a
functionally tissue organized state to
an inflammatory state underlies the loss
of tissue organization in lymphoma,
right? These are declarative. They are
telling me what they want me to take
away. Where again, in figure one, that's
much more descriptive. It's describing
the type of data in the figure rather
than telling me what they want me to
take away from these different panels.
Something I always like to do before
digging into the figure is trying to
understand how they made the figures.
So, if I do a search for Prism, nothing
comes up. If I look for R, I see that
some of the analysis was done in R
um and that they also have some
citations to various R packages and
analysis that you can do in R and
methods about how to do things in R. I
notice there's also a code availability
section where they have put their custom
code on GitHub. If we go to that
repository, we can see that they have R
markdown documents for each of the
figures. And figure one here has an R
markdown document stepping us through
how they made their code, right? And a
lot of this is of course using R because
it's R markdown. Something I see that
they use was complex heat map, which is
a heat map generating R package that
does things more sophisticated than uh
like a normal heat map like you might
get out of ggplot2. When I was looking
through this code, something I I that
didn't quite align with the actual
rendered figures. Down here they're
using different shades of green, orange,
and purple. And this is in figure F
left, which again down here this is F
left. They're using purple, blue, and
green, not purple, orange, and green. So
that's a little bit odd. Also, the
number of colors, 8, 10, and 10, doesn't
quite match what they had in the final
paper. And so this is probably code that
they developed along the way to generate
the analysis and it was probably
modified as they went along. So I don't
think this is the actual code that they
used to generate the figures that wound
up in the paper. Also in here there's
comments like for C saying it's Fiji
output, no need for code. And so it
appears that these panels were made in
different parts and then assembled
together probably in something like
Illustrator or PowerPoint. So we get a
little bit of an insight into how they
made these panels with in this figure
largely using R. I did see some mention
of Python in some of their work, but
it's largely made in R and then probably
compiled somewhere else with some
probable other modification along the
way. So continuing on the analysis phase
of this critique, let's now look at
panels E and F and let's start with E. E
is a heat map, right? And it's got
dendrograms showing how the different
rows and columns relate to each other.
The coordinate system I would say is a
discrete coordinate system. Sure we
could say it's Cartesian, but it's got
discrete values on the X axis, the
different subpopulations of cells, and
then on the Y axis it has the different
categories being the different zones,
neighborhoods, cell types, or whatever
that they described up above in panel D
or in panel B on the right, right? Where
they used that UMAP ordination of the
MIF data. The values on the X and Y axes
are being sorted by the clustering of
the dendrogram. And I would say that the
dendrogram is a statistical layer. So,
the fill color within each of the cells
of the heat map here on the left is an
indication of the enrichment, and we see
its legend over here in the bottom right
of panel E, where it says enrichment log
2 OR. That's the odds ratio. So, it's
taking a log 2 of the odds ratio. So,
what zero would mean here would be one,
and what four here would mean would be
two to the four, which is 16, or two to
the negative four, which would be 1 over
16. So, again, the dark red are more
enriched, and the dark blue are highly
depleted. I don't know that I would
necessarily call this an annotation, but
it's perhaps a sense of organization in
the data, is that they've taken these
cellular subpopulations
and broken them apart into five
different groups, or five different
facets, I would think of them as. And
again, that is being defined by the
clustering across the columns of data. I
think that's interesting to have that
type of organization within the
visualization, to say that these five
subpopulations go together, this one's
on its own, these four go together, and
so forth. All right, so that was 1E.
Let's now turn to 1F, and here there are
two parts to 1F. There's a left, and
there's a right. For F on the left,
we'll start there. On the X axis, they
have the relative abundance, or the
proportion of cells in this particular
zone or neighborhood,
and for this particular patient and
disease entity, okay? And so, each
rectangle, like this rectangle here,
corresponds to one patient, who in this
case had DLBCL,
and the fraction of their cells that
showed up in the B prol neighborhood,
okay? So, it's maybe a lot to wrap your
head around there, but that's kind of
how it breaks down. The Y axis, again,
is the same as what we saw in E, which
are these six different clusters of
overall zones or regions or cell types
that their ordination analysis allowed
them to break things down into. Again,
the position on the Y axis is being
dictated by the clustering in the
dendrogram shown on the left side of 1
E. The fill color then, again,
corresponds to the three different
disease groups, whether it's a purple, a
blue, or a green. And that different
shades of purple or different shades of
blue or different shades of green
corresponds to the different patients
that were representatives of those
different disease entity groups. In
terms of the geometry or the type of
plot, this is a stacked bar plot laid on
its side. And so, within, say the
ggplot2 framework, you could generate
this using something like geom_col.
Of course, you would have to flip your
traditional X and Y axis coordinates to
get them to lay down on their side. I
don't see any statistical layers or
annotation layers like we did in panel 1
E. So, let's move on to the right side
of panel F, where here they have another
heat map. And again, heat maps can be
generated in ggplot2 using the geom_tile
function, where again, you give it
discrete values on the X and Y axis and
a continuous variable, or I guess you
could do discrete, but in this case
continuous variable
as the fill color. And so, here on the X
axis, again, is discrete, we have the
three disease entities, right? So, the
reactive lymph node, the follicular
lymphoma, and the diffuse large B cell
lymphoma.
And the fill color is telling us what
percentage of the cells within that
neighborhood, which again is defined on
the Y axis, as it was for the left side
of 1 F and in 1 E, proportion of cells
that could be assigned to that disease
entity. And so, basically what we're
seeing with this very pale color, uh
where my cursor is, is that there
weren't many cells at all from the
reactive lymph node, right? And so we
can see that over here in purple that
very few of the cells are represented in
these purple rectangles and perhaps a
similar amount by the blue rectangles,
but that most of them are in these green
rectangles. And so that's why
the B pearl neighborhood is dark red
under DLBCL and quite faint for these
others. That basically this heat map on
the right as a summation of say the
purple the blue and the green going to
the three different columns in this
right hand heat map. I would concede
that in some ways this right hand heat
map is a statistical summary of the data
on the left side. They're basically
summing the percentage of cells across
the patients within a disease entity and
then representing that as a different
shade of red in this right hand heat
map. So now we've broken things down
into their constituent parts and try to
put them back together again. What we'd
like to do now is to interpret the data.
And so I always like to interpret the
data without reading what the paper
wants me to take away from it. In E, my
takeaway is that we have these six
different general populations of cells
or regions within a lymph node.
And what they're trying to do is they're
saying like these were defined by our
MIF data, right? By their
immunofluorescence data. And they want
to better understand like what went into
defining these clusters of MIF data,
right? And so they're breaking it down
then by these different cell populations
that their assay, their multiplex
immunofluorescence, was able to pull
out. And so trying to say like what
defines the plasma cell group? Oh, it's
plasma cells, right? This PC. And it's
also these to some degree, right? But
it's certainly not these four cell types
because those are quite blue. And also,
again, the the B prol is pretty
definitive of the pre prol neighborhood,
right? And so, I think what they're
trying to do here is to lend some
information to understand the sub
populations of cells that are going into
making these six different categories.
Then on the right side in F, what
they're trying to do is then to the link
that to the disease entities by saying,
"What of these six different categories
define the three different disease
entities, right?" So, again, the DLBCL
is quite abundant in the B prol
neighborhood. And so, if you have a lot
of cells in the B prol neighborhood,
then it's likely that you've got
something from a DLBCL patient. And that
if you've got something in the plasma
cells or the follicles, that that's
going to be more indicative of the
follicular lymphoma, right? And and so
forth. And you know, in some places it
gets a little bit more muddled uh as to
what is being defined. I think some of
the other things that I take away,
especially in panel F on the left side
stack bar plot, is that there's a lot of
variation in the data, right? So,
looking at like the plasma cells, you
know, I had just said that that's kind
of indicative of the FL state, but
there's really only one patient that has
a large number of plasma cells. Uh and
so, is that indicative of FL or not? I
don't know, right? And so, again,
there's a fair amount of variation in
the data among patients within a disease
entity. So, in their results section,
they go on to describe this MIF data.
So, we're looking at, again, panels E
and F, where they say, "As expected,
follicular neighborhoods were expanded
in both size and number in FL." And I
believe that's this data down here,
looking at the follicles, that it is
more common in patients with FL. They go
on to say that moreover, significantly
enlarged plasma cell neighborhoods and
enrichment of BECs in the T cell zone
were observed in FL lymph nodes. They
talk about those plasma cells being more
pronounced in FL. Again, that's that's
one subject. I don't know that I would
go for that. They then go on to say that
these BEC-rich T zone cells were more
pronounced or expanded in the FL. And I
don't know that I totally see that.
Again, there are five patients in the
blue where there's four in the green and
the purple. And so, it might be that
that looks expanded because there's five
rather than four. And so, if we had
four, maybe it wouldn't seem so
expanded. Maybe it would just seem kind
of consistent across all three disease
entities. They go on to say that
moreover DLBCL lymph nodes displayed a
near complete depletion of lymphatic
vessels and an expansion of non-FDC FRCs
pointing toward a potential role for of
stromal cell remodeling in the
structural reorganization of diffusely
growing lymphomas. And again, this I
think points to the cases up here where
like the T zone lymphatic sinus, plasma
cells, and follicles are much more
present among individuals with DLBCL
relative to the other disease entities.
So, I want to be a little bit guarded in
saying that my interpretation doesn't or
does agree with theirs just because
again, this is way out of my wheelhouse,
this type of analysis and this type of
biology that they're interested in. But
I do have some questions that I would
ask the authors about their analysis and
kind of already mentioned one of those
being kind of this idea that there's
five FL patients versus four for the
others. And this large variation. Can we
really say that plasma cells are
expanded in people with FL relative to
others when there's only one patient
that had any appreciable level of FLs in
them. Now, let's move on to the judgment
phase of the critique. What did I like
about this? Let's always start with the
positives. So, I really was intrigued by
this faceting of their heat map. Again,
I was aside from kind of my own personal
interest in lymphoma when I saw this
heat map, I was really intrigued by how
would I do this in R? And so that really
got my attention that they have these
fascinating of
these different cell sub populations
into the five different groups that are
being defined by the dendrogram
clustering those different columns. I
thought that was pretty cool. Something
else I really liked was in panel F that
they use different shades of the same
color to indicate replication of those
different disease entities, right? So
they use different shades of purple or
blue or green. I thought it was pretty
cool. Um and because I don't care
necessarily about who the sample is from
or which replicate number it is, I don't
really need to worry a whole lot about
was this the same shade of blue across
these different um you know, the the the
six different groupings or not. I just
need to know it's blue or it's a
different shade of blue. And and perhaps
you could go ahead and be like, yeah,
like this blue and this blue are the
same and that and that's the same,
right? But at the end of the day, it's
blue. And so I know that comes from a
patient with FL. And I think that's
really what is most important in
interpreting panel F. So those are some
positive things that I could take away
from these panels to help me think about
my own future data visualization efforts
and things that I might want to
incorporate. As a general negative, I am
left wondering why this is two panels
instead of one panel or perhaps even
three panels, right? That the left and
the right really do go together and it's
the very much the same data. I mean,
they share a Y axis across all three of
these different plots. So why isn't that
one panel? I don't know. Or why isn't it
three panels, right? That we've got this
other heat map off here on the right.
And perhaps it is, as I mentioned, that
this is kind of a summarization of the
data in the left side of F. So maybe
that's why it's two. Anyway, this is
something I've been thinking about
lately is why do we label panels as
different panels? When do we label them
as different panels or when do we
collect different plots together into
the same
panel letter? I don't know. It's If
you've got thoughts on how to letter
different panels or how to pull
different panels apart, uh let me know
because I would I would love to see
different people's logic for why they
make
uh
multiple figures under the same letter
or
different letters for different panels,
right? And so this is something I
wrestle with with my own stuff and as
I've been looking through more and more
figures with many panels, it's something
that just strikes my curiosity. So let's
look at E. And E, I feel like is an
alternative approach to something that
we often do in microbial ecology or
ecology in general, which is called a
biplot. Where you have an ordination and
you then try to explain what are the
factors that are pulling clusters into
different parts of that ordination
space. So if we look at this MIF
ordination, which is again where the
data that we're looking at in E come
from, what are the cell types that are
pulling NK up here or what's pulling PC
out here or what's pulling BEC down
here? That with a biplot, they generally
would draw a vector coming down with
effectively these different cell
populations. That you'd have a different
vector for each of these cell population
and they would be centered in the middle
here and they would be going out to the
different cell populations describing
how much that subpopulation is driving
the separation of these different
populations within this ordination. I
don't I don't know that I totally buy
this approach to doing it, but I think
it's an interesting alternative to a
biplot. Where in this case at least,
you'd have like 21 different vectors.
And perhaps some of those vectors would
just get removed because they're not
that important. But um
again, it's an interesting alternative
and I'll have to think more about what I
think about this. One thing that I'm
already questioning is with that biplot,
the length of the vector tells you how
important it is in pulling that
population out of the center of the
ordination. And so, if you have a very
short arrow, it's not that important.
And so, oftentimes the short arrows we
might remove. But if it's a very long
arrow, we know that that arrow is
representing some factor that's really
pulling that cluster out. In this case,
I don't know how important these 20 or
whatever different subpopulations are at
an individual level in defining these
six different categories. And so, I feel
like that's something perhaps that's
missing. There might be something I'm I
am missing
in the interpretation of this, but it
would be nice to have some measure
indicating how important these different
factors are in discriminating between
these six different groups. Something
else I wonder about with panel E is is
this odds ratio, and how was the odds
ratio calculated? I suspect it was
looking at all the cells within a group
and then saying, "How much more enriched
is this?" So, say for example, the
granulo and the LEC are more enriched
across all of the columns within this
row. I think, right? And so, that you'll
see that rows always have something red
and something blue. And if you kind of
average across all of them, it gets to
zero, right? And so, you know, you've
got these very dark red colors and you
have these very dark blue colors
balancing things out. And so, perhaps
that goes back to my question of like,
"How important are these columns?" Or
you might look at something like T prol
or T reg, but they're very muted colors
and perhaps aren't really important in
distinguishing between these six
different groups. I don't know. So, it
would have been nice to have some
indication of how the odds ratio was
calculated. So, if I'm right about how
they're calculating the odds ratio, I
wonder what this would look like if they
calculated the odds ratio relative to
what they were seeing in the RLN, the
the reactive lymph nodes, which are not,
as I understand it, disease. They're
They're not cancerous, right? And so,
what would it look like to take all the
data and to scale it,
again, in panel E, according to the RLN
data as a baseline, instead of taking
each row as its own baseline and looking
at kind of the average, um what would it
look like to use that reactive lymph
node as the baseline? Also, there wasn't
any description of how the dendrograms
were drawn.
I could see that they used the R package
complex heatmap. My understanding is
that complex heatmap is using Hclust
under the hood to drive this clustering,
and the default for Hclust is a complete
linkage or furthest neighbor clustering
algorithm. Again, it would be nice to
say that. Um also, how are the distances
calculated? That's not indicated, and
again, I'll assume that they were
Euclidean distances. Also, related to
these dendrograms and the heatmap, is
that it's not clear what decision
criteria was used to pull apart the
subpopulations into the different
facets. And so, while I like the
faceting, I think that's cool, I would
like to know how did they pull that
apart? Was there a certain level of
dissimilarity
between the different clusters that they
used say, "Okay, this is one group, this
is another group, this is another
group."
Uh
that's missing, right? Or did they just
say, "Biologically, these four go
together. Look, they cluster together.
We're going to make them a cluster." and
so forth. That type of transparency, I
think, would be helpful to understanding
the visualization. So, I always like to
read left to right, uh cuz that's my
culture,
and so I always kind of complain when
labels, like we see on the x-axis here,
are written perpendicular to the axis. I
would love for these to be written
horizontally, and so one strategy I
always have is to pivot the whole thing
90°. But, of course, we have long names
on the y-axis, and so that wouldn't
work. And so, I think this worked out
pretty well. One thing that kind of
drives me nuts is that they are
abbreviating things that I don't think
need to be abbreviated. So, like PC, why
not write plasma cells like you have
here? If they wrote out plasma cell
here, it wouldn't take up any extra
space than any of the other
sub-populations that they already have
on the x-axis. And I suspect the same
could be true for many of these, right?
Uh I know immunologists love their
jargon and their abbreviation, but I
really do think it gets in the way. And
it would be really nice to write some of
these out so that when you're looking at
something like PC, you know that's the
same thing as plasma cells, and so it's
not surprising that that is so red.
Within the names of these six different
groups, I know there is some kind of
physiology going in here where they're
looking at the spatial structure, but
the fact that they have like zone, so
like the BEC rich T zone, the T zone,
and that they also have like the B prol
neighborhood, why isn't this a zone?
Is this like a different word for the
sake of having a different word, or is
there something else going on? Again,
this might be something in the analysis
or the biology, I don't know, but why
not use zone here or use neighborhood up
here? I think they use neighborhood a
lot throughout the paper, so maybe use
neighborhood, but neighborhood is a
longer word, so maybe zone would be
better. I don't know. I don't
understand. Um
it would be nice to be consistent though
for people like me who get easily
confused by these things. One final
thing I'll mention about panel E is this
legend enrichment log 2 OR. I really
think could go over here
um and be more closely connected with
the data versus kind of sitting in the
middle here where it's perhaps not
immediately clear what side it goes with
or if it's both, but saying oh yeah,
that goes over here uh because the
colors match, right? Well, why not just
move the legend over here? You've got a
nice big gap here. Let's put it there
cuz that seems like a like a logical
place for it. All right, so that's panel
E. Now, let's turn to panel F.
One thing I'll start out with panel F is
that there's no title on the x-axis. And
so, you have to go down through the
caption which is rather lengthy and in
the PDF version of the figure and
there's no caption. It's the entire
figure takes up the entire page. So is
the caption down here on the next page?
No, it's back up here.
And then I'm looking through here. F
left bar plot illustrating the
proportions of identified neighborhoods
across patient samples and disease
entities. It'd be far preferable to give
something to your audience to know what
this x axis represents. I get that it's
a proportion but proportion of what,
right? Go ahead and say that here.
You've already got it down in the
caption. Maybe come up with a more pithy
statement of saying that but put it
right here. Again, there's plenty of
room. Also, there's making more room
would be taking these numbers on the x
axis and making them parallel to the x
axis. There's no need for these values
to be perpendicular. It actually looks
kind of weird for them to be
perpendicular. Sure everything else is
perpendicular but generally when you
have a continuous x axis, those values
are parallel to the x axis as a
convention. So go ahead and flip those.
That'll get you more room here to then
put in your title. Also thinking about
the axis, I see that again they probably
used ggplot2 which naturally adds this
expansion factor to the x axis and so I
think we're seeing that here on the left
and right where we've got the tails of
the x axis sticking out beyond the data.
I'd go ahead and remove that expansion,
expand equals false, whatever
to go ahead and remove those tails so
that the axis starts at zero and ends at
one. As I've already mentioned a few
times, the FL category, the disease
entity has five patients represented
whereas the others only have four and so
I think it is clouding the
interpretation of the data. I don't know
how much this extra patient is biasing
my interpretation of the data. If all
the data were pooled at a equal patient
level, then I would expect the blue
category to be wider than the other
categories all other things being equal.
But, if they pull based on disease
category, then I would know that okay,
there's five, but it's been scaled so
that the three disease entities are
equal. And so, maybe to that, it would
be interesting to know what panel F on
the left would look like across all six
of the categories. Again, they're taking
these six categories and pulling them
apart. We don't totally get a sense of
how many cells fall into each of these
six different categories. And that would
be that would be informative. But again,
pulling all six together to give me a
sense of is it a third, a third, a
third, or is it like a 13th, a 13th, a
13th, a 13th, one per patient. And so,
it would be good to see that and to see
how evenly distributed the data were
across the 13 patients and the three
different disease groups. Otherwise, I
think they are biasing the data towards
overemphasizing the FL data. The last
thing I'll say about panel F on the left
here is that I do not like stacked bar
plots. I'm pretty well on the record
with that.
I would I would prefer to see is what
would this look like as a jitter plot,
right? And so, we could take the 13
different patients and represent them
each as a colored point, perhaps using
the same colors, and then we could
jitter the points within each of the six
different categories using the same
x-axis. And that way then it'd be easier
to see
how much the purple
you know, how much variation there is in
the purple relative to the blue,
relative to the green, and how different
the overall proportion is of the blue to
the green to the purple, right? Uh as it
is here, it's much more difficult to
interpret. I also think that this will
help a little bit with that problem with
the FL having five versus four patients.
Finally, on the right, we have this
extra heat map. Not totally sure I see
what the point of this is since it is
summing up these other bars that we see
on the left side of panel F. Regardless,
it's here and if we're going to use it,
let's make the best of it. One
suggestion I would have would be to use
a different color. They have white to
red, which matches the white to red that
they have with the enrichment going from
zero to four, here going from zero to
80. Again, use a different color because
you're representing different type of
data.
The other thing I notice about this
legend for panel F on the right is that
it has white tick marks and a white
border or no border at all, whereas the
legend for the heat map for E has black
tick marks and a black border.
It's a little subtle thing, but it'd be
nice to have consistency between the
two. So, either have white tick marks
and border on both or a black tick marks
and black border on the others. I kind
of like having the black border and the
black tick marks. It kind of matches
having the black border within the panel
F on the left legend as well. Well, that
is my critique of panel E and F today. I
hope you've gotten something out of this
and thinking about again how we
represent complex data using these types
of visualizations. I know that these
data sets are really complicated and so
I commend the authors for their efforts
in trying to simplify things. At the
same time, hopefully my suggestions are
helpful to them if this ever gets back
to them or to you as you're thinking
about analyzing your own data. I hope I
haven't made a fool of myself because of
my ignorance about this type of analysis
or about the underlying biology, but I
think again, even as an outsider, I can
have opinions, I can have perspectives
that will help to improve the
interpretation of the data
visualization. Well, that's enough for
today. Thanks for watching. Please
subscribe. Please tell your friends what
we're doing here on Code Club and I will
see you next time for the live stream
where we will try to recreate this data
visualization.