Promised Rain.

Why Fictional “Ranking” is Ass.

article thumbnail

EXCERPT

ARTICLE

An Introduction, and What I Mean by “Ranking.”

First, this will involve my own “ranking” system, and I am not making any claim that my system is perfect, but it works for me.

Second, by “ranking,” I am discussing the act of placing works (of any medium) into a linear order according to their perceived superiority. More specifically, by “linear ranking,” I mean an ordering in which every pair of works receives a determinate relative position.

A ≻ B ≻ C ≻ D ≻ …

Each work occupies a determinate position relative to the others.

Generally, this is done by evaluating aspects of a work and comparing those evaluations across works.

Evaluation ≠ Comparison ≠ Ranking

The ability to evaluate and compare works does not necessarily produce a determinate linear order.

Evaluation: “A is a 9.5/10.” Comparison: “I consider A better than B.” Ranking: “A is #3 and B is #7.”

These are related, but they are not fully equivalent.

What a Linear Ranking Assumes

A linear ranking implicitly requires several things:

  • Comparability: works can be meaningfully compared.
  • Completeness: every pair can be assigned a relative position.
  • Transitivity: if A > B and B > C, then A > C.
  • Determinate ordering: comparisons produce sufficiently precise positions.

The problem with this is that our actual judgements about fiction do not necessarily satisfy all of these conditions.

The Problem With Linear Rankings

Fiction is evaluated across multiple dimensions—writing, characterization, themes, symbolism, emotional impact, etc. Each medium further complicates this, possessing its own medium-specific dimensions: prose through its use of language and narration, poetry through rhyme, rhythm, meter, sound, and so on.

This is not a metaphysical claim that any of these evaluative aspects are ontologically objective—that is, of being completely independent of conscious minds, perception, and belief. The point is simply that these are dimensions along which we can form judgements.

Difficulty comes when we attempt to reduce all of them into a single overall ordering.

Criterion Dependence

The question “Which is better?” is underspecified.

We might instead ask, “Which is better written?”, “Which has better characterization?”, “Which is more emotionally impactful?”, or “Which is more thematically sophisticated?” These different criteria can produce different comparisons:

A ≻writing B

While:

B ≻emotional impact ​A

This becomes particularly relevant in mediatok, where statements such as “A has better complexity than B,” “A > B in characterization,” or “A ncods B in writing” can be presented as though they were objective measurements (or not—I make no claim on the intents of said individuals, but you know who you are; I also won’t pretend I haven’t participated in such actions myself, albeit I do consider—and hope—that everyone participating in them is doing so in jest).

Yet what exactly counts as “better writing” or “complexity”?

Psychological complexity? Development? Consistency? Dialogue? Internality? Thematic integration? Internal dialectical conflict? Validity? Existential scope? Identity palimpsest? Argument density? Criticism? Cognitive depth? I could go on, but you get the point.

The point isn’t really that these criteria are illegitimate. Rather, such comparisons can conceal a large number of unstated assumptions about what is being measured, how it is being measured, and why those criteria should determine superiority in the first place.

Essentially, comparisons are not discoveries of an objective relation between works; they are judgements constructed from a particular set of intersubjective criteria.

Incommensurability & Aggregation

Even once we establish our [intersubjective] criteria, another problem arises: how exactly are we to combine them?

Works can excel in radically different dimensions. There is no obvious universal unit by which we can say that “one amount” of characterization compensates for “one amount” of prose quality.

For example, comparing something like the maximalist, densely layered prose of Pynchon within Gravity’s Rainbow with the visual composition and paneling of a manga like Vagabond makes the difficulty of reducing their respective strengths to a common scale even more apparent.

Even if we evaluate each dimension individually, it does not follow that those evaluations can be cleanly aggregated into one overall measure of superiority. Like, suppose A has better characterization while B has greater thematic depth. What determines how much one compensates for the other? We could assign weights to each criterion, but then we introduce more assumptions: why those weights?

Thus, even when individual dimensions can be meaningfully evaluated, there is no necessarily neutral rule for combining them into one overall rule.

Intransitivity

A linear ranking requires comparisons to be transitive. As said earlier:

(A ≻ B ∧ B ≻ C) → A ≻ C
If A is preferred to B, and B is preferred to C, then A is preferred to C.

Yet preferences can potentially produce an intransitive cycle:

A ≻ B
B ≻ C
C ≻ A

Or equivalently:

A ≻ B ≻ C ≻ A

A relation is non-transitive if there are cases in which A > B and B > C, but A > C does not hold. A cycle such as the above is one way this can occur.

Thus, pairwise comparisons can exist without yielding a consistent linear order.

Parity & Incomparability

Two works can be comparable without one being better than the other, and without them being considered equal in value. Roughly, A and B are on par when:

  • A is not better than B
  • B is not better than A
  • A and B are not equal in value.
  • Yet they remain genuinely comparable in their evaluative standing.

So, for fiction, imagine two works that excel in radically different ways:

A: has exceptional characterization, but weaker thematic depth
B: has exceptional thematic depth, but weaker characterization

You may indeed have good reasons for comparing them, while still lacking enough grounds to say that A > B or reverse, while still rejecting A = B. Thus: A || B.

Where || represents parity rather than equality.

Also to clarify the distinction: incomparability means that there is no sufficiently appropriate basis on which A and B can be compared. Parity means that A and B can be meaningfully compared, but neither is superior to the other and they are not equal.

Both challenge the assumption that every pair of works must receive a determinate position.

False Precision

A numerical rating can create the appearance of a degree of precision that our actual judgements do not possess. For example, if I were to rate several works within the range of 9.0–9.9/10, the numerical scale might imply:

9.9 ≻ 9.8 ≻ 9.7 ≻ … ≻ 9.0

Yet I may actually have no clear basis for saying that a 9.9 is actually superior to a 9.8, or that a 9.8 is superior to a 9.7. The numbers can simply indicate how highly I evaluate a work without establishing a precise ordering between works.

This becomes even more obvious when, let’s say, multiple works receive the same score. If ten works are all rated 9.9/10, assigning them positions from #1 to #10 would suggest distinctions that my actual evaluations may not necessarily contain.

Consider the same problem on a larger scale. If i consider fifteen works (ω) to be roughly within my 9.0–9.9 range, forcing them into:

ω1 ≻ ω2 ≻ ω3 ≻ … ≻ ω15

would make it appear as though I have fifteen distinct comparative judgements. In reality though, I may simply regard all fifteen as exceptionally close in evaluative standing, without any clear or defensible basis for determining their precise order.

This is what I mean by false precision. The issue is not necessarily that the ordering is false. Instead, the degree of determinacy communicated by the ordering can exceed the degree of determinacy present in the underlying judgement.

A ranking can therefore manufacture distinctions that the evaluations itself never made.

My “Ranking” System

This is essentially where my own system comes from.

Instead of forcing every work into a definitive linear order, I separate evaluation from ordering. My numerical scores indicate the level at which I evaluate a work, while the comparative ordering of individual works remain deliberately indeterminate where I lack sufficient grounds to distinguish them.

The basic structure is an ordered set of evaluative ranges:

1.0–1.9 ≺ 2.0–2.9 ≺ 3.0–3.9 ≺ … ≺ 8.0–8.9 ≺ 9.0–9.9 ≺ 10

A work within the 9.0–9.9 range is evaluated above a work within the 8.0–8.9 range. Within those ranges, however, I do not necessarily impose an order between individual works.

For example:

9.0–9.9: A, B, C, D
8.0–8.9: E, F, G
7.0–7.9: H, I, J

The ranges themselves are ordered, while the works within them can remain unordered in respect of parity. The ordering between ranges accordingly expresses an ordering of my evaluations, rather than necessarily a pairwise ordering between every work within those ranges.

A || B || C || D ≻ E || F || G ≻ …

Importantly, the numerical score and the comparative relation represent different things. Say, if E is rated 8.5 and F is rated 8.9, then 8.9 > 8.5 numerically. This does not necessarily mean that I consider F superior to E overall. I may still consider E || F.

The numerical score expresses the level at which I evaluate a work; the comparative relation expresses whether I regard one work as superior to another. So in essence, my system combines an ordered evaluative scale (1.0–1.9, 2.0–2.9, … 9.0–9.9, 10) with an intentionally less determinate comparative relation between works.

The 10/10 Exception

My 10/10s are the sole special case.

I reserve 10/10 for works that have genuinely touched my heart. As of making this, there have only been two:

10/10: A, B

A and B (and whatever future work that may be placed here) remain unordered relative to one another. The criterion for entering this category is specifically personal impact, rather than an attempt to establish that these works are comprehensively superior to everything else I have experienced.

So 10/10 is not simply “the work with the highest technical quality.” It represents a special category that I have defined as of personal significance.

Quality vs. Favouritism

I do think there is a reasonable distinction to make between evaluative quality, and personal preference. These can overlap heavily, but they are not necessarily the same thing. I can consider A more accomplished in its writing, characterization, themes, etc., while personally still preferring B because I have a stronger emotional attachment to it.

So, when someone asks me “Which do you think is better?” and “Which do you prefer?”, I consider these different questions, well, obviously, I suppose it more so pertains to that of “X or Y?” Regardless, the former concerns my overall evaluation of the work; the latter concerns my personal preference or attachment to it.

This is also why my 10/10 criterion is specifically personal impact. A work does not receive that score merely because I consider it technically or comprehensively superior. It receives it because, above all else, it has genuinely touched my heart.

The Problem of a Static Ranking

“Complex journey through art.” You’ve probably heard it before if you’ve been in the space [of mediatok] for a decent amount of time. Anyhow, regardless of the guy I’m talking about, I do think there is truth in the saying. Our evaluations of fiction can change minimally or drastically as we experience more works, revisit old ones, develop different standards, or simply change as people. A strict ranking, however, presents itself as a fixed ordering: A > B > C > D.

If my judgement later changes to: B > A, then the original ordering no longer represents what I actually think. And if I maintain my old ranking simply for the sake of consistency, then I end up preserving an ordering that no longer reflects my current judgement.

This is also another reason as to why I prefer a looser structure. My “ranking” is a representation of my present evaluations. It may change as my judgements change, and it is therefore not a permanent or even enduring hierarchy of everything I have experienced.

What a “Ranking” Actually Should Preserve

So, what should a “ranking” actually preserve you may ask?

In my view, it should preserve the distinctions that our evaluations actually contain, while avoiding distinctions that they do not.

If I regard A as clearly above B, the system should be capable of representing that.
If I regard A and B as roughly equal in standing, it should allow me to preserve that as well.

If I have no meaningful basis for deciding whether A should occupy position #3 or #4, then there is little reason for me to manufacture some superficial reasoning for that distinction simply because the format demands it.

So, rather than asking yourself “How can I put every work into an order?”, I think a better question to ask is: “How can I represent my judgements as faithfully as possible?”

Limitations of My System

Of course, I do not claim that my system solves every problem.

The ranges themselves are ultimately constructed conventions. Their boundaries cannot possess perfect precision, and my judgements can change over time. A work I currently place at 8.9 may later become 9.1, while another work may move in the opposite direction.

The system itself shouldn’t really be understood as a permanent or perfectly calibrated measurement of artistic value. It is simply a way of representing my present evaluations while preserving the indeterminacy that exists within them.

Conclusion

I don’t actually think fictional rankings are inherently “useless” per se. Bit of an engagement bait, I suppose.

I think the problem is that the structure of a ranking can communicate more certainty, precision, and comparability than our actual judgements contain. And my own system is simply an attempt to preserve that uncertainty: ordered where my evaluations establish distinctions, and unordered where they do not.


Thank you for reading.