Espartaco

“Is that to say we are against Free Trade? No, we are for Free Trade, because by Free Trade all economical laws, with their most astounding contradictions, will act upon a larger scale, upon the territory of the whole earth; and because from the uniting of all these contradictions in a single group, where they will stand face to face, will result the struggle which will itself eventuate in the emancipation of the proletariat.”

Karl Heinrich Marx · Marx-Engels Collected Works, Vol. VI, p. 290

25,834 views since December 2020

25,834 visitas desde diciembre de 2020

EnglishEspañol

Category: Mathematics

  • DIFFERENCES BETWEEN LINE INTEGRALS, MULTIPLE INTEGRALS, AND SURFACE INTEGRALS

    DIFFERENCES BETWEEN LINE INTEGRALS, MULTIPLE INTEGRALS, AND SURFACE INTEGRALS

    Differences Between Line Integrals, Multiple Integrals, and Surface Integrals
    Vector Calculus Notes · Expository Edition
    University Calculus · Conceptual Guide

    Differences Between Line Integrals, Multiple Integrals, and Surface Integrals

    A gradual reading of what integration means when the domain is no longer an interval, but a curve, a region, or a surface.

    · · ·

    Preliminary Analysis

    When calculus is first studied, the integral usually appears as a procedure for accumulating quantities over an interval. Later, the same idea is extended: integration is no longer carried out only along an axis, but also over curves, regions of the plane, volumes, and surfaces. Line integrals, multiple integrals, and surface integrals are precisely some of the forms taken by this generalization.

    In general terms, line integrals allow us to accumulate a quantity along a curve. Multiple integrals allow us to do so over regions of two or more dimensions. Surface integrals, in turn, allow us to integrate over a surface located in space. At the level of university calculus, this distinction is enough to understand why each kind of integral requires different techniques while keeping in view that all of them belong to the same mathematical idea: infinitesimally summing a magnitude over a given domain.

    The main theorems associated with these integrals are also related. For line integrals there is the Fundamental Theorem for Line Integrals; when studying multiple integrals one encounters Fubini’s Theorem, the Pappus–Guldinus Theorems, and Green’s Theorem; and in the study of vector fields on surfaces one encounters Stokes’ Theorem and Gauss’ Theorem, also known as the Divergence Theorem.

    The fundamental difference does not lie in the idea of integration itself, but in the geometric object over which accumulation is carried out.

    Green, Stokes, and Gauss are not literally the same theorem in an elementary calculus course, but they express a common structure: they relate an integral carried out in the interior of a region to another integral carried out over its boundary. In a more general treatment, this unity is formalized through the generalized Stokes theorem; however, for an undergraduate course it is enough to understand them as different manifestations of the same principle.

    Line Integrals

    A line integral is an integral whose domain of integration is a curve. Instead of moving through the points of an interval on the real axis, one moves through the points of a path that may lie in the plane or in space. To perform the calculation, the usual procedure is to parameterize the curve.

    To parameterize means to describe the coordinates of every point on the curve through an auxiliary variable, usually \(t\). In the case of a straight line, the parameterization may take a simple form such as \(x=c+kt\). For a general curve, however, it is more appropriate to think in terms of expressions such as \(x=x(t)\), \(y=y(t)\), and, when necessary, \(z=z(t)\). The procedure therefore does not require a single constant direction vector; that feature belongs to the particular case of a straight line.

    Line integral of a scalar function along a spatial curve
    FIGURE 1 · A scalar line integral accumulates the values of a function along a curve. Original source indicated in the article: Greg School.

    There are two especially important cases. In a line integral of a scalar function, a magnitude is accumulated along the curve. If a wire has variable density, for example, a line integral can be used to obtain its total mass. In a line integral of a vector field, by contrast, one is usually interested in the effect of the field along the displacement: the classical example is the work done by a force along a path.

    Scalar line integral: ∫C f ds Vector line integral: ∫C F · dr
    Different types of line integrals
    FIGURE 2 · Different kinds of line integrals: integration with respect to arc length and integration of vector fields along a path. Original source indicated in the article: YouTube.

    When the curve is closed, the symbol \(\oint\) is often used. The circle drawn over the integral sign indicates precisely that the path returns to its starting point. This detail is particularly important when studying circulation, because the integral is taken around the entire closed boundary.

    The Fundamental Theorem for Line Integrals

    The Fundamental Theorem for Line Integrals plays a role analogous to the Fundamental Theorem of Calculus in one variable. If a vector field is conservative and can be written as the gradient of a potential function, then the integral between two points depends only on the initial point and the final point, not on the particular path followed from one to the other.

    If F = ∇φ, then ∫C F · dr = φ(B) − φ(A)

    This is one of the first indications of an idea that will appear repeatedly: certain complicated integrals can be transformed into simpler expressions when the geometric structure of the problem allows the appropriate theorem to be applied.

    Multiple Integrals

    Multiple integrals generalize the process of integration to domains involving more than one variable. A double integral is carried out over a two-dimensional region and a triple integral over a three-dimensional region. Their geometric interpretation depends on what is being integrated.

    For example, if the constant function \(1\) is integrated over a region of the plane, the double integral gives the area of that region. If a nonnegative function \(f(x,y)\) is integrated over a region \(R\), the integral may be interpreted as the volume lying below the surface \(z=f(x,y)\) and above \(R\). Likewise, a triple integral of \(1\) over a region of space gives its volume. One should therefore not mechanically identify “double integral” with “volume” or “triple integral” with “hypervolume”: the meaning depends on the integrand and on the domain.

    Geometric interpretation of a double integral
    FIGURE 3 · Interpretation of a double integral as the accumulation of infinitesimal sections over a region. Original source indicated in the article: AlgebraHD.

    Fubini’s Theorem

    Fubini’s Theorem is fundamental because, under the usual conditions studied in calculus, it allows a multiple integral to be computed by means of iterated integrals. In practical terms, this means integrating first with respect to one variable and then with respect to another. In triple integrals there are as many as six possible orders of integration, and choosing a suitable order can simplify the calculation considerably.

    ∬R f(x,y) dA may be computed, depending on the region, as ∫ [ ∫ f(x,y) dy ] dx or as ∫ [ ∫ f(x,y) dx ] dy

    Conceptually, Fubini’s principle is not limited to triple integrals. It is a general result concerning integration on product spaces. In an undergraduate calculus course, however, it is most often applied to double and triple integrals, which are also the cases that can be represented geometrically most easily.

    Orders of integration in triple integrals
    FIGURE 4 · A triple integral can be set up using different orders of integration. The geometry of the region determines which order is most convenient. Original source indicated in the article: SlidePlayer.

    The Pappus–Guldinus Theorems

    The two Pappus–Guldinus theorems relate centroids to objects generated by revolution. Intuitively, both say that when a figure rotates around an axis, a geometric magnitude generated by that rotation can be calculated by multiplying the original magnitude by the distance traveled by its centroid.

    The first theorem relates the length of a plane curve to the area of the surface of revolution generated by that curve. The second theorem relates the area of a plane region to the volume of the solid of revolution generated when it is rotated around an axis that does not intersect the region. This is the relevant formulation of the two classical theorems in calculus.

    Green’s Theorem

    Green’s Theorem establishes a relation between a double integral over a plane region and a line integral along the closed curve forming its boundary. If \(C\) is the positively oriented boundary of a region \(R\), then, under the usual hypotheses:

    ∮C P dx + Q dy = ∬R (∂Q/∂x − ∂P/∂y) dA

    The conceptual importance of the theorem is greater than the formula itself. Green’s Theorem allows an accumulation performed around a boundary to be replaced by one performed throughout the interior, or vice versa. It therefore forms a natural bridge between line integrals and double integrals.

    Surface Integrals

    A surface integral is an integral carried out over a surface located in space. Here it is useful to distinguish two cases, in a way analogous to what occurred with line integrals.

    In a scalar surface integral, a scalar function is accumulated over the surface. If a curved sheet has variable surface density, for example, an integral of this kind can provide its mass. By contrast, when a vector field is integrated through the dot product with the normal vector, one computes the flux of the field through the surface.

    Scalar surface integral: ∬S f dS Flux of a vector field: ∬S F · n dS
    Scalar integral over a parametrized surface
    FIGURE 5 · Surface integral of a scalar function. The surface may be described by a parameterization, and the area element must be adjusted to its geometry. Original source indicated in the article: SlidePlayer.

    If a surface can be written as \(z=g(x,y)\), it may be parameterized by \((x,y,g(x,y))\). The differential surface element is not simply \(dx\,dy\), because it must account for the local inclination of the surface. For the graph \(z=g(x,y)\), the factor that appears is:

    dS = √(1 + (∂g/∂x)² + (∂g/∂y)²) dA

    This factor corrects the area projected onto the \(xy\)-plane so that it becomes the actual area on the inclined surface. If the surface is described in another way—for example, by \(x=g(y,z)\)—the corresponding parameterization is used.

    Gauss’ Theorem or the Divergence Theorem

    When \(S\) is a closed surface enclosing a volume \(V\), Gauss’ Theorem relates the total flux of a vector field through the surface to the divergence of the field throughout the interior volume. If \(\mathbf F=(P,Q,R)\), its divergence is:

    div F = ∂P/∂x + ∂Q/∂y + ∂R/∂z

    and the theorem states:

    ∯S F · n dS = ∭V div(F) dV

    The symbol \(\oiint\)—or, typographically, a double integral sign with a circle—indicates that the surface is closed. Conceptually, Gauss’ Theorem says that the net flux leaving the boundary of a volume is determined by the total contribution of the “sources” and “sinks” of the field within that volume.

    Green, Stokes, and Gauss: One Underlying Idea

    At this point we can return to the initial observation and formulate it more precisely. Green’s, Stokes’, and Gauss’ theorems should not be treated as isolated results. All three connect what happens in a region with what happens on its boundary, although each one does so in a different geometric situation.

    Green works in the plane: it relates a double integral over a region to a line integral along its boundary. Stokes works with an oriented surface in space: it relates the circulation of a field around the curve forming its boundary to an integral of the curl over the surface. Gauss works with a volume: it relates the flux through the closed surface enclosing it to the divergence throughout the volume.

    Conceptual Scheme

    Green: plane region ↔ boundary curve.

    Stokes: surface ↔ boundary curve.

    Gauss: volume ↔ boundary surface.

    This resemblance is not accidental. In more advanced courses, the generalized Stokes theorem allows these results to be gathered within a single mathematical structure. Even without differential forms, however, the central intuition can already be understood at undergraduate level: under suitable conditions, an integral over the interior can be transformed into an integral over the boundary, and vice versa.

    In this sense, the progression from line integrals to multiple integrals and surface integrals is not an arbitrary collection of techniques. It is a gradual extension of the same idea of accumulation and, at the same time, an introduction to one of the deepest relations in vector calculus: the relation between a region and its boundary.

    Comparative Summary

    Type of integral Domain Typical interpretations Related theorems
    Line integral Curve Mass of a wire, work, circulation Fundamental Theorem for Line Integrals; Green
    Double integral Plane region Area, volume under a surface, mass of a lamina Fubini; Green
    Triple integral Region in space Volume, mass, volumetric accumulation Fubini; Gauss
    Surface integral Surface in space Weighted area, surface mass, flux Stokes; Gauss

    The difference among these integrals can therefore be summarized simply. A line integral accumulates along a path; a multiple integral accumulates over a region; and a surface integral accumulates over a surface. What changes is the geometry of the domain and, with it, the differential element that must be used. The fundamental idea of integration, however, remains the same.

    To integrate is to accumulate. What changes from one integral to another is where the accumulation takes place and what magnitude is being accumulated.
  • OFF-THE-CUFF REFLECTIONS: TOPOLOGY AND GEOMETRICAL CHANGE (2021)

    OFF-THE-CUFF REFLECTIONS: TOPOLOGY AND GEOMETRICAL CHANGE (2021)

    Philosophy of Mathematics · Topology · Dialectics

    What Remains When Everything Changes Shape?

    A guided reading of Toward a Dialectical-Materialist Interpretation of Topology, by José Mauricio Gómez Julián
    · · ·

    Imagine a coffee mug made of rubber. We can stretch it, compress it, twist it, and gradually deform it into something resembling a doughnut. To ordinary geometry, the mug and the doughnut are very different objects: they have different curvatures, lengths, and proportions. Topology, however, asks a deeper question: did their essential structure really change, or did we merely alter their metric appearance?

    That question — what may vary without an object ceasing to belong to the same structural class — lies at the conceptual heart of topology. But José Mauricio Gómez Julián’s essay seeks to go one step further. Rather than merely presenting mathematical definitions, it reconstructs the history of the discipline in order to ask what philosophical meaning lies in studying precisely those properties that remain through certain transformations.

    The journey moves through Leibniz, Euler, Cantor, Dedekind, Poincaré, Peano, Brouwer, and Hausdorff; through the bridges of Königsberg, the paradoxes of dimension, set theory, continuity, and homeomorphisms. Eventually, these threads are brought together in a proposal: to read topological structure from a dialectical-materialist perspective, as a mathematical way of thinking about the relationship among transformation, invariance, structure, and qualitative change.

    01 · The Fundamental Problem From Measuring Objects to Studying Relations

    For centuries, thinking geometrically meant above all thinking in terms of magnitudes: lengths, areas, angles, distances, and proportions. Topology introduces a change in perspective. What matters is no longer exclusively how much something measures, but also how its parts are related.

    The essay finds a decisive antecedent in Gottfried Wilhelm Leibniz. In the seventeenth century, Leibniz imagined a geometria situs, a “geometry of position”: a discipline in which the relative arrangement of elements would take priority over their magnitude. The intuition was remarkably modern. Two configurations might differ in their metric dimensions and yet share something deeper in their organization.

    Topology asks less about how much an object measures than about which relations survive when its shape changes.

    This is also the first useful key for readers coming from economics or the social sciences. An economic network may change enormously in the volume of its transactions without necessarily changing its basic pattern of connections; an institution may grow or shrink while preserving certain internal relations; a political structure may undergo quantitative modifications without yet experiencing a qualitative transformation in its organization.

    This does not automatically turn such questions into problems of mathematical topology. It does, however, help us grasp the intuition that interests the essay: distinguishing between changes of magnitude or appearance and changes of structure.

    02 · Königsberg, 1736 Euler and the Birth of a New Way of Seeing

    One of the foundational episodes in this history takes place in the Prussian city of Königsberg. The city was divided by the Pregel River and connected by seven bridges. The problem was easy to state: was it possible to take a continuous walk crossing every bridge exactly once?

    Leonhard Euler realized that the distances, the sizes of the islands, and the lengths of the bridges were irrelevant. Each landmass could be replaced by a point, and each bridge by a connection between points. The physical problem was thus transformed into an abstract structure.

    The Mathematical Idea

    What we would now call a graph preserves only the information relevant to the problem: which vertices are connected by which edges. The Königsberg problem is a problem of an Eulerian traversal: it asks whether every edge can be traversed exactly once.

    Euler showed that this was impossible. All the relevant vertices had odd degree, whereas a traversal using each edge exactly once can have only zero or two vertices of odd degree.

    Yet for the historical argument of the essay, the decisive point is not merely the solution. It is the method of abstraction. Euler deliberately removed information about magnitude in order to preserve a structure of relations. A real city, with water, bridges, and distances, became a mathematical object whose organization could be studied independently of scale.

    03 · A Productive Crisis Cantor and the Strange Problem of Dimension

    The next major leap appears in the nineteenth century with Georg Cantor and set theory. Cantor discovered that the points of a line segment and the points of a square can be placed in one-to-one correspondence: both sets have the same cardinality.

    This result was profoundly counterintuitive. A segment appears one-dimensional and a square two-dimensional. How could they contain, in a precise sense, the “same number” of points?

    An Essential Distinction

    Having the same cardinality does not mean having the same topological dimension. That was precisely the problem: counting points is not enough to capture what intuitively distinguishes a line from a surface.

    In philosophical terms, the contradiction between geometric intuition and set-theoretic result forced mathematics to reformulate the question. If dimension could not simply be reduced to the number of coordinates or to the cardinality of points, a deeper structural property had to be discovered.

    The essay places particular emphasis on episodes of this kind: contradictions do not appear merely as unpleasant accidents in science, but as engines of conceptual development. A notion that once seemed self-evident — “dimension” — becomes problematic, and by becoming problematic it forces the construction of a deeper theory.

    1676 · Leibniz

    Imagines a geometry based on position rather than magnitude.

    1736 · Euler

    Reduces the Königsberg problem to a structure of vertices and connections.

    1877 · Cantor

    Correspondence between sets of different apparent dimensions destabilizes the old geometric intuition.

    Late 19th Century

    Dedekind, Peano, and others force mathematics to distinguish among cardinality, continuity, and dimension.

    Early 20th Century

    Poincaré, Brouwer, and Hausdorff consolidate the problems that will shape modern topology.

    04 · When a Definition Also Asks About the World Poincaré: Continuity, Cuts, and Meaning

    Henri Poincaré occupies a special place in the story because his questions about dimension did not arise solely from technical difficulties. He was also interested in understanding why we experience space as three-dimensional, what relationship exists between mathematical geometries and physical space, and where our geometric intuitions come from.

    His idea was to think about dimension through cuts. In intuitive terms, the dimension of a continuum could be investigated by asking what kind of object must be removed in order to divide it. A line can be disconnected by removing a point; separating a surface generally requires something of higher dimension.

    The essay grants this idea particular philosophical significance. A dimension no longer appears merely as a coordinate drawn along an axis, but begins to be related to the way in which the parts of a space are connected.

    A structure is defined not only by its components, but by the system of relations that makes those components into a whole.

    Poincaré did not thereby provide the final mathematical word on dimension. His proposal encountered difficulties and would eventually be replaced by more robust formulations. Yet for the historical reading developed in the article, that is precisely the point: a formulation may be mathematically superseded while still preserving a fertile philosophical intuition.

    05 · The Consolidation of the Discipline Brouwer, Hausdorff, and Modern Topology

    L. E. J. Brouwer brought the problem of dimension to a new level of rigor. Among his fundamental contributions was the invariance of dimension: Euclidean spaces of different dimensions cannot be equivalent through a homeomorphism. A line and a plane do not become structurally identical no matter how ingenious the correspondence between their points may be.

    The result is important because it separates two ideas that Cantor had forced mathematicians to confront: two sets may have the same cardinality and yet possess different topological structures.

    Felix Hausdorff, in turn, contributed to transforming topology and set theory into increasingly abstract and systematic disciplines. By the beginning of the twentieth century, mathematical “space” no longer had to be imagined as a physical room filled with geometric points. Its elements could be functions, sequences, or other abstract objects.

    This generalization is decisive. Topology ceases to be merely a strange geometry of deformable surfaces. It becomes a language for speaking about continuity, neighborhoods, convergence, connectedness, and structure across enormously broad classes of mathematical objects.

    06 · The Mathematical Core What Is a Topology, Really?

    We can now state the idea precisely. Let X be a set. A topology on X is a collection τ of subsets of X — called open sets — satisfying certain rules.

    ∅, X ∈ τ
    arbitrary unions of members of τ belong to τ
    finite intersections of members of τ belong to τ

    The pair (X, τ) is called a topological space. What matters is that τ determines what it means to be “near,” what continuity means, and how the space is organized without requiring any numerical notion of distance.

    A metric may tell us that two points are 3.7 units apart. A topology can study relations of proximity and continuity even when no distance function exists at all.

    Concept Intuition
    Metric Allows distances between points to be quantified.
    Topology Describes a structure of neighborhoods, continuity, and more general spatial relations.
    Homeomorphism A continuous bijection with continuous inverse between two topological spaces.
    Topological invariant A property that remains unchanged under homeomorphisms.

    The Famous “Rubber-Sheet Geometry”

    From here comes the classical metaphor. We may stretch, compress, or twist an object as long as we do not cut it or glue together parts that were previously separate. A circle can be deformed into an ellipse without leaving its topological class.

    A sphere can be deformed into an ellipsoid. Creating a hole in the sphere in order to transform it into a torus, however, requires a topologically radical modification: we are no longer merely changing distances and curvatures, but the structure of the object itself.

    The Central Point

    In topology, “preserving structure” does not mean preserving visual appearance or distances. It means preserving those relations encoded by the topological structure. Mathematically, the relevant notion of equivalence is the homeomorphism.

    07 · From Formalism to Meaning The Dialectical-Materialist Reading

    Up to this point, we have topology. The specifically philosophical move of the essay begins when it asks what this kind of mathematics tells us about the relationship among structure, transformation, and permanence.

    The author’s proposal begins by distinguishing between changes that affect certain properties of a system without destroying its fundamental structure and changes that do alter that structure. The distinction immediately recalls a central category of dialectics: not every quantitative modification yet constitutes a qualitative change.

    A topological object may be stretched, twisted, or deformed within certain limits while retaining its invariants. But when it is torn, when a new connection is created, or when an essential connection is eliminated, a different class of structure appears.

    The essay interprets this difference through the dialectical relation between form and essence. Form may vary considerably while certain internal relations remain stable; when transformations reach the very organization constitutive of the system, change ceases to be merely formal and becomes qualitative.

    Invariance does not mean immobility: something may change profoundly in appearance while preserving, through those transformations, a determinate structure.

    This is perhaps the most interesting conceptual bridge proposed by the text. “Remaining” and “changing” cease to be mutually exclusive absolutes. A system can change precisely because it possesses a structure within which certain changes are possible. That same structure also determines which transformations would cease to count as internal modifications and instead become a rupture.

    Homeomorphism and Structure

    The homeomorphism therefore acquires special philosophical importance for the author. Technically, two spaces are homeomorphic when there exists between them a continuous bijection whose inverse is also continuous. Philosophically, the article interprets this as a formalization of the idea that externally different configurations may share the same structural organization.

    This is not because a mug and a torus are “the same thing” in every possible sense, but because a particular level of abstraction permits them to be treated as equivalent with respect to the properties studied at that level.

    This connects with another important epistemological thesis of the essay: every science abstracts. Physics, chemistry, biology, economics, and mathematics isolate particular relations in order to investigate them. To abstract does not necessarily mean to deny the rest of reality; it means provisionally selecting which relations will be treated as essential for a particular problem.

    08 · An Excursion Beyond Mathematics From Abstract Space to DNA

    To show that topological language is not confined to geometric exercises, the essay turns to a particularly suggestive example: the structure of DNA.

    DNA molecules can form coiled, knotted, and interlinked structures. During real biological processes, enzymes known as topoisomerases can temporarily cut a strand, allow changes in the molecule’s entanglement, and then reconnect it. Knot theory and other topological tools are useful precisely for describing aspects of these configurations.

    Here, the old metaphor of “cutting and gluing” ceases to be merely a pedagogical image. The connectivity of a molecular structure can undergo physically real modifications.

    Why the Example Matters

    The article uses DNA as an epistemological illustration: changing a quantity — length, twist, distance — is not the same thing as changing the constitutive relations of a structure. When a connection is broken and recomposed, the kind of transformation is qualitatively different.

    From the dialectical perspective developed by the author, this case illustrates a more general idea: systems possess relatively stable properties, but that stability exists within processes of transformation. Some transformations may accumulate or reach a point at which a qualitatively different organization emerges.

    09 · The Thesis in Perspective What the Essay Proposes — and What It Does Not

    It is useful to distinguish carefully between two levels. The first is strictly mathematical: topological spaces, continuity, homeomorphisms, invariants, and dimension have formal definitions and results that do not depend on accepting a Marxist philosophy.

    The second level is interpretive. The article argues that the history and conceptual structure of topology can be understood particularly fruitfully through dialectical-materialist categories: structure, relation, transformation, invariance, essence, form, and qualitative change.

    In other words, the argument is not that a theorem of topology can be derived from Marx. Nor does it claim that a homeomorphism and a dialectical contradiction are literally the same concept. The project is to seek a structural correspondence: to show that certain relations formally discovered by mathematics may acquire epistemological meaning when placed within a more general conception of change and structure.

    Seen in this way, the historical reconstruction is not decorative. Cantor challenges an inherited intuition about dimension; Peano shows that continuity can produce phenomena that intuition did not anticipate; Poincaré attempts to redefine the problem; Brouwer introduces new proofs and new abstractions; Hausdorff helps systematize the language. The modern concept emerges through conflicts, reformulations, and successive theoretical developments.

    That historical movement is precisely what makes the article’s dialectical reading attractive: a scientific theory does not appear finished from the outset. Its categories develop through concrete contradictions that force earlier concepts to be revised, some of their elements preserved, and others abandoned.

    The Question That Remains

    Perhaps the most powerful intuition a non-specialist reader can take away is this: knowing something does not consist solely in measuring its visible properties. We may also ask which relations make that thing the structure it is, which modifications it can undergo without ceasing to preserve that structure, and what kind of transformation would have to occur for a different structure to emerge.

    Topology provides an extraordinarily precise mathematical language for one version of that question. Gómez Julián’s essay proposes that dialectical materialism, in turn, provides a way of interrogating its philosophical meaning.

    · · ·

    The journey that begins with bridges, points, and lines thus ends with a much broader question. What does it mean to say that something remains “the same” while changing? Which transformations are accidental with respect to a structure, and which alter what constitutes it? How can continuity be distinguished from rupture?

    These are mathematical questions when we speak of topological spaces. But they are also questions that reappear, in different forms, when we study physical, biological, economic, or social systems. The philosophical wager of the article is precisely that this recurrence should not be treated as a merely verbal coincidence: it deserves to be investigated as a correspondence among forms of structure, transformation, and invariance.

  • BAYES ESTIMATORS, CLUSTER ANALYSIS, AND GAUSSIAN MIXTURES

    BAYES ESTIMATORS, CLUSTER ANALYSIS, AND GAUSSIAN MIXTURES

    Probability · Statistical Theory · Unsupervised Learning

    From Bayes Estimators to Gaussian Mixtures

    A guided reading of José Mauricio Gómez Julián’s 2020 study on Bayesian decision theory, cluster analysis, density estimation, the EM algorithm, and model-based clustering in R.

    Research conducted in 2020 José Mauricio Gómez Julián Approx. 15-minute read

    Statistics often becomes difficult not because any single idea is impossible to understand, but because several ideas must suddenly be held together at once. A probability distribution leads to an estimator; an estimator leads to a loss function; a loss function leads to an optimization problem; an optimization problem leads to an algorithm; and the algorithm finally produces something that looks deceptively simple on the screen: a handful of clusters. Gómez Julián’s 2020 study, On Bayes Estimators, Cluster Analysis and Gaussian Mixtures, is essentially an attempt to reconstruct that entire chain of reasoning before asking R to perform the calculation.

    Its destination is model-based cluster analysis with finite Gaussian mixture models. But the paper deliberately takes the long road. Before arriving at Gaussian mixtures and the Mclust() routine, it moves through posterior probability, Bayesian estimators, loss functions, mean squared error, information criteria, data mining, machine learning, supervised and unsupervised learning, density estimation, k-means, the expectation-maximization algorithm, categorical and Dirichlet distributions, optimization, parameters and hyperparameters. The point is not simply to list definitions. It is to show why these concepts meet inside the same statistical machine.

    Reading note · what this article will do

    The original 2020 study is much broader than a conventional software tutorial. This guided reading therefore concentrates on its main statistical architecture: how Bayesian reasoning, cluster analysis, Gaussian mixture models, the EM algorithm, and BIC-based model selection fit together, and what happens when that framework is applied in R.

    · · ·

    01 · Context Why this 2020 study was written

    The immediate motivation is educational research. Gómez Julián begins from a doctoral study by Villegas Barahona concerned with academic performance, latent dimensions, directly observed student variables, CUR matrix decomposition, and the construction of a statistical model capable of supporting academic and administrative decision-making. That earlier project provides the practical problem; the 2020 study asks what statistical theory one must understand in order to follow the machinery being used.

    This matters because a statistical package can make a difficult procedure look trivial. A researcher can type a command, obtain a classification, inspect a graph, and move on. Yet the command silently presupposes answers to difficult questions. What is being estimated? What does it mean for observations to belong to different groups? What happens when those memberships are not observed? How should the number and geometry of the groups be chosen? What quantity is the algorithm maximizing? And how much uncertainty remains after a point has been assigned to a cluster?

    These are not merely programming questions. They are questions about probability, inference, geometry and decision-making. The paper’s distinctive strategy is therefore foundational: instead of treating Gaussian mixture models as a black box, it reconstructs the conceptual staircase leading to them.

    A cluster on a computer screen is the final visible result of a much longer argument about probability, hidden structure, estimation and optimization.

    Conceptual summary of Gómez Julián’s 2020 framework

    There is another reason this approach is useful outside statistics. Political scientists, economists, sociologists and public-policy researchers frequently work with populations that are heterogeneous. Countries, households, firms, voters or students may appear in one dataset while actually belonging to several statistically distinct subpopulations. If those subpopulations are not directly labelled, the analytical task is no longer simply to estimate an average. It is to infer the hidden structure that may have generated the observations.

    That is where cluster analysis and Gaussian mixture models eventually enter. But Gómez Julián begins one layer deeper: with Bayes and the logic of updating knowledge when new evidence arrives.

    · · ·

    02 · Bayesian foundations Bayes as a rule for learning from evidence

    At its simplest, Bayes’ theorem tells us how a probability should change when we acquire relevant information. Suppose we have a hypothesis \(H\) and observe some data \(D\). Bayesian updating connects four quantities:

    \( P(H\mid D)=\dfrac{P(D\mid H)\,P(H)}{P(D)} \) posterior = likelihood × prior / marginal probability of the data

    The prior, \(P(H)\), represents the state of information before the new evidence is incorporated. The likelihood, \(P(D\mid H)\), tells us how compatible the observed evidence is with the hypothesis. The denominator \(P(D)\) normalizes the calculation. The result, \(P(H\mid D)\), is the posterior: the probability conditional on having observed the new evidence.

    For readers coming from economics, there is a useful analogy. Imagine beginning with a set of beliefs about the likely position of an economy, then receiving new information about employment, inflation or production. The point of Bayesian updating is not that the old information disappears. Rather, prior information and new evidence are combined according to a precise probabilistic rule. The posterior then becomes the informational starting point for whatever decision comes next.

    The key intuition

    Bayes’ theorem is not yet a clustering algorithm. Its importance here is more fundamental: Gaussian mixture models repeatedly ask conditional-probability questions. Given an observed data point, how probable is it that the point came from component 1, component 2, component 3, and so on? Once group membership is hidden rather than directly observed, posterior probabilities become a natural language for reasoning about that uncertainty.

    This is one of the conceptual bridges that makes the paper coherent. What begins as an abstract discussion of conditional probability will later reappear in a very concrete form: each observation can carry a probability of membership in each possible Gaussian component. That is already a major difference between a Gaussian mixture model and the familiar hard assignment produced by ordinary k-means.

    03 · Statistical decision-making From posterior probability to a Bayes estimator

    Updating probabilities is only part of the story. Eventually, an analyst has to do something with the posterior distribution. A parameter must be estimated, a prediction must be produced, a model must be selected, or an observation must be assigned—perhaps provisionally—to a group.

    This is why Gómez Julián’s 2020 study moves from Bayes’ theorem into decision theory. Once several possible estimates or actions are available, the statistical problem can be expressed as a question of consequences: if the unknown quantity is really \(\theta\), what is the cost of reporting some estimate \(\hat{\theta}\)?

    That cost is represented by a loss function. The exact form of the function depends on what kinds of errors matter in the problem under study. One particularly important case, and the one emphasized in the paper, is squared-error loss:

    \( L(\theta,\hat{\theta})=(\hat{\theta}-\theta)^2 \) a larger distance between the estimate and the unknown parameter produces a disproportionately larger loss

    Squaring the error has two immediate consequences. First, positive and negative deviations no longer cancel each other. Second, large errors are penalized more heavily than small ones. The associated expected loss is therefore closely connected with the familiar mean squared error.

    \( \mathrm{MSE}(\hat{\theta}) = E_{\theta}\!\left[(\hat{\theta}-\theta)^2\right] \) mean squared error as an expected measure of estimation error

    Bayesian decision theory adds one decisive ingredient: rather than evaluating loss while treating the parameter as an unknown fixed object outside the probability calculation, the posterior distribution is used to average the possible consequences of a decision. The relevant quantity becomes the posterior expected loss.

    \( \rho(a\mid x) = \int_{\Theta} L(\theta,a)\, p(\theta\mid x)\,d\theta \) posterior expected loss: consequences averaged over current uncertainty about the parameter

    A Bayesian decision rule selects the action that minimizes this quantity. Under squared-error loss, something especially elegant happens: the optimal estimate is the posterior mean.

    \( \hat{\theta}(x) = E(\theta\mid x) = \int_{\Theta} \theta\,p(\theta\mid x)\,d\theta \) Bayes estimator under quadratic loss

    The intuition is straightforward. The posterior distribution describes what values of \(\theta\) remain plausible after observing the data. If squared distance is what we wish to minimize, then the posterior mean is the point that minimizes the average squared distance to all those possible values.

    A useful distinction

    Bayesian updating tells us how uncertainty changes after observing evidence. Bayesian decision theory tells us how to convert that updated uncertainty into an action. The first produces a posterior distribution; the second combines that posterior with a loss function.

    This distinction becomes important later. A Gaussian mixture model does not merely compute probabilities. It uses probabilities as part of an iterative estimation problem in which unknown component memberships and unknown component parameters have to be inferred together.

    The study also discusses the broader statistical idea of Bayes risk: the expected loss associated with a decision rule when uncertainty about the parameter is itself represented probabilistically. Within the decision-theoretic framework adopted in the paper, a Bayes estimator is the estimator chosen because it minimizes the relevant expected loss.

    Probability describes uncertainty; a loss function gives that uncertainty consequences.

    The bridge from inference to decision theory
    · · ·

    04 · Unsupervised learning Why clustering is different from classification

    The paper then changes scale. It moves from the estimation of an unknown parameter toward a broader machine-learning question: how can structure be discovered in a dataset when the observations do not already come with known class labels?

    This is the defining setting of unsupervised learning. In supervised learning, the training data contain an outcome or label that the algorithm is asked to reproduce or predict. A model may learn, for example, whether a loan applicant defaulted, which party a respondent voted for, or what numerical value a dependent variable took.

    In unsupervised learning there is no such answer key. The algorithm receives observations and their measured characteristics, but not a pre-existing declaration that observation 17 belongs to type A while observation 18 belongs to type B. The structure itself must be inferred from patterns in the data.

    Classification versus clustering

    In classification, classes are known during training and the model learns how to assign new observations to them. In clustering, the groups are not given in advance. The method attempts to discover a useful grouping from similarities, differences and distributional structure within the observed data.

    Gómez Julián places cluster analysis at the center of this unsupervised-learning problem. In its most general form, clustering means partitioning observations according to shared characteristics so that observations within a cluster are relatively similar and observations belonging to different clusters are relatively dissimilar.

    That formulation sounds simple, but it conceals one of the deepest difficulties in clustering: there is not always one uniquely obvious way to divide a dataset. The same cloud of points may plausibly be described as containing two broad groups, several narrower subgroups, or a hierarchy in which larger groups contain smaller ones.

    The study uses this ambiguity to distinguish two broad families. Partitional clustering divides observations into non-overlapping groups at a selected level. Hierarchical clustering, by contrast, organizes groups within groups, producing a nested structure that can often be represented as a tree.

    This is more than a technical distinction. It reminds us that a cluster is not simply an object waiting in the data to be photographed. A clustering procedure embodies a definition of what similarity means, how distance is measured, what geometry is permitted, and at what scale differences are considered important.

    To ask how many groups are in a dataset is already to ask what counts as a group.

    Why clustering is fundamentally a modelling problem

    This is particularly important for economists and political scientists. Suppose countries are represented by unemployment, literacy, poverty and public education expenditure. A clustering algorithm may detect statistically distinct configurations of those variables. But the resulting groups should not automatically be treated as substantive political or economic “types.” Statistical grouping is evidence about structure; interpretation still requires theory and knowledge of the phenomenon being studied.

    Before clustering, the paper also emphasizes the importance of data preprocessing. Outliers, different measurement scales, irrelevant variables and missing values can alter the apparent geometry of the dataset. Normalization may therefore matter when distance is central to the algorithm, while variable reduction can be useful when irrelevant dimensions obscure rather than clarify structure.

    · · ·

    05 · A first clustering model What k-means actually assumes

    To understand why Gaussian mixtures are useful, Gómez Julián first introduces one of the best-known clustering algorithms: k-means.

    The basic idea is geometric. Choose a number of groups, \(K\). Associate each group with a center, or centroid. Then assign every observation to the group whose centroid is closest. The centroids are updated from the observations assigned to them, and the process is repeated until the configuration stabilizes.

    \( \displaystyle \min_{C_1,\ldots,C_K} \sum_{k=1}^{K} \sum_{x_i\in C_k} \lVert x_i-\mu_k\rVert^2 \) the familiar k-means objective: minimize within-cluster squared distance from each observation to its centroid

    Even readers who have never implemented the algorithm can visualize its logic. Imagine placing \(K\) pins on a map. Each observation is sent to the closest pin. The pins are then moved to the centers of the observations assigned to them, and the assignment is repeated. Eventually the pins and memberships stop changing substantially.

    This procedure is powerful and computationally convenient, but its simplicity imposes a geometric structure. The distance-to-centroid logic works most naturally when clusters are compact and roughly spherical—or circular when visualized in two dimensions.

    Real datasets need not cooperate. A cluster may be long and narrow, tilted diagonally through the feature space, tightly concentrated in one direction and widely dispersed in another. Two groups can also overlap. Once these possibilities appear, distance to a single center may no longer describe the structure adequately.

    There is another limitation that becomes central to the paper. Ordinary k-means makes what is known as a hard assignment. An observation is placed in cluster 1 or cluster 2 or cluster 3. The algorithm does not naturally say: “there is a 72% probability that this observation belongs to cluster 1 and a 28% probability that it belongs to cluster 2.”

    Two limitations to remember

    The transition from k-means to Gaussian mixtures in the study is motivated by two especially important ideas: cluster geometry and uncertain membership. Gaussian mixtures can model clusters with covariance structure and can assign probabilistic, rather than purely deterministic, membership.

    This prepares the central conceptual turn of the study. Instead of thinking of a cluster merely as a collection of points around a centroid, we can think of it as a probability distribution.

    · · ·

    06 · Latent structure The central idea: a population can be a mixture

    Suppose we observe the distribution of some variable across an entire population. At first glance we see only one dataset. But what if that population is actually composed of several subpopulations generated by different statistical processes?

    This is the fundamental intuition behind a mixture model. The overall probability distribution is represented as a weighted combination of several component distributions.

    \( \displaystyle f(x_i;\Psi) = \sum_{k=1}^{G} \pi_k f_k(x_i;\theta_k) \) finite mixture model · each component has its own parameters and contributes according to its mixture weight

    Here \(G\) is the number of components. \(f_k(x_i;\theta_k)\) is the density of component \(k\), determined by its own parameters \(\theta_k\). The quantity \(\pi_k\) is the component’s mixture weight, satisfying \(\pi_k>0\) and \(\sum_{k=1}^{G}\pi_k=1\).

    The weights are important. If 70 percent of the population appears to have been generated by one component and 30 percent by another, the two component densities should not contribute equally to the overall population density. The mixture weights encode their relative prevalence.

    But the deepest feature of the model is something we do not directly observe: the component identity of each observation. The dataset contains \(x_i\), but it does not ordinarily contain an additional column supplied by nature saying: “this point was generated by Gaussian component 3.”

    Component membership is therefore a latent variable. It is hidden structure inferred from the observed data.

    Observed and latent quantities

    The observations \(x_1,\ldots,x_n\) are visible. The component labels that generated them are not. A finite mixture model therefore links observable data to unobservable group membership, while estimating the parameters and relative weight of each component.

    This is why the paper treats mixture models as naturally connected to hierarchical and latent-variable modelling. There is one level at which an observation belongs to some unobserved component, and another level at which the observed value is generated according to the probability distribution associated with that component.

    If the component distributions are Gaussian, the model becomes a Gaussian mixture model, or GMM:

    \( \displaystyle f(x) = \sum_{k=1}^{G} \pi_k\, \mathcal{N}(x\mid\mu_k,\Sigma_k) \) Gaussian mixture model · each latent group is represented by a Normal density with its own mean and covariance structure

    In one dimension, each component has a mean and a variance. In several dimensions, the mean becomes a vector \(\mu_k\), while dispersion and dependence among variables are represented by the covariance matrix \(\Sigma_k\).

    The covariance matrix is precisely what gives Gaussian mixture models their geometric flexibility. It allows one cluster to be narrow, another broad, another elongated, and another oriented along a diagonal direction in multivariate space.

    The paper therefore presents Gaussian mixtures as a probabilistic generalization of the more rigid centroid-based intuition associated with k-means. Instead of asking only which center is closest, the model asks a richer question: given the estimated distributions, how probable is it that this observation came from each component?

    \( \displaystyle P(Z_i=k\mid x_i) = \frac{ \pi_k\, \mathcal{N}(x_i\mid\mu_k,\Sigma_k) }{ \sum_{j=1}^{G} \pi_j\, \mathcal{N}(x_i\mid\mu_j,\Sigma_j) } \) posterior probability that observation i belongs to component k

    Now the earlier discussion of Bayes becomes visibly relevant. The model begins with component weights and component densities and, conditional on an observed point, calculates updated probabilities of component membership.

    An observation near the center of one component and far from all others may receive an assignment probability close to one. An observation lying in an overlapping region may receive substantial probability under two or more components. This is soft classification: the uncertainty surrounding membership is retained instead of being immediately discarded.

    A Gaussian mixture does not merely divide the data. It proposes a probabilistic account of how several hidden subpopulations could have generated the observed population.

    The core modelling idea of the 2020 study

    But we have now reached an apparent circularity. To estimate the mean, covariance and weight of each component, we would like to know which observations belong to which component. Yet determining which observations belong to which component is precisely what requires knowing those means, covariances and weights.

    Solving that circular problem is the task of one of the most important algorithms in latent-variable statistics: expectation-maximization.

    07 · Hidden information EM: learning when group membership is unknown

    The difficulty facing a Gaussian mixture model can now be stated precisely. We observe the data points, but we do not observe the component from which each point was generated. If those memberships were known, estimating the parameters of each Gaussian component would be relatively straightforward. But the memberships themselves depend on parameters that are still unknown.

    Gómez Julián’s 2020 study approaches this problem through the classical expectation-maximization algorithm, or EM, developed by Dempster, Laird and Rubin. The broader setting is estimation from incomplete data: there is information that would make the estimation problem easier, but that information is not directly observed.

    In mixture modelling, the missing piece is especially intuitive. Imagine that each row of the dataset secretly carries an additional variable saying which Gaussian component generated it. If that hidden variable were visible, we would have what can be thought of as the complete data. In reality, only the measured variables are observed.

    The missing-data interpretation

    In a Gaussian mixture model, the data point itself is observed, but its generating component is latent. EM treats this hidden information as the missing part of an otherwise more convenient statistical problem.

    The ingenious feature of EM is that it does not demand that this missing information somehow become directly observable. Instead, it alternates between two calculations. Each calculation makes the other possible.

    The E-step: estimate the hidden memberships

    Begin with some current values for the component parameters: the means, covariance matrices and mixture weights. Given those values, calculate how probable it is that each observation belongs to each component.

    These probabilities are often called responsibilities. Component \(k\) takes responsibility for observation \(i\) in proportion to how plausible that observation is under the component’s Gaussian density and how prevalent that component is in the mixture.

    \( \displaystyle \gamma_{ik} = P(Z_i=k\mid x_i,\Psi) = \frac{ \pi_k\, \mathcal{N}(x_i\mid\mu_k,\Sigma_k) }{ \sum_{j=1}^{G} \pi_j\, \mathcal{N}(x_i\mid\mu_j,\Sigma_j) } \) E-step intuition · estimate the probability that each observation belongs to each Gaussian component

    Notice what has happened. The hard, unknown statement “observation \(i\) belongs to cluster \(k\)” has been replaced with a set of probabilities. An observation may be overwhelmingly associated with one component, or it may sit in an overlapping region and divide its probability between several components.

    The M-step: update the model

    Once those expected memberships have been calculated, the algorithm turns the problem around. It now treats the probabilistic memberships produced by the E-step as information for re-estimating the parameters.

    The means, covariance matrices and mixture weights are updated so that the likelihood of the observed data increases under the newly estimated mixture.

    One EM cycle

    E-step: using the current model parameters, estimate the hidden component memberships.

    M-step: using those estimated memberships, re-estimate the model parameters by maximizing the relevant likelihood criterion.

    Then the algorithm goes back to the E-step. The new parameters imply new membership probabilities; those new probabilities imply new parameter estimates; and the process continues iteratively.

    \( \Psi^{(0)} \rightarrow \text{E-step} \rightarrow \text{M-step} \rightarrow \Psi^{(1)} \rightarrow \text{E-step} \rightarrow \text{M-step} \rightarrow \cdots \) expectation and maximization alternate until the fitted solution stabilizes

    The process stops when the parameter estimates—or equivalently the likelihood improvements—change so little that the algorithm is considered to have reached convergence.

    This makes EM a particularly elegant response to the apparent circularity encountered at the end of the previous section. We needed cluster membership to estimate the distributions, but we needed the distributions to estimate cluster membership. EM solves the problem by alternating between the two conditional tasks.

    Estimate what is hidden using the current model; then improve the model using what you have just estimated.

    The iterative logic of expectation-maximization

    There is an important qualification. EM is an optimization algorithm, not a magical guarantee that every possible starting point will lead to the globally best solution. Mixture-model likelihoods can contain multiple local optima. Initialization and model specification can therefore matter. What EM guarantees at the operational level is an iterative procedure for improving the likelihood until a stationary solution is reached.

    · · ·

    08 · Statistical geometry Why Gaussian mixtures can see ellipses

    The next step in Gómez Julián’s argument is geometric. In one dimension a Gaussian distribution is described by a mean and a variance. Move into two or more dimensions, however, and variance is no longer sufficient. The relationships among variables must also be represented.

    This is the role of the covariance matrix, \(\Sigma_k\). For component \(k\), the mean vector \(\mu_k\) determines its center, while \(\Sigma_k\) determines how the probability mass spreads through multivariate space.

    \( X\mid Z=k \sim \mathcal{N}(\mu_k,\Sigma_k) \) each Gaussian component possesses its own center and covariance geometry

    In two dimensions, contours of equal Gaussian density form ellipses. This provides an intuitive way to read covariance. A nearly circular ellipse indicates similar dispersion in different directions. An elongated ellipse indicates much greater variation along one direction than another. A tilted ellipse signals covariance between the variables.

    This is precisely where Gaussian mixture clustering becomes more flexible than the elementary geometric picture supplied by k-means. A centroid alone tells us where a cluster is centered. A covariance matrix also tells us its volume, shape and orientation.

    Think geometrically

    Two clusters may have centers that are equally far apart while still being statistically very different. One may be compact and almost circular; another may be broad and strongly elongated. Gaussian mixture models can represent this difference because the covariance matrix is part of the model.

    The mclust framework studied in the paper exploits this fact systematically. Instead of fitting only one possible covariance structure, it considers a family of Gaussian models obtained by placing different restrictions on the volume, shape and orientation of the component ellipsoids.

    Gómez Julián discusses the 14 multivariate Gaussian models available in the version of mclust studied in the 2020 research. Their compact names—such as EEE, VEV, VVI or EEV—encode restrictions on those geometric properties.

    Example Geometric idea
    EEE Equal volume, equal shape and equal orientation across components
    VEV Variable volume, equal shape and variable orientation
    VVI Diagonal covariance structure with variable volume and shape
    EEV Equal volume and shape, with orientation allowed to vary

    These codes are not decorative software jargon. They describe competing statistical hypotheses about the geometry of the hidden groups. Should all clusters have the same spread? May one be larger than another? Must their ellipses point in the same direction? Is a diagonal covariance matrix enough, or does the data require rotated ellipsoids?

    Seen this way, model-based clustering is doing more than deciding where to draw boundaries. It is comparing alternative generative descriptions of the data.

    In model-based clustering, the shape of a cluster is not an afterthought. It is part of the hypothesis being estimated.

    Covariance as statistical geometry
    · · ·

    09 · Model selection BIC and the problem of choosing a model

    Gaussian mixtures create a new problem precisely because they are flexible. We may fit different numbers of components, and for each number of components we may consider different covariance structures. Which model should be preferred?

    Maximized likelihood alone is not enough. Adding parameters usually gives a model more freedom to accommodate the observed data, so raw fit can improve simply because the model has become more complicated. If complexity is never penalized, the procedure is pushed toward increasingly elaborate specifications.

    This motivates the Bayesian Information Criterion, introduced by Gideon Schwarz and discussed at length in the 2020 study. In one common notation,

    \( \mathrm{BIC} = -2\log \hat{L} + k\log n \) conventional minimization form · fit is balanced against a penalty that increases with model complexity

    Here \(\hat{L}\) is the maximized likelihood, \(k\) is the number of estimated parameters and \(n\) is the sample size. The first term rewards fit; the second penalizes additional parameters.

    An equivalent sign convention is often written so that larger values are preferred:

    \( \displaystyle \log \hat{L} – \frac{k}{2}\log n \) Schwarz’s maximization form · the same fit-versus-complexity logic expressed with the opposite orientation
    A practical warning about signs

    Readers sometimes see “choose the smallest BIC” in textbooks and then encounter mclust output where the preferred model has the largest BIC value. This is a matter of convention. The criterion can be written with opposite signs. What matters is using the convention adopted by the software or source consistently.

    In mclust, BIC therefore becomes the mechanism for comparing combinations of component number and covariance parametrization. The software can fit a collection of candidate Gaussian mixture models and compare them rather than forcing the researcher to stipulate one geometry in advance.

    Conceptually, this is a competition among explanations. A one-component model says that a single Gaussian population is sufficient. A two-component model says that two latent subpopulations provide a better account after accounting for the additional parameters. A five-component model makes an even more elaborate claim. BIC asks whether the gain in likelihood is large enough to justify that extra complexity.

    What BIC is doing in this paper

    BIC acts as a bridge between estimation and model selection. EM estimates the parameters of a candidate Gaussian mixture. BIC helps decide which candidate structure—among different numbers and geometries of components—is comparatively preferable.

    This distinction is essential. EM does not, by itself, answer every modelling question. Given a specified mixture structure, it provides a way to estimate its parameters. Model selection operates at another level: it compares alternative structures.

    The result is a layered procedure. First define candidate probability models. Then estimate them. Then compare them. Finally inspect the resulting classification and ask whether the statistical structure is substantively meaningful.

    \( \text{candidate models} \rightarrow \text{EM estimation} \rightarrow \text{BIC comparison} \rightarrow \text{selected clustering structure} \) the model-based clustering workflow developed toward the applied section of the study

    We are now ready for the final step of the 2020 investigation: seeing what this machinery actually produces in R. Gómez Julián closes the substantive analysis with two types of application. The first uses the canonical Iris dataset; the second moves into social and economic data from the World Bank, combining indicators of education expenditure, literacy, unemployment and poverty.

    10 · Applied examples in R From Iris flowers to World Bank indicators

    After more than one hundred pages of theoretical preparation, Gómez Julián’s 2020 study finally lets the statistical machinery run. This last substantive section is useful precisely because the preceding discussion changes the meaning of what would otherwise look like a few lines of R code. By this point, a call to Mclust() is no longer merely a software command. It invokes finite Gaussian mixtures, latent membership, maximum-likelihood estimation through EM, alternative covariance geometries and BIC-based model comparison.

    The paper provides two kinds of illustration. First comes the canonical Iris dataset distributed with R. Then the analysis moves to a dataset assembled from World Bank indicators, bringing the method into a setting much closer to economics, political science and public-policy research.

    How to read the output

    When Mclust() reports a model such as VEV, EEE or VVI, it is describing the covariance structure selected for the Gaussian components. When it reports a number of components, it is describing the number of mixture components preferred by the model-selection procedure among the candidates fitted.

    The Iris example

    The first application uses the four familiar quantitative variables in the Iris dataset. Gómez Julián runs:

    mod1 <- Mclust(iris[,1:4])
    summary(mod1)
    Gaussian model-based clustering of the four measured Iris variables

    The reported solution is a VEV Gaussian finite mixture with two components. The 150 observations are partitioned into clusters containing 50 and 100 observations, respectively. The reported log-likelihood is \(-215.726\), while the output gives a BIC of \(-561.7285\) and an ICL of \(-561.7289\).

    Dataset Selected structure Clustering
    Iris VEV · 2 Gaussian components 50 / 100 observations

    The accompanying plots make visible the two layers of the procedure. One panel compares BIC values across candidate covariance models and different numbers of components. Another displays the resulting classification across pairs of the measured variables. The graph is therefore not merely showing clusters after the fact: it also gives the reader a view of the model-selection problem that produced them.

    This example is deliberately straightforward. Its role is to show that the theoretical discussion of mixture densities, EM, covariance structure and BIC can be condensed operationally into a remarkably short piece of R code.

    · · ·

    A social-science example using World Bank data

    The second application is more directly connected to the concerns of economists and policy researchers. Gómez Julián constructs an example using World Bank data and four variables for 2018:

    Variables used in the 2020 application

    Public expenditure on education as a percentage of GDP; unemployment as a percentage of the total labour force; poverty incidence according to the national poverty line; and the adult literacy rate for persons aged 15 and above.

    The R workflow imports the separate datasets, selects the 2018 observations, joins them by country, removes rows for which the required combination contains missing values, and then applies Mclust() to the resulting numerical variables.

    When all four indicators are considered together, only 13 complete observations remain in the dataset used by the code. The reported model is EEE with nine components: ellipsoidal Gaussian clusters with equal volume, equal shape and equal orientation.

    \( n=13,\qquad G=9,\qquad \text{model}=\mathrm{EEE} \) four-variable World Bank example reported in the study

    The output reports a log-likelihood of \(-77.62892\), 54 degrees of freedom, BIC \(-293.7651\) and ICL \(-293.7727\). The component counts are extremely small: the nine clusters contain respectively 1, 2, 2, 1, 1, 1, 2, 2 and 1 observations.

    Gómez Julián then repeats the model-based clustering exercise using smaller combinations of variables. This is particularly revealing because the available sample size changes sharply depending on which World Bank indicators must be simultaneously observed.

    Variables n Model Components BIC
    Education expenditure + literacy 39 XXI 1 −475.3468
    Education expenditure + unemployment 71 VVI 2 −679.1997
    Education expenditure + poverty 17 EEV 5 −193.6785

    The contrast is striking. For public education expenditure and adult literacy, the fitted solution contains only one component among 39 complete observations. For education expenditure and unemployment, the selected model is VVI with two components, containing 49 and 22 observations. For education expenditure and poverty, the result is an EEV specification with five components, whose sizes are 3, 3, 4, 4 and 3.

    These examples demonstrate something that can easily disappear when one speaks abstractly about “the number of clusters.” The number of components is not an intrinsic number attached forever to a set of countries. It depends on the variables being modelled, the available observations, the candidate covariance structures and the statistical criterion used to compare those models.

    An important inferential boundary

    The output reported in this section is cluster analysis. It describes statistical structure found by the fitted Gaussian mixture models. By itself, such an exercise does not establish that education expenditure causes unemployment, literacy or poverty, nor does it estimate the magnitude of a causal effect. Those would be different inferential questions requiring a different research design.

    This distinction is especially valuable for policy analysis. A cluster can reveal that some countries occupy similar regions of a multivariate statistical space. That may motivate substantive investigation. It does not, by itself, explain historically or causally why those countries occupy that region.

    The four-variable result deserves similar care. Nine components from only thirteen complete observations is exactly the kind of output that should be read together with the sample size and model complexity rather than reduced to the phrase “nine types of countries.” The study reports the statistical fit; substantive interpretation requires returning from the model to the empirical object being studied.

    · · ·

    11 · The larger lesson What the 2020 study is really teaching

    The most important feature of Gómez Julián’s investigation may be its refusal to begin with the software. The paper could have been a short tutorial showing how to call Mclust(), inspect BIC and plot a classification. Instead, it constructs a long conceptual route from probability and estimation to the final clustering output.

    That route matters because the elements are genuinely connected. Bayesian reasoning introduces conditional probability and updating. Decision theory explains how probability distributions can be connected to estimators and loss. Unsupervised learning introduces the problem of discovering structure without known labels. Cluster analysis gives that problem a statistical form. Mixture models reinterpret an apparently homogeneous population as the superposition of latent subpopulations. Gaussian mixtures give those subpopulations flexible probabilistic geometry. EM estimates models whose membership information is hidden. And BIC provides a way to compare competing specifications.

    \( \text{Bayes} \rightarrow \text{estimation} \rightarrow \text{latent variables} \rightarrow \text{mixtures} \rightarrow \text{EM} \rightarrow \text{BIC} \rightarrow \text{clustering} \) a compressed map of the conceptual route reconstructed in the 2020 study

    One can also read the paper as an argument for understanding statistical methods structurally. A model is not simply an equation. It includes assumptions about what is observable, what is latent, which probability family describes the data, how parameters are estimated, what geometries are permitted and how competing specifications are compared.

    Gaussian mixture models make that point unusually visible. The observable cloud of data is only the surface. Beneath it lies a proposed generative structure: component distributions, latent memberships, mixture weights, means and covariance matrices. The analyst does not observe this machinery directly. It is inferred.

    The visible dataset is the starting point. The statistical model is a hypothesis about the hidden structure capable of producing it.

    A central intuition running through the study

    This is also why the difference between hard and soft classification is so significant. Saying that a country, person or flower belongs to “cluster 2” suppresses information. A Gaussian mixture can instead preserve the fact that an observation may lie near the frontier between several plausible components. Probability makes ambiguity measurable.

    Likewise, covariance transforms clustering from the simple idea of distance from a center into a richer account of statistical geometry. Groups may differ not only in location but also in dispersion, shape and orientation. And BIC reminds us that greater flexibility comes at a cost: a model must earn its additional complexity through improved fit.

    For economists, econometricians and political scientists, perhaps the most transferable lesson is therefore methodological. If a population may contain qualitatively different statistical regimes, forcing every observation into a single homogeneous distribution can conceal structure. Mixture models offer one formal way of asking whether the aggregate pattern may instead be generated by several latent components.

    But the converse warning is equally important. Discovering a statistically preferred partition does not relieve the researcher of the obligation to understand the real phenomenon. A component is a component of a statistical model. Whether it corresponds to a meaningful social class, institutional regime, developmental configuration, biological population or merely a feature of the available sample must be established with substantive knowledge and further evidence.

    In one sentence

    Gómez Julián’s 2020 investigation is a theoretical guided tour of the ideas required to understand how Gaussian finite mixture models can discover latent structure in unlabeled data, how EM estimates that structure, and how model-selection criteria such as BIC help decide which probabilistic representation to retain.

    · · ·

    12 · Closing perspective Statistics before software

    There is a useful reversal at the heart of this study. Modern statistical computing encourages us to begin with a function and discover afterward what it does. Gómez Julián’s 2020 text proceeds in the opposite direction: first reconstruct the mathematical and statistical concepts, then approach the function.

    That choice makes the paper unusually broad. Posterior probability, Bayes estimators, loss functions, BIC, data mining, machine learning, clustering, density estimation, vector quantization, k-means, EM, categorical and Dirichlet distributions, optimization, parameters, hyperparameters and Gaussian finite mixtures all appear because the final procedure stands at the intersection of those ideas.

    For a technically trained reader, the value of this route is that it exposes the architecture hidden beneath a familiar R command. For a reader from political science, economics or another applied field, it offers something equally useful: an intuitive path into a method that otherwise arrives wrapped in matrix algebra and probability notation.

    And the practical lesson is simple. When an algorithm reports that the data contain one group, two groups or five, the interesting question is not merely what did the software return? It is: what statistical model made that answer possible, what assumptions gave the groups their shape, and what kind of statement about reality is the result actually capable of supporting?

    Good statistical practice begins where the automatic output ends: with an attempt to understand what the model has actually measured.

    Final reflection on Gómez Julián’s 2020 study
    José Mauricio Gómez Julián · 2020
    On Bayes Estimators, Cluster Analysis and Gaussian Mixtures: A General Theoretical Analysis of densityMclust in R and Statistical Theory
  • ON THE NEGATIVE BINOMIAL 2 DISTRIBUTION

    ON THE NEGATIVE BINOMIAL 2 DISTRIBUTION

    Probability · Count Data · Hierarchical Models

    When Counts Refuse to Behave: Understanding the Negative Binomial II as a Measurement Instrument

    A guided reading of José Mauricio Gómez Julián’s 2020 essay Some Reflections on the Negative Binomial Distribution II as a Measurement Instrument—tracing the argument from geometric series and probability mixtures to overdispersion, latent heterogeneity, and simulation in R.

    Mauricio Gómez Julián · Theoretical & Applied Probability · Approx. 15-minute read
    Reading note. This essay explains the paper on its own terms while keeping the mathematics technically precise. Where a qualification is needed—especially in the passage from an exponential mixing distribution to the general Negative Binomial II—it is marked explicitly rather than silently altering the paper’s argument.

    The question behind the paper

    Count data appear everywhere. An economist counts firm failures, strikes, defaults, patents, accidents, or entries into a market. A political scientist counts protests, cabinet changes, violent events, legislative vetoes, or international disputes. A biologist counts surviving organisms, mutations, infections, or offspring. The elementary model for many such problems is the Poisson distribution. But real counts often fluctuate more than a Poisson model allows. Their variance is larger than their mean: the data are overdispersed.

    The Negative Binomial II—usually abbreviated NB2—is one of the central statistical instruments for precisely that situation. Yet Gómez Julián’s paper is not content to present the NB2 probability mass function, list its moments, and move on. Its organizing question is broader: what kind of object is a probability distribution, where does it come from, what other distributions does it contain or presuppose, and what is gained when we understand its construction rather than merely its final formula?

    The paper therefore has two explicit axes. The first, and more important one, studies NB2 as the outcome of a wider theoretical structure involving hierarchical models and probability mixtures. The second studies NB2 as an individual distribution—its form, interpretation, moments, and practical use, including hand calculations and R. The intended unity between these axes is philosophical as well as mathematical: probability distributions are treated as instruments for measuring natural and social phenomena under uncertainty, and the paper argues that their history, formal structure, scientific interpretation, and application should not be torn apart.

    The distribution is easier to understand when we see not only the finished formula, but also the process that produces it.
    · · ·

    Why begin with the geometric series?

    The paper begins surprisingly far away from count-data regression: with the geometric series, its historical roots, and its relation to the binomial theorem. This is deliberate. Gómez Julián wants the reader to see the Negative Binomial II as part of a mathematical genealogy rather than as a formula that appeared fully formed.

    The route is roughly this: the geometric series provides a simple infinite expansion; differentiation exposes a recurring combinatorial pattern; that pattern is used to motivate the binomial expansion; and replacing the ordinary exponent with a negative one leads to the negative binomial series. The point is not merely algebraic. The paper repeatedly emphasizes the movement from simpler structures to more general ones and from one family of mathematical objects to another.

    1 + x + x2 + x3 + ··· = 1 / (1 − x),   |x| < 1 binomial expansion negative binomial series The paper’s algebraic genealogy in compressed form

    For a nontechnical reader, the important idea is simple: a probability distribution can be understood through the transformations and relationships that generate it. The paper later reinforces this visually with a large network diagram of probability distributions, reproduced from ProbOnto, in which distributions are connected by transformations, limiting relations, and special cases. NB2 is therefore presented as one node in a densely connected mathematical ecology, not as an isolated technique.

    Probability as a measurement problem

    Before building the NB2, the paper stops to ask what “probability” means. This is not a decorative philosophical detour. Gómez Julián’s position is that formal probability calculus and philosophical interpretation cannot be completely divorced, because statistical conclusions depend on what we think probabilities are measuring and on how the scientific problem is conceptualized.

    The paper adopts an explicitly objective and dialectical-materialist orientation. Randomness is treated primarily as an epistemological condition: events appear random because their causes are unknown, too numerous, or too complicated to represent completely. At the same time, the Kolmogorov axioms provide the formal mathematical framework that makes probability calculations coherent. The paper also draws on objective Bayesianism to argue that degrees of belief should be constrained by evidence, scientific theory, and the probability calculus rather than reduced to arbitrary personal opinion.

    Why this matters statistically

    A model is not chosen only because its formula is convenient. The scientific description of the process determines which random variables, conditional relationships, latent quantities, and parameterizations are meaningful. That principle becomes concrete once the paper turns to hierarchical models.

    A family, not an isolated formula

    The paper next introduces the distributions needed for its construction: Bernoulli, Binomial, Poisson, and Exponential. Each plays a distinct role.

    Distribution Plain-language role Role in the paper
    Bernoulli One trial with two possible outcomes. The elementary unit from which repeated success/failure experiments are built.
    Binomial Counts successes in a fixed number of Bernoulli trials. The first level of the hierarchical construction.
    Poisson Counts occurrences when events arrive with a given mean rate. Makes the number of opportunities or events itself random.
    Exponential Models waiting time or positive continuous variation associated with a Poisson process. Introduces variation in the Poisson rate across observational units.

    This sequence already contains the paper’s methodological intuition. A complex phenomenon can be decomposed into simpler probabilistic stages, each corresponding to a different part of the scientific story. Instead of forcing all uncertainty into a single flat formula, a hierarchical model lets uncertainty enter at more than one level.

    Hierarchies, mixtures, and latent variables

    A hierarchical model specifies variables and parameters in stages. A mixture model appears when a parameter in one probability distribution is itself treated as a random quantity governed by another distribution. The parameter that seemed fixed at the lower level becomes variable at the higher level.

    That is the key conceptual move in the paper. It is also why Gómez Julián brings in the language of latent variables: some of the forces producing observed variation may not be directly observed, but their consequences can still be represented probabilistically. In applied work, this is familiar. Two factories, municipalities, firms, hospitals, or individuals may face different underlying event rates even if we initially write one common Poisson equation for all of them.

    The paper links this mathematical construction to the Hegelian distinction between Being-in-itself and Being-for-itself. Stripped of the philosophical vocabulary, its statistical meaning is fairly intuitive. Studying NB2 “in itself” means studying the wider process and network of relationships from which it emerges. Studying NB2 “for itself” means taking the marginalized distribution as a distinct object and examining its own formula, parameters, moments, and applications.

    The first mixture: Binomial inside Poisson

    The first major construction uses a biological example. Imagine an insect that lays many eggs. Conditional on a mother having laid Y eggs, each egg survives independently with probability p. The number of surviving eggs, X, is therefore Binomial. But instead of fixing the number of eggs Y, the paper lets it vary according to a Poisson distribution with mean λ.

    X | Y ~ Binomial(Y, p)
    Y ~ Poisson(λ) First hierarchical model

    Marginalizing means asking for the distribution of X after summing out the intermediate variable Y. Algebraically, we combine all the possible values of Y, weighted by how probable each one is. The result is elegant:

    X ~ Poisson(λp) After marginalizing over Y

    In modern probability language, this is a version of Poisson thinning. If the total number of opportunities is Poisson and each opportunity independently survives with probability p, then the surviving count is also Poisson, with its mean reduced from λ to λp. The paper also derives the same conclusion through iterated expectations:

    E[X] = E{E[X | Y]} = E[pY] = pλ

    For the paper, this is more than a computational trick. It shows how a hierarchical process that appears to contain two random layers can be “compressed” into a simpler marginal law without erasing the scientific interpretation that motivated the hierarchy.

    When the Poisson rate itself varies

    The next step introduces another level of heterogeneity. Suppose there are many insect mothers, and the Poisson mean is not the same for every mother. The paper now treats the rate itself as random:

    X | Y ~ Binomial(Y, p)
    Y | Λ ~ Poisson(Λ)
    Λ ~ Exponential(β) Three-stage hierarchy in the paper

    The statistical intuition is important. Variation does not occur only in the observed count; it can also occur in the underlying rate that generates the count. Once the rate differs across observational units, the final distribution becomes more dispersed than a single-rate Poisson model. This is precisely the kind of latent heterogeneity that makes the negative binomial family useful in economics, epidemiology, demography, political science, and many other count-data settings.

    The paper shows that the mean of the full hierarchy can be obtained by repeatedly conditioning and averaging, arriving at E[X] = pβ under its parameterization. It then integrates out the random Poisson rate and identifies the resulting expression with the negative binomial form.

    Technical qualification added for accuracy

    There is an important distinction here. An Exponential distribution is a Gamma distribution with shape parameter equal to 1. Therefore, a Poisson–Exponential mixture produces the geometric distribution, which is a special case of the negative binomial with r = 1. The general NB2 with arbitrary dispersion parameter r arises from a Poisson–Gamma mixture. Thus, the paper’s core intuition—random heterogeneity in the Poisson rate generates negative-binomial-type overdispersion—is correct, but the fully general NB2 requires the Gamma mixing family rather than the Exponential distribution alone.

    This qualification actually sharpens the paper’s broader message. The geometric distribution, the exponential distribution, the gamma family, Poisson processes, and the negative binomial are not unrelated objects. They sit inside a network of special cases and mixture relationships. The more general Gamma mixing distribution preserves the same hierarchical logic while extending the model beyond the r = 1 case.

    · · ·

    The NB2 “for itself”: what the finished distribution tells us

    Once the hierarchy has been “compressed,” the paper turns to the Negative Binomial II as an object in its own right. One common parameterization writes the probability of observing x failures before the r-th success as:

    P(X = x) = Γ(x + r) / [Γ(r) Γ(x + 1)] · pr · (1 − p)x,   x = 0, 1, 2, …

    Here, p is the success probability and r is the target number of successes. The distribution answers a reversed version of the familiar Binomial question. The Binomial asks: how many successes occur in a fixed number of trials? The negative binomial asks: how many failures occur before a fixed number of successes is reached?

    The paper’s most important statistical property appears in its first two moments. With q = 1 − p:

    μ = E[X] = rq / p
    Var(X) = rq / p2 = μ + μ2/r

    That final equality is the bridge to modern count-data econometrics. The Poisson distribution imposes Var(X) = μ. NB2 allows:

    Var(X) = μ + αμ2,   where α = 1/r

    In other words, variance can grow faster than the mean. The smaller r is—or, equivalently, the larger the heterogeneity parameter α is—the more dispersion the model permits. As heterogeneity vanishes, NB2 approaches the Poisson benchmark. This is why the paper presents NB2 as a more flexible instrument for count data when the Poisson equality between mean and variance is too restrictive.

    Feature Poisson Negative Binomial II
    Mean μ μ
    Variance μ μ + αμ²
    Extra heterogeneity Not separately modeled Captured through α (or r)
    Typical use Equidispersed counts Overdispersed counts

    For an econometrician, this variance function is often the fastest route to understanding NB2. For a broader reader, an intuitive translation is enough: NB2 expects the world to be more uneven than a simple Poisson process. Some units have persistently higher event rates than others; unobserved conditions vary; clusters form; the same mean can coexist with much wider dispersion.

    The maintenance example: counting failures before the fourth alarm

    The paper gives a concrete industrial interpretation. Imagine fixed capital—a machine—producing parts. A part is either satisfactory or defective. A monitoring system treats a defective part as the event of interest because it signals the need for maintenance. Assume independent Bernoulli trials and a constant defect probability.

    Gómez Julián sets r = 4, p = 0.005, and asks for the probability of observing exactly x = 100 failures before the fourth success, using the negative binomial mass function. Substitution gives:

    P(X = 100 | r = 4, p = 0.005) ≈ 0.000067

    The number is tiny—about 6.7 chances in 100,000. The statistical interpretation is not that “100 failures happen and then four successes happen” as two separate blocks. Rather, among an ongoing sequence of independent trials, exactly 100 non-events occur before the fourth event of interest is reached.

    The paper checks the result manually, with a Texas Instruments calculator, and with R. It then plots the corresponding distribution, illustrating how a small success probability pushes substantial probability mass toward relatively large counts before the required number of successes is accumulated.

    What the R simulation is trying to show

    The final applied part of the main text shifts from evaluating a probability to generating pseudo-random data. The paper constructs a custom card-drawing experiment: repeatedly sample from a 52-card deck until a specified rank appears, record how many draws were required, and repeat the experiment many times. It then compares the histogram produced by that “from first principles” counting procedure with a histogram generated by R’s built-in rnbinom function.

    The figures for 50, 100, 150, and 200 repetitions show the same qualitative pattern: a strongly right-skewed count distribution with many small values and a long tail. The pedagogical purpose is clear. Software is not magic. A built-in random generator is implementing a probabilistic structure that can also be approximated through an explicit sequence of elementary trials.

    A convention to watch in R

    R’s rnbinom convention counts the number of failures before a specified number of successes. A hand-built routine that counts the total number of draws including the successful draw differs by one when size = 1. For an exact one-to-one comparison, the manual routine and the software call should use the same counting convention. This does not erase the pedagogical point of the simulation, but it matters for exact numerical equivalence.

    The broader lesson is useful far beyond R. Simulation can reveal what a distribution means operationally: not merely how its formula looks, but what repeated mechanism would generate data with that shape.

    NB1, Bayes, and the annexes: why the paper keeps widening the frame

    The annexes extend the same relational approach. The paper distinguishes NB2 from NB1, emphasizing that different negative-binomial parameterizations answer slightly different counting questions. In the NB1 presentation used there, the random variable is the total number of Bernoulli trials needed to reach r events of the chosen type; in the NB2 presentation, the random variable is restricted to the number of failures before those r successes.

    The paper also returns to conditional probability, total probability, inverse probability, and objective Bayesianism. This may seem far removed from overdispersed count data, but it serves the same philosophical thesis: statistical formulas must be understood through the relationships they encode. Conditional probability is not merely a ratio; it represents a structured dependence between events. Bayesian updating is not merely algebra; it connects prior knowledge, evidence, and posterior assessment.

    Finally, the paper gives a more general NB2 expression in terms of the Gamma function and reports estimators based on the first two sample moments. This again links the abstract distribution to empirical work: the population parameters acquire meaning only because sample information provides a route to estimation.

    What should we take away?

    Gómez Julián’s paper is best read as an extended argument against treating the Negative Binomial II as a black-box formula. Its distinctive contribution is not a new NB2 estimator or a new regression algorithm. It is an attempt to reconstruct the distribution through several layers at once: historical, algebraic, probabilistic, computational, applied, and philosophical.

    For the nontechnical reader, the central statistical lesson can be stated in one sentence: when counts vary more than a single-rate Poisson model permits, the extra variability can often be understood as heterogeneity in the underlying event rate, and the negative binomial family provides a natural way to represent it.

    For the econometrician, the key signature is the NB2 variance function, Var(Y) = μ + αμ². For the mathematician, the paper is an invitation to follow the transformations linking geometric series, binomial expansions, gamma functions, conditional distributions, and marginalization. For the philosopher of science, its central claim is methodological: the formal instrument, the scientific object, the history of the instrument, and the interpretation of uncertainty should be studied in relation rather than isolation.

    And for the applied researcher, perhaps the most useful question is the simplest one: what process would have to be operating for this distribution to be a sensible measurement instrument? Once that question is asked, NB2 stops being just a convenient correction for overdispersion. It becomes a hypothesis about how heterogeneity enters the data-generating process.

    A distribution is most informative when its probability law and its generating story tell the same scientific story.
    In one compact map

    Bernoulli gives the elementary success/failure trial; Binomial aggregates such trials at a fixed size; Poisson makes the number or rate of occurrences stochastic; Gamma heterogeneity lets that Poisson rate vary across units; marginalizing the latent rate yields the Negative Binomial II, whose variance can exceed its mean.

    This explainer follows the architecture and substantive aims of Gómez Julián’s paper while separating the paper’s own philosophical framing from the technical qualifications added here for mathematical precision.

    Read the Original Paper ↗
    Editorial-academic layout · Playfair Display · Lora · DM Mono
  • The Shape of a Crisis: A General Theory of Capitalist Cycles

    The Shape of a Crisis: A General Theory of Capitalist Cycles

    Thesis Release · Political Economy

    The Shape of a Crisis

    A general theory of the cycles of the dynamics of the capitalist system in the long run — now available in English

    Every few years the same story is told twice. First, that the economy has entered a new era in which the old rules no longer apply. Then, some months later, that what happened was an accident: a shock, a bubble, a virus, a war. Both tellings share a premise so quiet that it is rarely examined — that the rise and the fall are separate events, and that a good theory of the good years need not be a theory of the bad ones.

    The thesis released today argues the opposite, and then goes to some length to measure it. The boom and the crisis are not two phenomena but two moments of one: the crisis of overproduction is the mechanism by which capitalism restores the conditions of an accumulation that its own success had eroded. Devaluation clears the field; new methods of production are introduced under duress; profitability recovers on the ruins. The recovery is not the negation of the crisis. It is its product.

    That claim is old. What is new here is the attempt to make it decidable — to state it in a form that quarterly data on the United States economy between 1992 and 2024 could have contradicted, and then to check whether they do.

    Three questions, and why the order matters

    The investigation is organised around one general objective — to analyse the long-run cyclical behaviour of U.S. capitalism in the light of the dominant economic theories — and three specific ones, asked strictly in this order:

    • Which theory explains and predicts best? Not which is most elegant, or most widely taught, but which survives being pointed at the data.
    • Which factors generate the cycle? Economic and extra-economic alike — the thesis refuses in advance to treat wars and monetary policy as noise sitting outside a clean economic mechanism.
    • By which rules do those factors interact? A list of causes is not a theory. The theory is in the grammar that binds them.

    The order is not decorative. A great deal of applied economics answers the third question with machinery borrowed from a theory it never subjected to the first. Here the selection of the framework is itself a result, defended before it is used.

    Five families of an old argument

    Before measuring anything, the thesis maps the terrain. Economic thought on the cycle is sorted into five groups: the pre-Kondratieff non-heterodox schools; the Kondratieff school; the post-Kondratieff marginalist and neoclassical schools; the heterodox schools; and the historiographic vision of long waves, which reads the cycle through the archives rather than through the equations.

    With that map in hand, three long-running disputes are adjudicated rather than summarised. Does the crisis originate in overproduction or in underconsumption? Is a sustained expansion of credit a symptom of recovery, or of the exhaustion of the conditions that made recovery possible? Is there really an inverse relation between inflation and unemployment, or is the appearance of one an artefact of the precariousness of the labour market? Each is answered, and each answer carries consequences later, when the model is specified.

    A framework that states its own conditions of failure

    A substantial part of the theoretical apparatus is devoted to a materialist characterization of the dialectical method: its fundamental categories, a Marxist ontology built from a metalogical gnoseology, and an explicit treatment of verification, falsification and decidability. The purpose is unglamorous and indispensable — to fix, in advance, which propositions of the theory are empirically decidable and which are interpretive. Without that boundary, no amount of subsequent statistics can tell you what has been tested.

    Ten dials, seven of them internal

    The empirical core is a Bayesian generalized linear model of the growth of U.S. real output, estimated with Hamiltonian Monte Carlo and cross-validated against machine-learning and deep-learning competitors. It retains thirteen coefficients across ten factors. Seven are economic:

    FactorWhat it registers
    Net Average Rate of Profit (ARoP)The central variable of the accumulation process, and the one whose long-run tendency the theory predicts.
    Elasticity of the gross rate of surplus value to the average organic composition of capitalHow the exploitation of labour power responds when the technical structure of capital changes.
    Non-residential fixed investmentThe pace of accumulation in the productive sector; the hinge between boom and crisis.
    Inventory-to-sales ratioThe gap between producing value and realising it on the market.
    S&P 500Financialization, entering through a natural cubic spline with three degrees of freedom.
    Non-financial private sector creditThe credit system as the accelerator and the brake, splined with two degrees of freedom.
    Capitalist R&D spendingThe innovative impulse; the second largest coefficient in the model.

    And three are extra-economic: military spending (splined with three degrees of freedom), the federal surplus or deficit, and the effective federal funds rate. Their presence is not a concession to realism. It follows from the argument that an imperial economy counteracts the tendency of its own profit rate to fall by means that are not internal to its national accounts.

    The Average Rate of Profit carries the fourth largest coefficient of the thirteen — behind only the intercept, R&D spending, and one basis function of the splined S&P 500. The conclusion the author draws from its behaviour is worth quoting in substance: what is favourable to the global process of capital accumulation is not thereby favourable to the dynamics of aggregate growth. The two are not the same quantity, and treating them as one is precisely the confusion the cycle punishes.

    Note, too, what the splines are doing. Three of the ten factors would not sit still in a straight line. That is not a technical footnote: it is the first quantitative sign that the interaction of these factors involves thresholds and turning points rather than a stable proportionality.

    Not random. Chaotic.

    “Unpredictable” and “random” are not synonyms, and the difference decides what kind of science economics can be. A random system has no internal structure to find. A chaotic one is rigidly determined and still unpredictable at long horizons, because arbitrarily small differences in initial conditions grow exponentially apart.

    Three measurements place the U.S. economy in the second category. The Lyapunov exponent is positive (approximately $0.0515$): small perturbations amplify rather than dissipate. The correlation dimension is not an integer ($3.32798$): the attractor reconstructed by Takens’ theorem has a fractal structure, patterns repeating across scales of time and magnitude — which is what “cyclical, but not periodic” means when it is stated precisely. And recurrence quantification finds high determinism alongside variability in laminarity and in the maximum diagonal line length: underlying deterministic structures that themselves evolve.

    $\lambda > 0 \quad\text{with}\quad D_2 = 3.32798 \notin \mathbb{Z}$

    Read together, these say something a forecaster should find sobering and a theorist should find encouraging. The long-horizon forecast is not merely hard; it is structurally bounded. But the structure that bounds it is real, stable and measurable — which is exactly what a theory of the cycle needs to have something to explain.

    The shape of time

    The most unusual instrument in the thesis is topological. The idea is to stop asking how big the numbers are and start asking which observations can see which. Convert the series into a directed visibility graph — a link from one quarter to another when the second is visible from the first over the intervening data — and study the order structure that results.

    Two topologies are built on it, and they disagree in an informative way.

    • The coarser Alexandrov topology, built on temporal reachability, turns out to be connected. At the level of its order structure the economy is globally a single piece: every observation is bound to every other by chains of temporal visibility. There is no quarter that stands apart.
    • The finer Nada topology is locally fragmented — six components under the natural visibility graph, thirty-six under the horizontal one. Zoom in, and the fabric shows seams: structural discontinuities at the level of closed neighbourhoods.

    Global unity and local rupture at once. That duality is not a contradiction to be resolved; it is the object being described. And a third measurement gives the whole thing a direction: the bitopological analysis yields $D = +4$, meaning that expansions generate more temporal visibility than contractions. The cycle is not symmetric in time. Growth accumulates gradually and in view; collapse happens abruptly and blind. Run the film backwards and it is recognisably the wrong film.

    ⚠️ Why you must not “clean” the crises

    There is a habit in applied work of treating extreme values as contamination and smoothing them away by discontinuous imputation. Here that habit is shown to be a category error with a measurable price. The extreme fluctuations of the 2020 crisis belong to a connected block even under the finer topology; severing them is a topological rupture, not a cleaning operation. The thesis reports the consequence directly: models fitted after such imputation performed worse, because one was using predictors suited to one phenomenon — real output growth — to predict a qualitatively different one: real output growth after the crisis had been removed from it. The crises are not noise around the cycle. They are the cycle.

    The grammar of the cycle

    The third question receives a seven-part answer. The factors interact through feedback (the rate of profit shapes investment, investment shapes the organic composition of capital, which feeds back into the rate of profit); time lags (R&D and fixed investment pay out on a delay, and the delay is itself cycle-generating); non-linearity (thresholds and regime changes, which is why three factors needed splines); deterministic chaos; sectoral interdependence between the department producing means of production and the one producing means of consumption; topological structure, global connectedness with local fragmentation; and the influence of the global context, which is how military spending and the S&P 500 enter a nominally domestic account.

    The unifying claim is that each phase of the cycle contains the seed of its own negation. New methods of production introduced during the crisis lay the foundations of the next boom; the overaccumulation of the boom prepares the ground for the next crisis. Innovation initially arrests the fall of the profit rate and ultimately deepens it — through the way the degree of exploitation of labour power responds, over time, to the very methods introduced to raise it.

    What a cycle is for

    The thesis closes on a question most treatments never pose. If the cycle is a mechanism, what does it accomplish? Two answers, at different depths. Its intermediate practical end is to restart the process of capital accumulation once instability has reached a critical level — this the mechanism achieves, repeatedly, at a cost borne unevenly. Its definitive practical end is to lay the material and spiritual conditions for a reorganization of the fundamental productive structure of society, one capable of a stability beyond what the capitalist mode of production can reach within its own limits.

    What this establishes, and what it does not

    The evidence supports the claim that classical Marxist economic theory possesses the greatest explanatory and predictive capacity for long-run cycles among the theories examined here, on this economy, over this period. It is a comparative result on the United States between 1992 and 2024, quarterly — not a universal proof, and not a forecast. The thesis is explicit about the cost of its own data: the Average Rate of Profit and the average rate of surplus value were available only annually through 2020, and completing the series to 2024 required temporal disaggregation and prediction, which puts a wider band of uncertainty around the most recent quarters. The philosophical, historical, conceptual and statistical scope of each result is distinguished in the text, and results unfavourable to the hypotheses are reported alongside the favourable ones.

    About this edition

    This is the English edition of a thesis originally written in Spanish and submitted to the Universidad Latina de Costa Rica for the degree of Licentiate in Economics. It is interdisciplinary by construction, drawing on Marxist political economy, dialectical and historical materialism, the history and historiography of economic thought, the philosophy and methodology of science, econometrics, Bayesian statistics, the theory of complex systems and topology.

    The edition carries a Note on the Translation that fixes the rendering of the terms whose Spanish usage is technical and not interchangeable with their nearest English cognates — gnoseology, sublation, long wave, solvent demand, technique — and records the editions from which quotations are taken, including the two distinct English and Spanish editions of the Soviet philosophical dictionary, which are cited under different transliterations because they are different books with different pagination.

  • SOME REFLECTIONS ON MARX’S PRICES OF PRODUCTION

    SOME REFLECTIONS ON MARX’S PRICES OF PRODUCTION

    Was Marx Wrong About Prices of Production? — A 260-Page Investigation Says No

    Political Economy • Econometrics • Marx

    Was Marx Wrong About Prices of Production?
    A 260-Page Investigation Says No.

    How one researcher spent years showing that the most famous critique of Marx’s economics rests on a mistake Marx never made.

    Based on: Gómez Julián (2026), “Some Reflections on Marx’s Prices of Production” — Introduction, Conclusions & the Formal-Empirical Chapter · DOI 10.5281/zenodo.21842251

    A Fatal Flaw, or a Fatal Misreading?

    For over a century, a single mathematical argument has been wielded as the definitive proof that Karl Marx’s economics doesn’t work. It goes like this: Marx claimed that the value of goods is determined by the labor that produces them, and that market prices eventually gravitate toward “prices of production” — modified versions of those labor values, adjusted for how capital-intensive each industry is. But when you try to verify this with a system of simultaneous equations, the numbers don’t add up. The sums of values don’t equal the sums of prices. The theory, critics have said since the early 1900s, contains a fatal algebraic error.

    This paper — spanning 260 pages and drawing on philosophy, history, sociology, and statistics — argues that the error was never Marx’s. It was the error of the people who checked his math using a method he never used.

    The Photograph vs. the Movie

    Imagine you’re trying to understand a river. You could take a photograph of it — capturing one frozen moment — or you could film it as a movie, watching how the water flows over time. For over a hundred years, the economists who criticized Marx took a photograph of his theory and then complained that it didn’t look like a movie.

    Here’s the specific issue. Marx described a two-step process: first, a general rate of profit forms across the entire economy; then, each industry’s price deviates from its pure labor value according to how much capital it ties up relative to the average. The standard critique — originating with Ladislaus von Bortkiewicz in 1907 and repeated ever since — takes all of Marx’s accounting identities and solves them simultaneously, as if input prices and output prices were determined at the same instant. Under that framework, Marx’s three aggregate equalities cannot all hold at once.

    The “inconsistency” that has been attributed to Marx for over a century is the inconsistency of the simultaneous-dualist framework that was imposed on him, and it dissolves as soon as time is restored. — Gómez Julián, summarizing the central thesis

    But here’s the catch: solving everything simultaneously is equivalent to assuming that the economy is a photograph — that there is no time. And Marx’s entire framework is built on the opposite premise: that the economy is a process, an unfolding sequence in which the prices that exit one period become the input prices that enter the next. Once you restore that temporal dimension, the “inconsistency” vanishes. The three equalities hold simultaneously — not because Marx was secretly consistent in some miraculous way, but because the contradiction was an artifact of the framework imposed on him, not of his own logic.

    The paper calls the simultaneous approach “Walrasian Marxism” — a phrase that captures the irony: economists imported the logic of Léon Walras’s general equilibrium theory and used it to read Marx, then blamed Marx when the result didn’t work.

    In Plain Language

    Marx was accused for over a century of getting the arithmetic wrong. What actually happened is that someone redid his arithmetic under an assumption he never made — that the prices of things you buy to produce and the prices of things that come out of production are the same prices, set at the same time. If you assume that, Marx’s accounts don’t close. But that assumption is equivalent to saying the economy doesn’t happen in time.

    But Was the Movie Real?

    Pointing out that Marx’s logic works when you read it correctly is necessary but not sufficient. The “temporalist” school has been making this argument for nearly fifty years. But the author noticed a critical gap: nobody in that school had ever taken real-world data and actually estimated the three types of prices Marx described — direct labor values, prices of production, and market prices — and then tested whether market prices actually gravitate toward prices of production as the theory predicts.

    This matters because, as the paper puts it, leaving the correct reading of Marx “in the territory of conceptual argumentation while the incorrect reading occupies alone the territory of measurement” is a strategic vulnerability. If you can’t show that real prices behave the way your theory says they should, your theory remains a philosophical argument, however internally consistent.

    But before presenting any numbers, the paper devotes substantial space to establishing that the process Marx described actually happened in history. This is not an appendix; it’s a foundational part of the argument.

    Before Capitalism

    In pre-capitalist societies, exchange was regulated by labor time — not because someone enforced a theory, but because the material conditions made it so. Barter was dominant, inflation did not exist, and prices could only reflect production costs given available technology. Evidence from anthropology (Malinowski’s Trobriand Islands studies), sociology (Mauss on gift exchange), accounting history (Kula’s analysis of feudal estate records), and even paleogenomics all converge: objects were valued in proportion to the labor they embodied.

    The Transition

    The dissolution of feudal relations, the monetization of exchange, and the destruction of pre-industrial normative frameworks created the conditions for capital to move freely between industries. Thompson’s work on the “moral economy” documents how the new free-market ideology had to be violently imposed, destroying customary protections and creating an unprecedented relationship of exploitation.

    Capitalism Established

    Once barriers to capital movement were destroyed, capital flowed from commerce to industry chasing higher profits, and generalized competition forced a redistribution of total surplus value across sectors. The crisis of 1873 — which destroyed nearly half the blast furnaces in major iron-producing countries — is presented as concrete evidence of the mechanism: firms whose costs were still based on older, individually more labor-intensive methods went bankrupt when they couldn’t compete with prices of production dictated by modern technology.

    In Plain Language

    Prices of production didn’t appear the day someone wrote an equation. They appeared the day capital could freely move from one industry to another chasing the highest profit — which didn’t happen until legal, moral, and political barriers were destroyed. Before that, things were exchanged roughly according to the labor they cost, and there is more than enough evidence — ethnographic, accounting, archaeological, and genetic — to show it.

    What Is a Production Price, Exactly?

    This is where the paper moves into its most technically original territory. The author carefully separates two things that must not be confused:

    What a production price is (the explanandum): it is the expected value, over the distribution of economic perturbations, of the long-run time average of market prices. In plain language: it’s the center of gravity around which actual market prices keep spinning. Not the price they arrive at and stay at (that would be equilibrium), but the average around which they never stop oscillating.

    Key Concept

    The production price is neither an eternal, timeless equilibrium (the error of the simultaneous approach and of Walrasian economics, which takes the law as such for the whole and eliminates time) nor a chaos of prices without law (the error of empiricism, which stays at the level of individual prices and loses the law). It is the law of the whole realizing itself through the contingency of the parts.

    How each step of the process works (the explanans): a rule that determines this year’s market price from last year’s market price and last year’s latent production price, and nothing else. This is modeled as a hierarchical Ornstein-Uhlenbeck process — a three-level cascade in which the production price is itself a latent state with its own dynamic gravitating toward value, and market prices gravitate toward that latent state rather than toward a fixed, noisy index.

    The uncertainty is built into the model explicitly: uncertainty in the average rate of profit, uncertainty in the advanced capital, uncertainty in the disaggregation of national accounts into 37 sectors (handled through multiple imputation with 25 imputations combined by Rubin’s rule), and parametric uncertainty estimated through Bayesian Markov Chain Monte Carlo methods.

    One crucial point: no magnitude is obtained by solving a simultaneous system. Value is constructed empirically and directly as $V = c + v + p$ (cost plus surplus value), and production price as $\Phi = c + K \cdot G’$ (cost plus capital times the general rate of profit). There is no Leontief inversion, no simultaneous algebra, anywhere in the construction.

    In Plain Language

    Think of a production price as the “gravitational center” of a spinning object. The object (a market price) never stops moving — it wobbles, it swings, it drifts — but over time its average position is pulled toward that center. The math describes both what the center is and how each wobble happens, and it does so while honestly accounting for all the uncertainty in the measurement.

    The Defining Equations: (9) Through (11)

    Here is where the metaphor turns into mathematics. The paper writes the definition of a production price in three successive steps — each one making explicit an assumption the previous step left implicit — numbered (9), (10), and (11) in the original text. None of the three generates a trajectory by itself; together they define the explanandum — what the object is — that the cascade below then generates.

    Equation 9 — What a Production Price Is
    $$ \lim_{t\to\infty} E\!\left[\varphi^i_t\right] \;=\; k^i_t + K^i_t\, E\!\left[G'(t,X)\right] \;=\; \Phi^i_t $$

    Here $\varphi^i_t$ is sector i’s market price at time $t$, $k^i_t$ is its cost price (constant capital consumed plus variable capital), $K^i_t$ is the total capital advanced, and $G'(t,X)$ is the general rate of profit — itself a stochastic process indexed by a perturbation $X$ that bundles the exodus of capital between branches and technological innovation.

    In words: a production price is the long-run limit of the average market price. Not the price itself at any instant — that keeps oscillating forever — but where its time-average settles as the horizon stretches out. Notice the object on the right-hand side, $k + K \cdot E[G’]$: it is the same accounting identity introduced earlier (cost price plus the average profit rate applied to capital advanced), except the profit rate is now written as an expectation, because it fluctuates.

    Equation 10 — Making the Averaging Explicit
    $$ \Phi^i_t = \lim_{t\to\infty} E\!\left[\varphi^i_t\right] = \int_{-\infty}^{\infty} \!\left(\lim_{t\to\infty} \varphi^i_t(x)\right) f_X(x)\, dx \;=\; k^i_t + K^i_t \int_{-\infty}^{\infty} G'(t,x)\, f_X(x)\, dx $$

    $f_X$ is the probability density of $X$. The equation says the expectation is an average over every possible state $x$ of the system’s turbulence, weighted by how likely that state is.

    Equation (10) earns its keep by making a subtle move legitimate: swapping the order of the limit and the expectation. That looks harmless, but it hides a real question — does the market price $\varphi^i_t$ even converge to anything as $t \to \infty$? The paper’s answer is no: a capitalist system doesn’t settle into a fixed point, it settles into a limit cycle — perpetual oscillation. So the convergence the argument needs isn’t of the instantaneous price, but of its cumulative time-average. That average does converge, for almost every state of the world, precisely because the system is ergodic — the fraction of time the cycle spends in each region of its orbit stabilizes. This is the Birkhoff ergodic theorem doing, in mathematical language, exactly what Marx says in economic language: the production price isn’t the value the market price reaches and stays at, it is the average around which it never stops oscillating. The oscillation isn’t an obstacle to the average — it is the average’s condition of existence.

    Why the Order of Operations Matters

    The paper invokes Lebesgue’s Dominated Convergence Theorem to justify swapping “limit of the average” for “average of the limit.” This requires bounding market prices by some integrable envelope — economically, that no price can grow without limit, which technological ceilings and competitive pressure guarantee — and, crucially, it does not require that the convergence be uniform across sectors. Uniform convergence would mean competition equalizes profits instantly and identically everywhere, with no room for a shock to hit one industry harder than another. Marx’s theory says the opposite, and the math is built to allow it.

    Equation 11 — When the Capital Base Is Also Uncertain
    $$ \Phi^i_t = \lim_{t\to\infty} E\!\left[\varphi^i_t\right] = \int_{-\infty}^{\infty}\!\!\int_{-\infty}^{\infty} \left[k^i_t + K^i_t(y)\, G'(t,x)\right] f_{X\mid Y}(x\mid y)\, f_Y(y)\; dx\, dy $$

    Equation (10) still treated the capital base $K^i_t$ as known exactly. Equation (11) drops that simplification: $Y$ is a second random variable carrying the estimation error in $K$, with density $f_Y$, and $f_{X \mid Y}$ lets the profit-rate perturbation depend on which realization of that error occurred. The object is the same double average — only now uncertainty is propagated from two sources instead of one.

    This last equation is not a mathematical flourish; it is the reason the empirical section spends so much effort on multiple imputation. National accounts don’t hand anyone a clean measurement of capital advanced by sector — it has to be reconstructed from incomplete data, and that reconstruction carries its own error. Equation (11) is the license to treat that error as a random variable to be averaged over rather than a nuisance to be ignored. The uncertainty is propagated externally — by a generator outside the statistical model itself — rather than estimated as an internal parameter of the dynamic model: estimating $K$’s error inside the model would confound it with the model’s own measurement-noise term, opening a ridge of non-identification between two magnitudes that the data alone cannot tell apart. Kept external, twenty-five complete reconstructions of the data are generated first, each respecting the Marxian aggregate identities to machine precision, the dynamic model is fit on each, and the twenty-five fits are combined by Rubin’s rule. That is the outer average of equation (11), computed by literally drawing from the distribution of $Y$ instead of assuming it away.

    The Engine: A Three-Level Ornstein–Uhlenbeck Cascade

    Equations (9)–(11) define the target; they don’t generate a path toward it. The explanans — the mechanism that actually produces a year-by-year trajectory consistent with that target — is a hierarchical Ornstein-Uhlenbeck process with up to three nested levels, fit as a single Stan program (the same program handles one, two, or three levels, which guarantees that adding levels can never silently break the simpler cases nested inside them). All series enter standardized; time is discretized one year at a time using the Euler–Maruyama scheme.

    Level 1 — The Market Price
    $$ dev_{t,s} = \varphi_{t-1,s} – \Phi_{t-1,s} $$
    $$ \kappa^m_{t,s} = \kappa_{\mathrm{cap}} \cdot \mathrm{invlogit}\!\left(\kappa_s + \beta_1\, z^{TMG}_t\right) $$
    $$ \Delta\varphi_{t,s} = \kappa^m_{t,s}\!\left(-\,dev_{t,s}\right) \;+\; a_{3,s}\, dev_{t,s}^{\,3} \;+\; \gamma\, COM^{std}_{t,s} \;+\; \varepsilon_{t,s} $$

    Subscripts $s$ (sector) and $t$ (year) run throughout. $dev$ is last year’s gap between market price and the latent production price. $\kappa^m$ is the sector’s reversion speed, passed through a logit link that caps it inside $(0, \kappa_{\mathrm{cap}})$ and lets the general rate of profit ($z^{TMG}$) modulate it without ever pushing the system out of the stable region of the discretization. $\varepsilon$ is a fat-tailed (Student-t), stochastic-volatility innovation, so volatility can cluster in time without destabilizing the mean.

    Read the Level 1 line as a spring. The term $-\kappa \cdot dev$ is the restoring force: it pulls the market price back toward the production price with a force proportional to how far it has drifted. The cubic term $a_{3,s} \cdot dev^3$, with $a_{3,s}$ constrained negative by construction — not estimated, imposed — makes that restoring force grow faster than proportionally once the deviation gets large: the further the market strays, the harder it snaps back. This is a declared stability assumption, not a discovery: it guarantees the model can never generate an explosive regime, at the real cost that if such a regime existed in some sector of the actual economy, this particular specification could not detect it.

    Levels 2–3 — Where the Latent Center Itself Reverts
    $$ \mu_{s,t} = m_{0,s} + m_1\, G’_t + m_v\, V_{s,t} $$

    The production price $\Phi$ is not treated as a fixed, observed index; it is itself a latent state that reverts — more slowly, with its own sector speed $\kappa_p$ — toward this mean $\mu$. $m_1$ is the channel running through the general rate of profit; $m_v$ is the coefficient measuring how strongly the production price tracks the directly-constructed value $V_{s,t} = k + p$ (Level 3, and the reason the cascade goes up to three levels rather than stopping at two).

    This is the bridge back to the abstract equations above, term by term. $\mu_{s,t}$ is the estimable stand-in for the right-hand side of (9): $m_{0,s} + m_1 G’_t$ plays the role of $k + K \cdot E[G’]$, and $m_v V_{s,t}$ is the specific functional form chosen for the value-tracking channel that the abstract definition deliberately leaves open (the paper is careful to say that capitalist competition as a function of the value structure is declared at the level of equations 9–11, not derived; giving it the concrete shape $m_v V$ is a modeling choice made at the cascade level, defended by how it performs under validation rather than deduced from the definition). And the expectation of $G’$ from equation (9) has its operational counterpart in the profit rate averaged across the twenty-five multiple imputations — the mechanism equation (11) licenses.

    The coefficient $m_v$ carries real theoretical weight: it is the empirical stand-in for Chapter 9’s claim that prices of production gravitate around values. It is given a neutral prior, $m_v \sim \mathcal{N}(0,\, 0.5)$ — centered at zero, symmetric, assigning equal plausibility to $m_v > 0$ and $m_v < 0$ before seeing any data. That matters for the same reason a fair coin matters in a coin-flip experiment: if the data carried no signal, the posterior would sit wherever the prior put it, hugging zero. It doesn’t. It lands at $m_v \approx 1.0136$ with $P(m_v > 0) = 1$ — evidence that the data moved it there, not the prior. The anchoring to value is found, not assumed into the setup.

    In Plain Language

    The cascade is three springs stacked on top of each other. The market price is tied by a spring to the latent, unobserved production price. The production price is tied by its own, slower spring to a moving target that blends the general rate of profit with the directly-measured labor value. Pull any one spring and let go: it doesn’t snap to a fixed point, it settles into the kind of perpetual, decaying oscillation that equations (9)–(11) describe as an average. The springs are estimated from sixty-one years of real U.S. data, not assumed; the coefficient tying prices of production to values, specifically, could have come back negative or zero — the model gave it every chance to — and it didn’t.

    What the Numbers Say

    The empirical core of the paper is a panel of 37 productive branches of the United States economy over 61 years, from 1960 to 2020. The hypothesis tested encloses three distinct relationships, and the paper is meticulous about not conflating them. Each is stated, tested, and reported separately.

    Market Prices ↔ Prices of Production: The Strongest Link

    This is the relationship with the firmest statistical support, confirmed through six independent lines of evidence:

    Central Finding

    Gravitation exists, and it is slow. The median speed across sectors is $\kappa_m = 0.0770$, equivalent to a half-life of approximately 9 years. Market prices take about a decade to cover half the distance toward their production-price center. This is consistent with Marx’s characterization of gravitation as a tendential, mediated regulation, not an instantaneous fit.

    The number is remarkably stable under stress tests:

    • Removing five of the six productive blocks from the panel barely moves the estimate — it shifts in the third decimal place. The sixth, which gathers 18 of the 37 sectors, does produce a shift (from 9 years to 6 years), and the paper decomposes it: about half the acceleration is the generic effect of halving the panel — removing 18 sectors at random already gives 0.0929 — and not the block itself.
    • Dismantling the value anchor in three different ways — including permuting surplus value across spheres — moves the speed in the third decimal place. This is significant: it means the conclusion about market-to-production gravitation does not depend on the less robust production-to-value link.
    • The market deviation has its own dynamic signature. Compared against a random walk matched in variance, three out of six test statistics separate cleanly (the weighted-sum convergence reaches a tolerance of 0.01 while the null never reaches a tolerance ten times more lenient; recurrence analysis laminarity triples the null; recurrence entropy doubles it). The ones that don’t separate are recurrence-analysis determinism and the two deterministic-chaos invariants — the Lyapunov exponent and the correlation dimension — which the paper never claimed to find.
    • The estimate is invariant to secondary methodological choices. Sweeping the latency regularizer across three values produces life medias of 9 years in all three arms (speeds of 0.0774, 0.0770, 0.0772).
    • The known bias of disaggregation pushes against the result. Splitting a national figure among 37 branches is underdetermined and biases speed estimates downward — meaning the true half-life is probably 7–8 years rather than 9. A bias that works against your conclusion is one you can live with, because the result holds despite it, not thanks to it.

    Prices of Production ↔ Values: The Thinnest Leg

    This is the weakest part of the empirical argument, and the paper states so with complete transparency. The problem is not a defect of the instrument but a property of the object:

    Methodological Transparency

    The coupling coefficient estimated within the dynamic model is $m_v = 1.0136$ with a 95% credible interval of $[1.0096,\; 1.0176]$ — but the same procedure returns 1.0365 when surplus value is permuted across spheres, preserving all annual aggregates. Why? Because production price and value share the cost price, which explains 66.1% of the variance of the former and 72.0% of the latter, and their correlation in levels is 0.9987. The coefficient would land near one even if the law of value didn’t hold at all. The paper therefore reports it as a consistency check, not as evidence.

    The real support for this relationship comes from cross-sectional tests, not from the dynamic coupling. When temporal common trends are removed and analysis is conducted within-year, the slope of the markup on own surplus value is 0.675 with the true data versus 0.090 under permutation, with intervals that don’t come close to overlapping. The sectoral ordering of the wedge between $\Phi$ and $V$ has an inter-annual rank correlation of 0.986 and a 60-year value of 0.558 — highly persistent structure, not noise.

    A collateral finding worth noting: the coefficient of variation of sectoral profit rates is 0.669 — meaning profit rates across industries show considerable and persistent dispersion. Far from contradicting the theory, this dispersion is the condition of existence of the mechanism: if profit rates were already equalized, there would be no differential to drive capital migration, and gravitation would have nothing to operate on. Marx postulates equalization as a tendency, not an accomplished fact.

    Market Prices ↔ Values: Sustained in Form, Adjusted in Existence

    The structural modification across sectors exists and is nonlinear (the nonlinearity step holds comfortably at 6.8 null deviations). But the existence step is adjusted: 44% of its gain is obtained equally with sectoral characteristics unpaired from their spheres, and the gap against the maximum null is on the order of one paired standard error. The coefficients survive a deliberately severe correction for serial dependence (tripling the error).

    The Instrument Behind That Number: A Nested Ladder in gdpar

    That test is a small ladder of nested distributional-regression models, fit with gdpar (Gómez Julián, 2026b), the author’s own R package for generalized distributional parameter regression, published on CRAN on July 15, 2026. The ladder climbs from a bare model — “the market-to-value ratio has no sector-specific correction at all” — through a model where organic composition, wage share, and sector size shift that ratio linearly, up to a model where the correction is a flexible spline rather than a straight line. Two gains matter, measured in units of predictive density: adding the linear correction buys 207.3 units; letting it curve buys another 215.1. Both were checked against a control built to be hard to pass — shuffling which sector gets which characteristics 99 times, refitting each time, with the spline’s knots held fixed across every shuffle so the comparison can’t be won by a better basis alone. The curvature gain clears its null with room to spare (6.8 null standard deviations; the best of 99 shuffles reaches only 114.8 against 215.1 observed). The existence gain is honestly reported as thinner: shuffled sectors still buy about 44% of the real gain merely by having some characteristics to fit — three covariates and an intercept give a model room to accommodate noise even when it is being told nothing true — so the genuine margin over the null sits at about one paired standard error (23.2, against a gap of roughly 24 units). Both numbers are reported together, precisely so the large one isn’t read alone.

    A companion specification, estimated in the same gdpar fit, asks the same question about dispersion rather than location: not where the market-to-value ratio is centered, but how tightly it clusters. Larger sectors and sectors with higher capital composition show systematically less relative dispersion — elasticities of $-0.226$ and $-0.104$ — consistent with equalization operating more effectively where capital is more concentrated. Both effects clear a “breaking factor” (the multiple of the standard error at which the 95% interval would first touch zero) north of six and four respectively, past the 2.94 ceiling reached anywhere else among this paper’s location coefficients, and the finding reproduces under a completely different likelihood family (a gamma distribution on the price ratio) to within 5.2%.

    Three Failures That Confirm the Theory

    One of the most intellectually striking features of this paper is how it handles results that, at first glance, look bad for its thesis. There are three, and the paper reports all of them without softening — then shows deductively why each one was expected if the theory is correct.

    Negative Result No. 1

    The model does not out-of-sample predict better than a random walk. But this was deductively implied by the slow form of the thesis. At a horizon much shorter than the half-life, a mean-reverting process is, to first order, a random walk. If something takes a decade to get halfway back, looking at a single year won’t let you see it return.

    Negative Result No. 2

    The value term is predictively indistinguishable. Again, this follows from the slow coupling between prices of production and values: with half-lives on the order of decades and only 61 years of data, univariate root-unit tests are structurally underpowered.

    Negative Result No. 3

    No univariate test separates the true wedge from its permuted placebos. But this was predicted before measuring, by the persistence of sectoral ordering itself (inter-annual rank correlation of 0.986). A highly persistent time series is hard to distinguish from its permuted version using tests designed for shorter memory.

    Finding these signatures is corroboration of the slow form of the thesis, and not finding them would have been the real problem. — Gómez Julián, on the negative results

    The paper’s stance on this is worth highlighting: “Lejos de refutar la tesis, los tres están deductivamente implicados por su forma lenta” — far from refuting the thesis, all three are deductively implied by its slow form. A single mechanism (slow gravitation) explains both the substantive thesis and all the apparently negative results, and it also survives in the validated posterior. “That a single cause explains the thesis and all the apparently negative results, and that it additionally survives in the validated register, is the opposite of a petitio principii: it is a unified, falsifiable, and internally validated narrative.”

    Temporalism Isn’t a Preference — It’s a Condition of Measurement

    Perhaps the most consequential result in the entire paper is not a number but a statement about what can and cannot be measured. It concerns the “modulator” — the component of Marx’s argument in which the general rate of profit enters into the structural modification of each sphere, meaning the deviation of each sphere is not independent of the reference but generated by it.

    The Identifiability Argument

    When the model was run with a single, fixed general rate of profit for all 61 years (as a simultaneous approach would require), the posterior exhibited a flat ridge: two completely different functional bases (a degree-two polynomial and a spline basis) produced the same pathology to the third decimal place, with an effective sample size of only six draws. The diagnostic got worse with more sampling (R-hat rising from 1.33 to 1.73). This is the unmistakable signature of a direction in parameter space along which the likelihood does not change.

    The cause is theoretical, not computational. With one fixed reference, the modulator can only be identified evaluated at that single point — a single number, not a function over the space of references. You cannot estimate three coefficients from a polynomial if you have one data point.

    When the reference was allowed to vary year by year (61 different general rates of profit), the model converged within minutes, with a large improvement in both time and effective sample size, and zero divergences.

    Named, Not Improvised: Theorems 1A and 1E

    This diagnosis isn’t an ad hoc read of a misbehaving sampler. gdpar (Gómez Julián, 2026b) — the same package behind the nested ladder above — ships a formal identifiability result for exactly this situation. Its Theorem 1A establishes that, with a single fixed reference point, a distributional modulator is identified only at that point: as one number, not as a function over the space of possible references. Theorem 1E is the positive counterpart: letting the reference vary restores identifiability of the modulator as a function. Fitting a degree-two polynomial (three coefficients) or a five-knot spline basis (five coefficients) against one single, unmoving reference asks for more than a single data point in that dimension can support — which is exactly what a flat likelihood ridge looks like from the sampler’s side.

    The figures behind the improvement, precisely: a fixed reference with a degree-two polynomial gives an R-hat of 1.7333, an effective sample size of 6, and 8 divergent transitions in 39 minutes; a one-knot spline basis reproduces the same pathology — R-hat 1.7335, effective sample size 6, 14 divergences, 5.6 hours. Letting the reference vary year by year (61 distinct annual values of the general rate of profit), centering the additive component and raising the sampler’s adaptation parameter to 0.99, gives an R-hat of 1.0035, an effective sample size of 1332, and zero divergent transitions — in 2.9 minutes. That is the 115-fold improvement in time and 222-fold improvement in effective sample size referenced above, and it is a theorem, not a tuning trick: no amount of additional sampling closes that gap under a fixed reference, because the object being asked for — the modulator as a function — simply is not there to find.

    The consequence is stated precisely: with a single fixed general rate of profit obtained by solving the system simultaneously, the claim of Chapter 9 of Volume Three of Capital is unverifiable by construction. It is not that the data are insufficient — the object is not identified, and no amount of data would identify it. The argument does not establish that simultaneism is false as a description of capitalism (that is established by historiography and sociology); it establishes that a simultaneous procedure cannot, even in principle, empirically verify the specific part of Marx’s argument that this work estimates.

    In Plain Language

    Marx says: first a general rate of profit forms, then each industry deviates from it according to how capital-intensive it is. To check whether the deviation depends on the general rate, you need to see what happens to the deviation when the general rate changes. If you calculate one general rate for the entire 61-year span, it never changes, and there is nothing to observe. That is exactly what happened: the model with one fixed rate doesn’t converge — not because of computational limitations, but because it is being asked to measure a relationship with a single observation of one of the two variables. Calculating one rate per year — which is what the temporal reading says you should do — the same model converges in three minutes.

    What This Is, and What It Isn’t

    The paper is careful, almost painstakingly so, about the limits of what it claims. This section matters because a reader coming from the “pro-Marx” or “anti-Marx” side might be tempted to over-read the results. The author doesn’t let you.

    What the evidence authorizes: In the United States between 1960 and 2020, market prices gravitate toward prices of production with a decadal half-life that is sectorially heterogeneous, and this speed survives three independent assaults (removing five of the six productive blocks, destroying the value anchor, varying secondary methodological decisions). This is a measured, calibrated, and falsifiable fact.

    What the evidence does not authorize:

    • It does not claim superior predictive power (the model does not out-predict a random walk, which was expected).
    • It does not claim that univariate root-unit tests confirm gravitation (they are structurally underpowered at this time scale).
    • It does not claim uniqueness or categorical novelty. The contribution is the explicit integration and canonization of a slow gravitation cascade with value anchoring, measured on real data, with propagated uncertainty, validated, and subjected to a diagnostic whose unfavorable results are reported alongside the favorable ones.
    • It does not claim that this statistically demonstrates the law of value, “and not for rhetorical prudence but because it would be false: a price series can show that a magnitude behaves as the law predicts, and cannot explain why that magnitude exists or whether the category with which we name it is the correct one.”

    That last point is the paper’s deepest epistemological commitment. Questions about whether “value” is the right category for what prices ultimately measure are not answerable by any price series, no matter how long. They are answered by history, sociology, and philosophy — and the firm answer is the one obtained when all four disciplines (those three plus statistics) point in the same direction. The four-dimensional convergence is the argument, not any single leg of it.

    The paper also addresses the homology that unifies its seemingly disparate halves — the historiographical-filosofical first chapter and the econometric second chapter. The relationship between necessity and contingency that governs the transition from feudalism to capitalism (where the same demographic shock produced opposite outcomes in different regions of Europe) is structurally identical to the relationship between prices of production and market prices. A law determines the center; circumstances determine each particular outcome. Neither fact negates the other, because they describe different levels of the same reality.

    What It All Adds Up To

    Here is the simplest version of what this 260-page paper establishes:

    Marx was reproached for a century for having done an arithmetic calculation wrong. What happened is that his calculation was redone under an assumption he never made: that the prices of things bought to produce and the prices of things that come out of production are the same prices, fixed at the same time. If you assume that, Marx’s accounts indeed don’t close. But that assumption is equivalent to saying the economy doesn’t happen in time. As soon as you accept that what exits the factory this year is what enters the factory next year, the accounts close without anyone having to fix anything. — Gómez Julián, Summary for the Reader

    But recognizing the conceptual error was only the first half. What had been missing — and what this paper contributes — is doing those accounts with real data instead of with fictitious numerical examples, which is what the school that had the correct conceptual reading had never done.

    The empirical results show that prices in the U.S. economy over six decades do behave as the theory predicts: they gravitate, slowly, toward prices of production calculated with Marx’s theory and no other. This finding survived every attack the author could devise — removing productive sectors, destroying the value anchor, permuting surplus values, varying methodological decisions, and running diagnostics whose unfavorable results are reported in full alongside the favorable ones.

    The part of the argument linking prices of production to labor values is also supported by real evidence, though less firmly, and the paper says exactly where the weak points are and why they are properties of the object, not defects of the instrument.

    And the paper does not claim to have demonstrated the law of value with a series of numbers, because “questions of that kind are not answered with numbers: they are answered with history, with sociology, and with philosophy, and the firm answer is the one obtained when the four things (the previous three, together with statistics) all point in the same place.”

    That convergence doesn’t make the result eternal — better evidence can overturn it tomorrow. But it makes it, for now, “our best possible approximation to the truth.”

    — — —

    “In science as in life, overcoming adversity is what makes us truly strong.”

    This post summarizes the introduction, conclusions, and the formal-empirical chapter (§2.4) of Gómez Julián, J. M. (2026). Some Reflections on Marx’s Prices of Production: Historicity of the Law of Value, Dialectical-Materialist Foundation, and Dynamic Formalization Under Uncertainty. Zenodo. https://doi.org/10.5281/zenodo.21842251. The full paper spans approximately 260 pages across two chapters covering philosophy, historiography, mathematical formalization, and empirical econometrics. Equations (9)–(11) and the model specification cited here reproduce that chapter’s notation; gdpar is cited separately as Gómez Julián (2026b).

    Written for the curious. An invitation to read.

  • BITOPOLOGICAL SPACES: LISTENING TO THE DIRECTION OF TIME WHEN IT MATTERS

    BITOPOLOGICAL SPACES: LISTENING TO THE DIRECTION OF TIME WHEN IT MATTERS

    Bitopological Spaces: Listening to the Direction of Time When It Matters

    Bitopological Spaces: Listening to the Direction of Time When It Matters

    How two topologies — built from the same data — can hear the difference between past and future

    Most of the tools we use to analyze sequences of data — averages, correlations, spectral analyses — treat time as a label that could run in either direction without changing the answer. Reverse the order of your data points and many standard methods give you identical results. But in the real world, the direction of time matters profoundly. Economies expand slowly and crash suddenly. Heartbeats rise smoothly and fall steeply. A method blind to direction is a method blind to one of the most fundamental features of how systems change.

    A recent paper by independent researcher José Mauricio Gómez Julián introduces a construction that addresses this gap. Taking a known method from graph theory and extending it to directed graphs, the paper produces a pair of topologies — mathematical frameworks for understanding structure and connectivity — whose divergence is a topological fingerprint of temporal irreversibility. Applied to three decades of American economic data, the method recovers a picture that is both mathematically precise and economically interpretable. Here is a walk through the main ideas.


    Seeing and Being Seen

    The starting point is a beautifully simple idea introduced by Lucas Lacasa and collaborators in 2008. Imagine plotting a time series — say, 129 consecutive quarterly growth rates of U.S. GDP — as points above a timeline. Now connect two points with a line if they can “see” each other: the straight segment between them passes above every intermediate data point, as if you stood at one point and shone a flashlight toward the other with no obstacles in the way.

    The result is a visibility graph: a network whose nodes are time points and whose edges encode a geometric relationship. Visibility graphs have been used to classify chaotic systems, detect heartbeat anomalies, and distinguish between types of economic regimes. They translate the shape of a time series into the structure of a graph, opening the door to the vast toolkit of network science.

    But standard visibility graphs are undirected: an edge between two points does not record which one came first. If you orient each edge from the earlier time point to the later one, you obtain a directed visibility graph — a directed acyclic graph in which the arrows always point forward in time. This orientation carries information about temporal asymmetry that the undirected graph throws away entirely.

    From Networks to Structure

    Here is where the paper’s contribution begins.

    In 2018, Huda Nada and collaborators introduced a procedure for turning any undirected graph into a topological space. For those unfamiliar with the term, a topology is a mathematical framework that defines what it means for groups of points to be “open,” for sets to be “connected,” and for spaces to have “structure.” It operates at a level more abstract than distances or coordinates — it captures the pattern of how sets overlap and separate.

    The Nada construction works as follows. For each vertex of the graph, compute its closed neighborhood: the vertex itself plus all of its immediate neighbors. Take this family of neighborhoods and generate a topology by closing it under two operations: finite intersections (combine neighborhoods by overlapping them) and arbitrary unions (combine neighborhoods by collecting them). The result is a topology on the vertex set, and its invariants — connected components, separation properties, component counts — capture structural features of the graph.

    This procedure is universal: it works for any family of subsets of any set. The mathematical content lies in identifying the right family to use.

    Gómez Julián’s key observation is that for a directed graph, you do not get one family of neighborhoods — you get two. For each vertex:

    The forward closed neighborhood includes the vertex itself and all the vertices it points to — the later time points it can see. The backward closed neighborhood includes the vertex itself and all the vertices that point to it — the earlier time points from which it is visible.

    Apply the Nada procedure to the forward neighborhoods and you get a topology τ+. Apply it to the backward neighborhoods and you get a topology τ. The resulting triple (V, τ+, τ) is what mathematicians call a bitopological space: a set equipped with two topologies simultaneously, a concept introduced by John Kelly in 1963.

    The extension is, in a precise mathematical sense, trivial — the topology axioms do not care where the generating family came from. But recognizing this trivial extension as the right thing to do, and showing that the resulting bitopological structure captures something real about temporal asymmetry, is the paper’s central insight.

    When Two Topologies Disagree

    If the process generating your data is symmetric — equally likely to go up as down, at the same speed — then the forward and backward neighborhoods are statistically exchangeable. The two topologies τ+ and τ look the same, and the bitopological structure adds nothing beyond the undirected construction.

    But if the process is asymmetric, the two topologies diverge. Consider the prototypical asymmetry of economic and physical systems: gradual expansion followed by sudden contraction. Forward visibility through a gradual rise connects many points — each point can see far ahead through the gentle slope. Backward visibility through an abrupt drop connects few — the sharp fall blocks the line of sight. The forward topology ends up more connected (fewer separate components) than the backward topology.

    This divergence is the topological fingerprint of temporal irreversibility. The paper defines three quantitative measures of it:

    The asymmetry direction Δ = C − C+, where C+ and C are the numbers of connected components in the forward and backward topologies. Positive Δ means the forward topology is more connected. The component-count irreversibility index IC, which normalizes the difference to lie between 0 and 1. And the base-size irreversibility index IB, which measures the analogous difference in the sizes of the generating bases.

    These are pure numbers — no calibration, no free parameters, no training data. They emerge from the structure of the data and the construction itself.

    If you reverse the direction of time in your data and the topologies change, something in the process that generated the data is irreversible — and the gap between the two topologies measures exactly how much.

    Peeling the Onion: Three Layers of Structure

    One of the paper’s most clarifying contributions is the identification of three nested layers of topological structure on the same time series, each revealing different information.

    Layer 1 — The Alexandrov topology (reachability). For a directed acyclic graph, the most basic topology treats as “open” any set that is closed under forward reachability: if a node is in the set, all its descendants are too. This topology has exactly one connected component for any weakly connected graph, because every node can reach every later node through some directed path. At this level, the system is globally indecomposable. It tells us what we already know: the economy is a single connected process in which each quarter influences every subsequent quarter through chains of causation.

    Layer 2 — The undirected Nada topology (local fragmentation). When you apply the Nada construction to the undirected shadow of the visibility graph, the topology fragments dramatically. The intersection closure of the neighborhoods reveals clusters of time points that share local structural similarity — groups of observations linked by overlapping visibility neighborhoods — that go beyond mere reachability. This layer uncovers genuine structure that the reachability topology hides entirely.

    Layer 3 — The bitopological layer (temporal asymmetry). When you split the construction into forward and backward, a further distinction emerges. The forward and backward topologies have different component counts — and the difference is invisible to the undirected construction and invisible to the Alexandrov construction. It lives only in the gap between the two directed topologies.

    Each layer is contained within the next: the Alexandrov topology is a subtopology of the Nada topology (a theorem proved in the paper), which in turn underlies the bitopological structure. But each coarser layer hides information that the finer layer reveals.

    What the American Economy Looks Like Through a Topological Lens

    The paper applies the full pipeline to the quarterly growth rate of U.S. real GDP from Q1 1992 to Q1 2024 — 129 observations spanning the dot-com bust, the Global Financial Crisis, and the COVID-19 shock. Two graph constructions are used: the Horizontal Visibility Graph and the Natural Visibility Graph, both in their directed forms.

    The headline finding is that Δ = +4 in both constructions. The forward topology has 4 fewer connected components than the backward topology, regardless of which visibility-graph variant you use. This positive value is consistent with the well-documented asymmetry of the American business cycle over this period: expansions are gradual and sustained (1992–2000, 2001–2007, 2009–2020), while contractions are sharp and short-lived (2001, 2008–2009, 2020). Forward visibility through a gradual expansion is unobstructed; backward visibility through an abrupt contraction is fragmentary.

    What makes this finding compelling is its invariance. The HVG and NVG produce very different graphs — 248 vs. 406 edges, different base sizes, different absolute component counts — yet they agree on the sign and magnitude of Δ. The signal appears robust: a feature of the underlying data, not an artifact of how you choose to draw the graph.

    Another detail worth noting: the base-size irreversibility index IB is exactly zero in both constructions. The forward and backward topologies are generated by bases of equal size (250 and 250 for the HVG, 239 and 239 for the NVG). The asymmetry lives entirely in the structure of those base elements and how their intersections distribute — not in how many there are. The two topologies are built from the same number of building blocks, but those blocks fit together differently depending on whether you are looking forward or backward through time.

    A Single Shock

    Perhaps the most striking empirical finding is what the topology says about the COVID-19 shock.

    The second quarter of 2020 recorded the sharpest contraction in U.S. GDP on record — an annualized rate of roughly −31%. The third quarter recorded the sharpest rebound — roughly +33%. These are the two most extreme observations in the entire 129-quarter series, opposite in sign and opposite in economic interpretation.

    A naive analysis would naturally separate them: one is the worst crash, the other the best recovery. They sit at opposite ends of the value spectrum.

    But the Nada topology classifies them together. Under both the undirected and directed topologies, under both the HVG and NVG, these two observations belong to the same connected component.

    Why? Because the topology is not a proximity measure. It does not group points by how close their values are. It groups them by the structure of their visibility neighborhoods — which other points they can see, and how those visibility patterns intersect. Despite their extreme and opposite values, the two quarters share neighborhoods that overlap substantially. The intersection closure, which drives the Nada construction, puts them in the same cluster.

    This matches the interpretation most economists give to the event: the contraction and the rebound are two phases of a single exogenous shock, driven by the same underlying cause — the pandemic and the policy response to it. The topology recovers this interpretation from the geometry of the data alone, without any economic priors built in.

    What the topology does not claim: it does not say that the two quarters are “similar” in value (they are the two most distant observations in the entire series). It says they are structurally linked — that no topological open set separates them. The construction responds to the combinatorics of visibility, not to the metric of distance.

    Certifying the Construction

    The paper takes reliability seriously at three levels.

    Machine-checked proofs. The central theorem and related core results have been formalized in Lean 4, a proof assistant, against the Mathlib mathematical library. A computer has verified that the proofs are logically correct, with no gaps or hidden assumptions. The formalization is archived alongside the paper as part of a reproducibility bundle on Zenodo.

    Polynomial-time algorithms. Every step of the construction has an explicit algorithm with proven complexity bounds. The connected components of each topology can be computed in polynomial time without enumerating the full topology, which can be exponentially large. The key trick is to work through a combinatorial proxy for the topology called the specialization preorder, using bitset operations that are highly efficient in practice.

    Honest uncertainty. A three-valued decision procedure for pairwise connectedness reports “pairwise connected,” “pairwise disconnected,” or “undecided” — the last when the computation exhausts its resource budget. Rather than guessing, the algorithm honestly reports that it has not finished the work. This is a methodological commitment as much as a technical one: a topological statement counts as established only when the computation has completed the work that establishes it.

    No free parameters. The construction has no tuning knobs. The topological invariants — component counts, base sizes, irreversibility indices — are determined entirely by the data and the definitions. There is nothing to calibrate, nothing to overfit.

    A New Lens

    The paper does not propose to replace existing methods of time series analysis. Correlation, spectral analysis, regime-switching models, and the many other tools of econometrics and statistics capture information that topology cannot see: amplitude, frequency, distributional shape. The paper is explicit about this complementarity.

    What the topological construction offers is a new lens — one that responds to the relational structure of a time series rather than its metric structure. It asks not “how big is this change?” but “what does this change connect to, and what does it disconnect from, and is the answer different depending on which direction in time you are looking?”

    For systems where temporal asymmetry is a defining feature — business cycles, climate dynamics, physiological signals, causal event sequences — this lens may reveal structure that traditional tools, by their very construction, cannot see.

    The application to U.S. GDP is a proof of concept. The construction is general: it applies to any time series that can be turned into a directed visibility graph, which is to say, any time series at all. Whether the invariants it produces are useful features for classification, prediction, or interpretation in broader contexts is an empirical question that the paper opens but does not close.

    What it does establish is this: there exists a construction that takes a time series, produces two topologies from it, and quantifies the gap between them as a measure of temporal irreversibility. The construction is mathematically sound, mechanically verified, algorithmically tractable, parameter-free, and — when applied to the American economy across three turbulent decades — gives answers that make economic sense.

    That is a foundation worth building on.

    “Bitopological Spaces from Directed Graphs: Extending the Nada Construction to Capture Temporal Irreversibility” by José Mauricio Gómez Julián is available at Zenodo (v1.0.2, April 2026). The complete research compendium — Lean 4 formalization, R package, empirical dataset, and reproducibility notebook — is archived alongside it.

  • A DIALECTICAL MATERIALIST ANALYSIS ON PATH DEPENDENCE, IRREVERSIBILITY AND TELEOLOGY IN PHYSICAL SYSTEMS

    A DIALECTICAL MATERIALIST ANALYSIS ON PATH DEPENDENCE, IRREVERSIBILITY AND TELEOLOGY IN PHYSICAL SYSTEMS

    Why Broken Eggs Don’t Unbreak — A New Physics of History, Direction, and Purpose
    Physics • Philosophy • Foundations

    Why Broken Eggs Don’t Unbreak

    A philosopher argues that history, not mechanics, is the deepest grammar of the physical world — and that matter itself pursues stability.

    Based on a paper by José Mauricio Gómez Julián ~9 min read

    Why does a broken egg never reassemble itself? Every law governing the motion of its atoms is perfectly reversible — run the film backward and nothing in the mathematics complains. And yet, in the real world, broken eggs stay broken. This gap between what our equations allow and what nature actually does has haunted physics for over a century. A new paper from the University of Costa Rica proposes an answer that is as philosophically bold as it is mathematically precise.

    The Paradox That Won’t Go Away

    In the 1870s, the Austrian physicist Josef Loschmidt challenged Ludwig Boltzmann’s statistical explanation of the second law of thermodynamics. His objection was devastatingly simple: if the equations of motion work the same forwards and backwards in time, how can entropy — disorder — only ever increase? This is Loschmidt’s paradox, and it sits at the intersection of thermodynamics, quantum mechanics, and the philosophy of science.

    Physicists have proposed many partial answers. Some invoke the statistical improbability of reversal (there are astronomically more disordered states than ordered ones). Others appeal to cosmological boundary conditions — the universe simply started in a very special, low-entropy state. More recent work, cited in this paper, turns to information theory and Landauer’s principle (the idea that erasing information has a minimum physical cost).

    But the author — José Mauricio Gómez Julián, writing from the University of Costa Rica — finds all of these solutions insufficient. His central claim is provocative: the paradox is not in nature. It is in our models. We build theories that are fundamentally ahistorical, then act surprised when the real world — which is fundamentally historical — doesn’t obey them. The problem, he argues, is not that reality misbehaves. It is that our theories refuse to remember.

    The Core Move

    Instead of asking “Why is the macroscopic world irreversible?”, the paper reframes the question entirely: “Why do we insist on building reversible models and then call irreversibility a paradox?”

    The Philosophical Engine: Dialectical Materialism

    The paper is grounded in dialectical materialism — a philosophical tradition rooted in Marx and Engels, developed further by Hegel (in its idealist form) and by Soviet physicists like Blokhintsev and Rosental. If that sounds unusual for a physics paper, the author would say: that’s exactly the point.

    Dialectical materialism holds that reality is made of matter in motion, that contradictions are not flaws in our thinking but features of the world, and that quantitative changes eventually produce qualitative leaps. It also insists on a distinction that modern physics has muddled:

    • Epistemology — what we can know and measure (our limitations as observers).
    • Ontology — what actually exists in the world, regardless of our ability to observe it.

    This distinction turns out to be the paper’s sharpest tool. When quantum mechanics says that a particle’s position is “uncertain,” the author asks: is that uncertainty a feature of reality (ontology), or a feature of our knowledge (epistemology)? His answer is nuanced and consequential. The Schrödinger equation itself is fully deterministic — give it an initial state and a Hamiltonian, and it will predict the future wave function with perfect precision. The randomness enters only when we try to measure — when the quantum system meets our macroscopic instruments.

    In other words, probability in quantum mechanics is an epistemological resource — a powerful tool for managing complexity and incomplete knowledge — not a statement that reality itself is fundamentally random. This does not mean quantum mechanics is wrong. It means that its probabilistic character tells us about us, not about the universe.

    The author is careful, however, to distinguish this from classical Laplacian determinism. Heisenberg’s uncertainty principle, he argues, is ontological — it reflects genuine structural features of reality (the complementarity between position and momentum). So the world is deterministic in a deep sense, but not in the naive clockwork sense. It is deterministic in the way a complex, path-dependent system is deterministic: constrained, structured, lawful — but too intricate for any observer to fully predict.

    When the terrain disagrees with the map, trust the terrain

    Path Dependence: History Before Time

    Here is where the paper makes its most original conceptual move. The author argues that path dependence is more fundamental than time itself.

    What does this mean? In complex systems — economies, ecosystems, living organisms — the current state depends not just on the present conditions but on the entire sequence of events that led there. An economy with the same GDP, population, and technology as another can behave very differently because its institutions, crises, and policy choices followed a different historical path. The paper claims this is not just a feature of complex systems. It is a feature of all physical reality, at every scale.

    In this framework, time is reimagined. It is not a background parameter through which things happen (as in Newtonian mechanics), nor a dimension woven into spacetime (as in relativity), but rather a structure that records material transformations — a kind of universal memory. Space and time co-emerge as the necessary fabric for matter to develop and preserve its evolutionary trajectory.

    The author traces this insight back to Hegel’s philosophy of nature (published decades before Einstein’s relativity), where place, space, and time are understood as a unity — and where motion (and therefore matter) arises from their contradiction. The passage is striking:

    Place is spatial singularity… This perishing and regenerating of space in time and of time in space… is movement. This becoming… is the immediate, identical and existing unity of space and time: it is matter.
    — Hegel, Encyclopaedia of the Philosophical Sciences

    If path dependence is fundamental, then irreversibility is not something to be explained — it is something to be assumed, just as mathematicians assume the existence of natural numbers rather than proving it. Every broken egg, every aging star, every evolved species is evidence that the universe remembers its own history.

    A New Equation for a Remembering Universe

    The author doesn’t stop at philosophy. He proposes a concrete mathematical reformulation of the Schrödinger equation — the foundational equation of quantum mechanics — to make path dependence explicit.

    Standard quantum mechanics writes:

    Standard Form

    iℏ ∂ψ/∂t = Ĥ ψ

    where the Hamiltonian Ĥ describes the system’s energy at a given instant.

    The paper’s reformulation introduces a history-dependent Hamiltonian:

    Path-Dependent Form

    iℏ ∂ψ/∂t = Ĥ(t) ψ

    where Ĥ(t) = Ĥ₀ + ∫ K(t, t′) ψ(t′) dt′

    Here, t is not clock time but an interaction index — an ordering of how physical interactions emerged. The kernel K(t, t′) encodes how every prior interaction influences the current one.

    This is a bold move. It says: the state of a quantum system at any moment is shaped by the entire chain of interactions that brought it there — not just by its instantaneous configuration. The author also defines a metric on this “interaction space,” capturing the distance between interactions in terms of both complexity and energy change, and proves it satisfies the standard properties of a mathematical metric (nonnegativity, symmetry, triangle inequality).

    The elegance of this formulation is that it reproduces known physics as special cases:

    • Classical regime (low energy, macroscopic): the metric reduces to ordinary Euclidean geometry.
    • Relativistic regime (high velocity): it reproduces the Minkowski spacetime interval.
    • Quantum regime (entanglement, superposition): the metric captures quantum correlations through entanglement entropy.

    Entanglement Without Spookiness

    The path-dependent framework also offers a fresh take on quantum entanglement — Einstein’s famous “spooky action at a distance.” In the standard picture, measuring one entangled particle seems to instantaneously affect its partner, no matter how far away. In the author’s framework, entangled particles don’t communicate across space. They share a common interaction history. They are “close” in interaction space even when they are far apart in physical space — much like two points on a folded piece of paper that look distant but are actually adjacent when the paper is unfolded.

    Bell’s theorem — which proved that no local realistic theory can reproduce all quantum predictions — is not violated but reinterpreted: the “nonlocality” is real in emergent spacetime but disappears when you consider the deeper interaction space. Locality is preserved at the fundamental level; it only appears broken in the effective spacetime we observe.

    Why Does Matter Seek Stability?

    IV

    The paper’s second major argument is about teleology — the idea that natural processes are directed toward ends or purposes. This is a concept that modern science has largely banished (with some notable exceptions in biology). The author argues it should be restored — but in a materialist, not a mystical, form.

    The claim: physical systems universally tend toward maximum achievable stability within their material constraints. This is not an external force or an intelligent design. It is an inherent property of matter itself, arising from the internal contradictions within material systems.

    The evidence spans every scale of reality:

    • Cosmological: The universe’s laws appear “fine-tuned” for structure and complexity. Cyclic cosmological models suggest a drive to preserve laws conducive to stability.
    • Stellar: Stars burn through nuclear fuel, then transform — into white dwarfs, neutron stars, or supernovae — each outcome representing a reorganization toward the next achievable stable state.
    • Chemical: The pressure-induced transformation of graphite to diamond. Autocatalytic systems that reorganize when reactants deplete.
    • Biological: DNA’s role as a stable transcription template. Gould’s punctuated equilibria — long periods of stasis followed by rapid change. Insect metamorphosis.
    • Neural: The brain reorganizing through neuroplasticity after injury or during learning.
    • Social: Revolutions occurring when existing structures of production become incompatible with productive forces.

    In each case, the pattern is the same: systems seek stability, achieve it temporarily, exhaust the conditions that made it possible, undergo a qualitative transformation, and resume the search in a new configuration. Stability is the attractor; transformation is the mechanism.

    Least Action and Ground States

    The author connects this teleological perspective to two pillars of physics:

    The principle of least action — the mathematical rule that physical systems follow paths that extremize (usually minimize) the “action” functional. This is usually treated as a computational tool. The paper reinterprets it as a teleological law: systems select trajectories in service of their drive toward stability, and the path of least action is the one that best serves this purpose. Sometimes, the system does not take the absolute minimum energy path — because the absolute minimum may not serve the broader goal of sustained stability.

    The ground state tendency — quantum systems’ natural inclination to settle into their lowest energy configuration. The author, drawing on Solovej’s work on the stability of matter, argues this is not merely a mechanical outcome but an expression of matter’s fundamental need for stabilization. Electrons don’t “accidentally” fall into lower energy levels. They are driven there by the internal logic of material reality.

    Key Insight

    Teleology here is not purpose in the human sense — no intentions, no intelligence. It is the tendency of matter to resolve its own internal contradictions by seeking the most stable configuration available. Purpose arises from the inherent contradictions of matter, not from any external guide.

    The Arrow of Time, Revisited

    With path dependence as the foundation, the arrow of time becomes almost trivial to explain. Time flows in one direction because systems are historical. The future depends not only on the present state but on the entire trajectory that led to it. Irreversibility is not a statistical accident or a cosmological boundary condition — it is a structural feature of reality, as basic as the existence of natural numbers in mathematics.

    The author draws an analogy to the Cosmic Microwave Background (CMB) — the faint radiation left over from the early universe, which provides a natural “preferred frame” for cosmic observations without violating relativity. Similarly, the paper proposes that a preferred temporal direction can emerge from path dependence without requiring absolute time. Each observer may have their own “proper time,” but the causal structure — the chain of dependencies — is invariant and objective across all reference frames.

    This bridges a gap between quantum mechanics and general relativity. Quantum theory works with a notion of time closer to the classical (absolute) picture, while relativity treats time as relative and observer-dependent. The path-dependence framework offers a way to reconcile both: the ordering of interactions is fundamental and observer-independent; the measurement of time is relative.

    Testable Predictions

    The paper does not remain in the realm of philosophy. It proposes four concrete experimental protocols:

    • Decoherence studies: Prepare identical quantum systems, give them different interaction histories, and measure whether their decoherence rates differ. The paper predicts they will.
    • Entanglement analysis: Generate entangled photon pairs, expose them to different interaction histories, and check whether entanglement strength decays exponentially with “interaction distance” as the metric predicts.
    • Modified double-slit experiment: Introduce controlled interaction histories before particles reach the slits and look for history-dependent deviations in the interference pattern.
    • Time emergence clocks: Prepare identical atomic clocks with different interaction histories and compare their temporal evolution rates.

    These are technically demanding experiments — requiring millikelvin temperatures, ultra-high vacuum, single-photon detection, and high-fidelity quantum tomography — but they are within reach of current laboratory capabilities. The predictions are specific enough to be falsified, which is exactly what good science requires.

    Why This Matters Beyond Physics

    FOR THE NON-PHYSICIST

    If you are an economist, a political scientist, or simply someone who thinks about how societies change, this paper’s conceptual framework should feel familiar — and provocative.

    The concept of path dependence is already central to institutional economics (think of Douglass North or Paul David’s QWERTY keyboard). The idea that history matters — that you cannot understand a system’s current state without knowing how it got there — is a staple of comparative politics and historical sociology. What this paper does is argue that path dependence is not just a useful metaphor borrowed from physics. It is a fundamental feature of physical reality itself.

    Similarly, the paper’s concept of teleology without intention — systems pursuing stability through the internal logic of their own contradictions — resonates powerfully with Marx’s theory of historical materialism, where modes of production develop, exhaust their potential, and undergo revolutionary transformation. The author draws this connection explicitly, noting that revolutions occur “when existing relations of production become incompatible with developing productive forces.”

    And the distinction between epistemology and ontology — between what we can model and what actually exists — is a question every social scientist should take seriously. When our econometric models fail to predict a financial crisis, is the crisis a “black swan” (an anomaly), or is it evidence that our models are too ahistorical to capture reality?

    The problem is not that reality “contradicts” theory but that theory is a limited abstraction of reality, creating tension when attempting to make reality fit the model instead of developing models that capture reality’s historical-contextual nature.
    — José Mauricio Gómez Julián

    A Bridge Between Worlds

    This is not a paper that will convince everyone. Its philosophical framework — dialectical materialism — is unfamiliar and, for some, politically charged. Its mathematical proposals, while rigorous, are exploratory and await experimental confirmation. Its claim that teleology is a fundamental feature of matter will strike many physicists as a step backward toward pre-modern thinking.

    But that is precisely what makes it worth reading. In a landscape where theoretical physics has fragmented into string theory, loop quantum gravity, and various interpretations of quantum mechanics that all reproduce the same experimental results, a paper that asks “What if we’re starting from the wrong assumptions?” is exactly the kind of provocation that science needs.

    The paper’s deepest contribution may be methodological: a demonstration that philosophy and physics can inform each other without either colonizing the other. The philosophical framework provides the conceptual clarity to ask better questions. The physics provides the experimental discipline to test whether those questions have real answers.

    Whether or not its specific proposals survive experimental scrutiny, the paper succeeds in something more modest but no less important: it makes you see the broken egg differently. Not as a problem to be explained away, but as evidence of a universe that remembers — and that, in remembering, moves irreversibly forward.

    Original paper: “A Dialectical Materialist Analysis on Path Dependence, Irreversibility and Teleology in Physical Systems” by José Mauricio Gómez Julián, University of Costa Rica.

    Available as a preprint: OSF Preprints

    This post is an explanatory summary and does not represent the views of the author or any institution. Errors in interpretation are the blogger’s own.

  • Outlining a Dialectical Hypothesis On The C-Value Paradox In The Light of Quantum Chemistry

    Outlining a Dialectical Hypothesis On The C-Value Paradox In The Light of Quantum Chemistry

    Why an Amoeba Has 200 Times More DNA Than You — A Philosophical Take on the C-Value Paradox
    Explainers · Philosophy of Science · Molecular Biology

    The C-Value Paradox:

    Why an Amoeba Has 200 Times More DNA Than You?

    A philosopher argues that the way we count genes is broken — and proposes a dialectical, quantum-informed fix.

    Blog Post 2025
    ~ 9 min read

    Imagine you are handed two books. One is a slim novella; the other is an encyclopedia the size of a suitcase. Intuitively, you’d guess the encyclopedia contains more information. Now imagine that the novella turns out to encode the instructions for building an entire human being, while the suitcase-sized volume merely describes how to be a single-celled amoeba. Welcome to the C-value paradox — one of the most stubborn puzzles in modern biology — and to a recent paper that proposes a genuinely unusual way of thinking about it.

    The article in question is “Outlining a Dialectical Hypothesis on the C-Value Paradox in the Light of Quantum Chemistry” by the philosopher José Mauricio Gómez Julián, published in the Pitt Philosophy of Science archive (available here). It is not a typical biology paper. It moves fluidly between Hegelian logic, quantum mechanics, selfish genetic elements, and the mathematics of how we measure sets. If that sounds intimidating, don’t worry: by the end of this post, you’ll see why the argument matters — even if you’ve never opened a biology textbook.

    1. The Puzzle: More DNA, But Not More Complexity

    Let’s start with the basics. Every living cell carries a complete copy of the organism’s DNA — its genome. Biologists measure genome size in base pairs (bp) or, for convenience, in megabases (Mb), where 1 Mb = one million base pairs. This measurement is called the C-value.

    In prokaryotes (bacteria and archaea — the simplest forms of life, without a cell nucleus), the relationship is fairly intuitive: bigger genome, more genes, somewhat more complex organism. But when we turn to eukaryotes (everything from yeast to humans, with cells that contain a nucleus), the intuition collapses.

    A Few Striking Numbers
    Organism Genome Size (Mb) Gene Count (approx.)
    Yeast12~6,000
    Fruit fly180~14,000
    Human3,400~20,000–25,000
    Onion18,000
    Amoeba (A. dubia)686,000

    Sources: Latorre & Silva (2013); Pray (2022).

    A single-celled amoeba carries roughly 200 times more DNA than a human being. An onion needs about five times more DNA than we do. Amphibians, as a group, show genome-size variations of up to 91-fold. As the paper notes, citing Latorre and Silva, “it is hard to believe that this may reflect variations of nearly 100 times the number of genes necessary to give rise to the corresponding amphibians.”

    Nor is it simply a matter of how many genes there are. Even the raw count of protein-coding genes doesn’t track complexity well: a pufferfish has roughly the same number as a human (~35,000), and the rice plant has more (~51,000). The disconnect between genome size, gene number, and organismal complexity is the C-value paradox.

    2. Why Should Anyone Outside Biology Care?

    If you’re an economist, a political scientist, or a mathematician, you might be wondering what amoebae have to do with your work. The answer lies not in the biological details but in the type of reasoning the paper employs. Gómez Julián is making an argument about how we measure complexity — and specifically, why our standard tools for counting and measuring break down when the system we’re studying is fundamentally nonlinear.

    This is a problem that recurs everywhere: in financial markets (where small shocks cascade unpredictably), in political systems (where a single event can reshape an entire order), and in ecology (where species interact in webs, not chains). The C-value paradox is, at its core, a case study of what happens when you try to impose a linear accounting framework on a nonlinear reality.

    3. The Philosophy: What Does “Dialectical” Mean Here?

    The paper’s philosophical backbone comes from dialectical materialism — a tradition rooted in Hegel and adapted by Marx, Engels, and later Soviet philosophers. For readers unfamiliar with the term, here is the essence in plain language:

    Things are not only what they are in terms of their current state of development, but also their potential.

    In this framework, reality is a totality: not just what currently exists, but what could exist, what is coming into being, and what is being annihilated. The concept of “contradiction” is central — but not in the colloquial sense of a logical error. A dialectical contradiction means that any complex thing contains opposing developmental tendencies that are simultaneously complementary and mutually exclusive. These tendencies can be nonantagonistic (stable, coexisting) or antagonistic (destabilizing, eventually forcing the system to transform into something qualitatively new).

    Gómez Julián draws an explicit parallel between this philosophical notion and Bohr’s complementarity principle in quantum mechanics: to understand a quantum phenomenon fully, you need both the wave description and the particle description, even though they are mutually exclusive. The paper argues that this isn’t merely an analogy — it reflects a deeper logical structure shared across physics, chemistry, and biology.

    For those with an economics background, the parallel to dialectical reasoning in political economy is direct. Just as a commodity is simultaneously a use-value and an exchange-value — and you cannot understand the commodity by examining only one aspect — so a gene is simultaneously a physical structure (DNA sequence) and a functional agent (information carrier, regulatory element, or “selfish” replicator). Reducing it to just one dimension is precisely what creates the paradox.

    4. The Mathematical Core: Why Linear Counting Fails

    Now we arrive at what will interest the mathematicians and econometricians. The paper makes a precise mathematical claim: the tools we use to count genes assume linearity, but the genetic system is nonlinear.

    Formally, a function φ is called sigma-additive (or countably additive) if the measure of a union of disjoint sets equals the sum of the measures of each set. This is the standard foundation of probability theory and measure theory — the Kolmogorov axioms that every statistician and econometrician relies on.

    A subadditive function, by contrast, only requires that the measure of the union be less than or equal to the sum of the parts. Additive functions are a special case of subadditive ones. In genetics, if you use an additive model, you are assuming a perfect linear relationship between the number of allele copies and the organism’s traits — no dominance, no interaction, no epistasis. As Huang and Mackay (2016) showed, this assumption is empirically inadequate for most quantitative traits.

    Gómez Julián’s argument is that counting genes with sigma-additive functions implicitly treats the genome as a linear system: more genes = proportionally more complexity. But the evidence shows this is false. The complexity emerges from how genes interact, not from how many there are. Therefore, the counting function itself must change.

    5. What Actually Generates Complexity? Eight Factors

    The paper proposes that any meaningful relationship between gene count and organismal complexity must account for eight key aspects of the underlying molecular processes. Here they are, translated into plain terms:

    1. What kind of information is encoded? — Not all genes carry the same type of instruction. Some code for structural proteins; others regulate when and where those proteins are made.
    2. What encoding system is used? — The “language” of the genome is not uniform; different regions operate under different coding rules.
    3. Should we weight protein-coding genes more heavily? — Protein-coding genes make up only about 1.5% of the human genome. Should the other 98.5% count equally?
    4. What type of transcription occurs? — Through alternative splicing, a single gene can produce multiple different proteins. Humans may produce over 500,000 distinct proteins from only ~20,000 genes. The process is not one-to-one.
    5. DNA is a nonlinear dynamical system. — The double helix doesn’t behave like a simple linear chain. Researchers have modeled it using nonlinear Hamiltonians since at least the 1980s, and solitary conformational waves (solitons) can propagate along the strand.
    6. What type of gene is involved? — There are protein-coding genes, RNA genes, regulatory sequences, transposable elements, and more. They don’t all contribute to “complexity” in the same way.
    7. What role do “negative genes” play? — This is one of the paper’s most distinctive contributions. Gómez Julián renames so-called “selfish genes” as “negative genes” — borrowing the concept of negativity from dialectical philosophy. These are genetic elements (like transposons) that replicate for their own benefit, even if they are harmful or neutral to the organism. They exist in a state of unity and struggle with the organism’s “ordinary” genes, and this conflict is, according to Werren (2011), “an important driver of evolutionary change and innovation.”
    8. What happens during and around transcription? — This is when the DNA double helix unwinds and single strands are exposed. It is the moment of maximum vulnerability and maximum creative potential: DNA editing, trans-splicing, and tandem chimerism all occur here. The source of nonlinear complexity, the paper argues, is concentrated in this phase.

    If these eight factors could be incorporated into a new kind of counting function — one that captures nonlinear interactions, gene regulation, and the dialectical interplay between “positive” and “negative” genes — the paradox might dissolve. Genome size and gene number would, at least approximately, map onto organismal complexity.

    6. Quantum Chemistry Enters the Picture

    You might wonder: where does quantum mechanics fit into all of this? The paper’s answer is that the covalent bonds holding DNA together are quantum-mechanical phenomena. As early as the 1920s, Heitler and London showed that covalent bonds can be understood through the Schrödinger equation. The nucleotides in each DNA strand are linked by strong covalent bonds, so the strand’s dynamics — its rigidity, its unwinding, its conformational changes — are ultimately governed by quantum mechanics.

    In practice, solving the full Schrödinger equation for a molecule as large as DNA is computationally staggering. But progress is being made. The paper points to three recent advances:

    Computational Progress

    Analytical and numerical solutions of the Peyrard-Bishop DNA model (a nonlinear model of DNA dynamics) now show strong convergence (Al et al., 2020). Kink and localized solutions for the helicoidal version of the same model have been found and could serve as tools for modeling DNA-to-RNA transcription (Zdravković et al., 2019). And quantum annealing has been applied to de novo genome assembly — solving the combinatorial problem of stitching DNA fragments together using quantum and quantum-inspired optimization (Boev et al., 2021).

    These are early steps, but they suggest that the computational barriers to modeling DNA as a quantum-mechanical, nonlinear system are not permanent. Quantum computing may eventually make the Schrödinger-based analysis of large molecules feasible.

    7. The Bigger Picture: A Self-Teaching Universe

    At this point, the paper makes its most ambitious philosophical move. Drawing on research by Alexander et al. (2021), Gómez Julián describes a universe that is self-organized, deterministic, historically determined, and autodidactic — one that “evolves learning in an autodidactic way its own laws,” applying a process physically equivalent to biological natural selection at a cosmological scale. The universe, in this view, is a system that adds new nonlinearities to itself over time — a kind of spontaneous increase in complexity.

    This is linked to the concept of emergence: the spontaneous appearance of new information (new structures, new behaviors) as a result of a system’s internal dynamics. The laws of physics may themselves be subject to higher-order laws, just as a logic of a certain order is subject to the rules of a higher-order logic.

    For the C-value paradox, the implication is this: you cannot understand the parts (genes) without understanding the whole (the organism and its evolutionary history), and you cannot understand the whole without understanding how it emerged from the parts. The truth, as Hegel would say, is in the totality.

    · · ·

    8. So What Would a Solution Actually Look Like?

    Gómez Julián is careful to say that his paper is a guide, not a solution. He proposes the construction of a “paradox-free gene counting function” (PFGCF) — a new mathematical object that would replace simple sigma-additive counting with something capable of capturing:

    • Nonlinear gene interactions
    • The role of alternative splicing and regulatory elements
    • The dialectical interplay between ordinary genes and “negative” (selfish) genes
    • Quantum-mechanical properties of DNA structure
    • What happens during and around transcription

    This function might not even be a single function at all, but rather a family of functions, each capturing different aspects of genomic complexity. The construction will require, the paper argues, “philosophers, chemists, geneticists, and physicists, as well as the use of high-capacity computational equipment.”

    It is, in the author’s own words, a “legitimate speculation” — grounded in established science but not yet experimentally verified. The value of the paper lies in its identification of which factors matter and what kind of mathematics is needed, rather than in providing a finished model.

    9. Why This Paper Matters (Even If You’re Not a Biologist)

    Let’s return to the question of why a non-biologist should care. Here are three reasons:

    The whole is more than the sum of its parts — and the tools we use to count the parts must reflect that.

    First, the paper is a case study in interdisciplinary thinking. It weaves together philosophy, mathematics, chemistry, and biology in a way that is rare in any field. Whether or not you agree with its dialectical-materialist framework, the attempt to build a bridge between Hegel and quantum chemistry is intellectually stimulating.

    Second, it highlights a general methodological problem: when linear tools fail, what replaces them? Economists face this when GDP doesn’t capture well-being; political scientists face it when vote counts don’t capture democratic health; mathematicians face it whenever measure theory meets real-world complexity. The paper’s call for new counting functions is, at bottom, a call for new mathematics.

    Third, it reminds us that paradoxes are productive. The C-value paradox has been around for decades and hasn’t been solved — but it has forced biologists to discover alternative splicing, transposable elements, non-coding RNA, and epigenetic regulation. The paradox was never a dead end; it was a signpost pointing toward deeper truths. That’s a lesson every discipline can take to heart.

    · · ·

    You can read the full paper by José Mauricio Gómez Julián at the PhilSci Archive: https://philsci-archive.pitt.edu/24513/

  • General Dynamic Parameter Models via Reference Anchoring

    General Dynamic Parameter Models via Reference Anchoring

    You can also find this library at CRAN and download it directly from R and RStudio.

    Also, we recommend viewing the mind map summary at the end of the article to better understand the relationship between the functions of the package.

    R Library Review

    Meet gdpar

    General Dynamic Parameter Models via Reference Anchoring

    In the fleeting calculus of a two-second decision—overtaking a car on a narrow road—the human brain performs a remarkable statistical trick. It does not build a model of the approaching driver from scratch. Instead, it retrieves a baseline: the average driver, representing typical reaction times and modal aggression. In a split second, it reads the specific signals of the actual driver—relative speed, vehicle type, micro-movements—and estimates how this specific driver deviates from the baseline. The decision to overtake emerges from that synthesis.

    This cognitive recipe—population reference + individual deviation—is the philosophical bedrock of the R package gdpar (General Dynamic Parameter models via Reference Anchoring) by José Mauricio Gómez Julián. The package takes this intuition, formalizes it as a rigorous statistical decomposition, proves the conditions under which it is mathematically identifiable, ships a Stan-based Bayesian engine to estimate it, and layers on causal inference, geometry-adaptive sampling, and dependence-robust inference.

    The Anatomy of Deviation

    Every layer of gdpar is an elaboration of a single, elegant equation. For each observation $i$ with covariates $x_i$:

    $$ \theta_i \;=\; \theta_{\text{ref}} \;+\; \Delta(x_i,\; \theta_{\text{ref}}) $$

    Read it as: the parameter of individual $i$ equals a population reference, plus a deviation that is itself a function of the individual’s covariates and of the reference itself.

    That final clause is where the architecture pivots from classical statistics. The deviation $\Delta$ does not merely depend on who you are (your covariates $x_i$); it depends on what the reference is. If you transplant the model to a new population, the deviation function behaves differently because $\theta_{\text{ref}}$ is one of its arguments. This structural dependence is the defining feature of “reference anchoring.” It distinguishes gdpar from random-effects or varying-coefficient models, where the deviation is structurally separate from the reference.

    So, what is the shape of $\Delta$? The package singles out a specific functional form called the Additive–Multiplicative–Modulated (AMM) decomposition:

    $$ \Delta(x,\theta_{\text{ref}}) \;=\; \underbrace{a(x)}_{\text{additive}} \;+\; \underbrace{b(x)\odot\theta_{\text{ref}}}_{\text{multiplicative}} \;+\; \underbrace{W(\theta_{\text{ref}})\,x}_{\text{modulated}} $$

    Three mechanisms, cleanly separated and independently interpretable:

    • $a(x)$ — A pure additive shift. Think of this as a traditional fixed-effect driven by covariates.
    • $b(x)\odot\theta_{\text{ref}}$ — A covariate-dependent scaling of the reference (using the Hadamard/elementwise product). This is where “the deviation depends on the reference” enters multiplicatively.
    • $W(\theta_{\text{ref}})\,x$ — Covariates are mixed through a matrix $W$ that is, itself, tuned by the reference. This is the explicit, structural reference-dependent channel.

    Standard models drop out as special cases. Set $\Delta \equiv 0$ and you have fixed-effects regression. Set $W \equiv 0$ and you have a hierarchical model with multiplicative interaction. Set $b \equiv 0$ and you have a varying-coefficient model. The AMM is the smallest natural family that contains all three and elevates the reference to an active argument of the deviation.

    The Three Estimation Engines

    gdpar defines three complementary engines for estimating $\Delta$. Crucially, only one is executable in the current release—a deliberate choice to promise a mathematical scope that exceeds the executable surface, and to say so honestly.

    Path Engine Representation Status
    Path 1 Hierarchical Bayesian (Stan) Parametric AMM ✅ Operational
    Path 2 Varying-coefficient (splines) Smooth $\beta(z)$ 🚧 Conceptual
    Path 3 Hypernetwork / Neural Net Net generates $\theta_i$ 🚧 Conceptual

    Paths 2 and 3 are documented to “reference grade”—full asymptotic theory (contraction rates, Bernstein–von Mises) is developed in the Wiki—but they abort with gdpar_unsupported_feature_error if invoked. Path 1 places priors on every component ($\theta_{\text{ref}}, a, b, W$) and samples the joint posterior with HMC, yielding native, full-posterior uncertainty.

    A Tale of Two Posteriors: EB vs. FB

    Within Path 1, gdpar offers two inferential regimes. Full Bayes (FB) via gdpar() samples the joint posterior, remaining most faithful to the cognitive analogy. Empirical Bayes (EB) via gdpar_eb() estimates the hyperparameters by maximizing a marginal likelihood via a Laplace approximation, then samples the remaining parameters conditionally.

    The EB vs FB Comparator

    Rather than forcing a choice, gdpar treats them as parallel routes. It ships a dedicated comparator, gdpar_compare_eb_fb(), which quantifies agreement on $\theta_{\text{ref}}$ and the reduced parameter vector $\xi$. The Wiki develops the theory to first-class depth: EB and FB lower-level posteriors agree asymptotically (Theorem 7A), while EB intervals under-cover by $O(n^{-1})$ (Proposition 7B). If you have ever wondered if EB is “good enough” for your data, gdpar lets you answer that empirically.

    Distributional Regression: Every Parameter is a Slot

    gdpar is not constrained to modeling the mean. A probability distribution has multiple parameters—location, scale, shape, tail index, zero-inflation probability—and each one can carry its own AMM decomposition. The package indexes these by $k = 1, \dots, K$:

    $$ \theta_i^{(k)} = \theta_{\text{ref}}^{(k)} + \Delta^{(k)}(x_i, \theta_{\text{ref}}^{(k)}), \qquad k = 1, \dots, K $$

    The built-in roster covers Gaussian, Poisson, negative binomial, Bernoulli, Beta, Gamma, Student-$t$, Tweedie, ZIP, ZINB, and hurdle families. Zero-inflated and hurdle models receive an especially elegant treatment: both the zero-inflation probability $\pi_i$ and the count parameter $\theta_i$ are anchored to their respective references—a dual deviation design.

    The Causal Bridge

    Because the AMM form produces individual parameters, individual treatment effects emerge naturally. gdpar_causal_bridge() implements a T-learner: fit the anchored model separately under treatment and control, then read the conditional average treatment effect (CATE) at $x_i$ as the difference of the anchored individual predictions:

    $$ \widehat{\tau}(x_i) = \widehat{\mu}_1(x_i) – \widehat{\mu}_0(x_i) $$

    A second layer, gdpar_compare_meta_learners(), benchmarks the AMM-based learner against external meta-learners via pluggable adapters: grf::causal_forest on the R side and EconML’s CausalForestDML on the Python side (via reticulate). The framework’s causal claims are benchmarked, not asserted.

    Mechanics & Clockwork

    Several engineering decisions elevate gdpar from a theoretical exercise to a serious computational environment:

    • Stan Code Generator: Composes programs from canonical pieces—AMM blocks for $p=1$ and $p \geq 1$, EB marginal/conditional blocks, distributional-$K$ blocks—selected by the resolved $(K, p, \text{family}, W, \text{parametrization}, \text{group})$. The $W$ basis supports B-splines with Stan-side Cox–de Boor evaluation, ensuring differentiability inside HMC.
    • Identifiability Pre-flight: Before any sampling, gdpar_check_identifiability() runs a Gram-matrix check (Proposition 1C), a per-coordinate cross-component check (C4-bis) for $p > 1$, and a per-group anti-aliasing check (C7). If your design is non-identifiable, you find out before the sampler burns your CPU, accompanied by a structured gdpar_identifiability_error naming the dependent directions.
    • Data-Driven Reparametrization: Treats the parametrization of $b(x) \odot \theta_{\text{ref}}$ as a pre-fit decision. A short pilot computes an information ratio, dispatching to CP, NCP, or—gdpar‘s root-cause resolution—a linear reparametrization that samples the product $\theta_{\text{ref}} \cdot b$ directly, sidestepping bilinear funnels altogether.

    Opt-in Power Tools

    Two advanced capabilities are switched off by default, documented as thoroughly as the core path.

    1. Geometry-Adaptive Sampling

    Hierarchical AMM posteriors can be geometrically hostile—funnels, near-determinism, heavy tails. The opt-in geometry engine climbs a ladder of Riemannian metrics: Euclidean → Fisher/SoftAbs → sub-Riemannian → relativistic/Finsler. A certifying orchestrator diagnoses the pathology, selects a metric, tunes the integrator, and emits a certificate. If full sampling is certified infeasible, a Laplace fallback provides a plug-in posterior with ELPD on par with mgcv-REML or INLA-Laplace.

    2. Dependence-Robust Inference

    gdpar does not model temporal or spatial dependence in its point structure; instead, it makes the inference robust to dependence (a working-independence + sandwich-variance stance in the spirit of Liang & Zeger, 1986). You receive diagnostics (Durbin–Watson, Ljung–Box, Moran’s $I$) and robust SEs via block bootstrap—moving or circular blocks in time (with the Politis–White flat-top automatic block length), tiled randomized-origin blocks in space. Point estimates remain pristine; only the uncertainty is made honest.

    ⚠️ Honest Limitations

    The Wiki is admirably forthright about scope. Only Path 1 is executable in 0.1.0. Dependence is not modelled—only the inference is made robust. The package’s mathematical scope exceeds its executable surface by design. Read the “Implementation status” notes carefully before relying on a feature.

    TL;DR

    gdpar takes one of the most natural ideas in human prediction—predict an individual as a deviation from a population reference, where the deviation itself depends on the reference—and transforms it into a fully specified, identifiability-checked, Stan-powered Bayesian regression framework. It is theoretically rigorous, computationally serious, and unusually honest about what it does and does not yet do. If your work involves individual heterogeneity, distributional regression, or causal effect estimation with principled uncertainty, gdpar demands a careful look.