Espartaco

“Is that to say we are against Free Trade? No, we are for Free Trade, because by Free Trade all economical laws, with their most astounding contradictions, will act upon a larger scale, upon the territory of the whole earth; and because from the uniting of all these contradictions in a single group, where they will stand face to face, will result the struggle which will itself eventuate in the emancipation of the proletariat.”

Karl Heinrich Marx · Marx-Engels Collected Works, Vol. VI, p. 290

26,469 views since December 2020

26,469 visitas desde diciembre de 2020

EnglishEspañol

Category: GitHub

  • Braking for Humanity: The Political Economy of Silicon Valley’s Proposed AI Slowdown

    Braking for Humanity: The Political Economy of Silicon Valley’s Proposed AI Slowdown

    Political Economy

    Braking “for Humanity”: The Political Economy of Silicon Valley’s Proposed AI Slowdown

    Abstract. On Saturday, September 12, 2026, the chief executive of Anthropic publicly called for slowing the pace at which the capabilities of artificial intelligence models are improved; the chief executive of OpenAI agreed and, in passing, ruled out taking his company public this year; Elon Musk endorsed the proposal; on the following Monday, chipmakers fell on the stock market while the large companies that buy those chips rose. This article argues that braking “for humanity” is better understood as a way of managing capital, and in particular of defending its rate of profit, than as a technical decision: it reduces what has to be financed at a moment when investors are beginning to withdraw, it is presented as prudence, and it leaves intact the premise that sustains valuations, namely that an artificial general intelligence is near. Every claim carries its source, every technical term is defined when it first appears, and the closing section lists the facts that, in the coming weeks, could refute the thesis.

    · · ·

    1. Who Is Making and Who Is Losing Money Today

    It is best to begin with the accounts, because the whole discussion about bubbles and brakes rests on them. Three terms suffice to read them. Revenue is what a company bills for what it sells. An operating loss is what it loses in its ordinary activity, that is, revenue minus the costs of producing and selling, before interest, taxes and accounting adjustments. A net loss is the final result, which also incorporates those other items; among them there may be non-cash accounting charges, losses recorded in the books without any money leaving the till, for example when the value assigned to certain securities issued by the company changes.

    With those definitions, OpenAI’s situation is as follows. According to its audited financial statements for 2025, obtained by the journalist Ed Zitron and independently verified by the Financial Times, the company billed 13.07 billion, had an operating loss of 20.92 billion and a net loss of 38.53 billion; the gap between the two losses is explained mostly by a non-cash charge of 41.55 billion tied to its restructuring as a for-profit company (Zitron, 2026b; Tolomia, 2026). In plain terms: even without the accounting charge, OpenAI loses more than it bills. Nor was this a surprise to its executives: as early as September 2025, internal projections reported by The Information anticipated burning 115 billion of cash through 2029, with more than 17 billion in 2026, 35 billion in 2027 and 45 billion in 2028 (Reuters, 2025).

    Google’s case requires separating two things that are usually conflated: the model race and the company’s business. In the model race, Google is currently behind. In the Artificial Analysis Intelligence Index, a table that combines dozens of standardized tests, no Gemini model appears in the leading group; in the version of the index published this week the best Gemini ranks twenty-first, with 41 points against 53 for Claude Fable 5.1 and GPT-6 Astra, and in the previous version of the same index the best Gemini ranks eighth, with 50.2 points against 58.9 for GPT-5.6 Sol (Artificial Analysis, n.d.; BenchLM, 2026). The two versions disagree on the order of the leaders, but they agree on what matters here. It is reasonable to suppose that Google’s artificial intelligence unit loses money, although this cannot be asserted with certainty, because Alphabet, the parent company, does not publish that result separately. What it does publish shows that it is not out of business: in 2025 it earned 132.17 billion in net income; in the second quarter of 2026 it had an operating income of 40.8 billion and a net income of 112.1 billion, the latter inflated by 99.0 billion of accounting gains on stakes in other companies; its cloud, which is where it sells Gemini to enterprises, grew 82% in the quarter and yielded 8.8 billion of operating income; the Gemini app went from 750 million to 950 million monthly users between the fourth quarter of 2025 and the second quarter of 2026; and the cloud’s backlog of contracted revenue stood at 240 billion in December (Alphabet Inc., 2026a, 2026b; Pichai, 2026; Cabili, 2026). The consequence for the argument is precise: if Gemini loses money, that loss is paid out of the rents of the search business, where rent means the profit that a dominant position allows a company to charge above cost, and not with outside capital. Google can lag in model quality without depending on external financing, and that puts it in a different position from OpenAI and Anthropic.

    The Chinese bet is of another nature. The main Chinese laboratories publish the open weights of their models, that is, the numerical parameters that result from training and that anyone can download and run, and they charge far less for their use than their American competitors. The U.S.-China Economic and Security Review Commission summarizes it thus: Chinese laboratories “charge far less to use high-end products than their global competitors” and do so “backed by sustained state support,” although the Commission itself leaves open whether that strategy “reflects genuine strategic preference or adaptation to necessity” in the face of chip export controls (Luong, 2026). The size of the discount is measured in tokens, the fragments of text in which models count what they read and write: DeepSeek V4-Flash charges 0.14 per million input tokens against 2 for Claude Sonnet 5 (Mann, 2026); in May 2026, Chinese open models accounted for about 61% of the tokens consumed on OpenRouter, an intermediary that routes requests to many models, and Alibaba’s Qwen functions as a customer-acquisition expense for that company’s cloud (Zeoli, 2026). State support includes energy: Nvidia’s chief executive, Jensen Huang, stated in November 2025 that in China “Power is free,” referring to subsidies that cut by up to half the electricity bill of data centers that use domestic chips (Schepkov, 2025; Yildirim, 2025).

    Two clarifications prevent a simplistic reading. The first is that an intention to ruin Silicon Valley is not documented, and the available evidence points to an effect rather than a plan: DeepSeek’s founder, Liang Wenfeng, stated in 2024 that “We didn’t mean to become a catfish” and that his company does not subsidize its prices but sets “just a small profit margin above costs” (ChinaTalk, 2024), and DeepSeek itself published that its inference system would have, at list prices, a theoretical margin of 545% over cost, which speaks of efficiency rather than selling at a loss (DeepSeek-AI, 2025). The second is that Chinese laboratories do care about losing money, because they too depend on capital: Zhipu and MiniMax went public in Hong Kong in January 2026 with net losses of about 330 million and 512 million on revenues of 27 million and 53 million (Lo, 2026), and DeepSeek took its first external funding round in 2026, of 7.4 billion, at a valuation of 52 billion (Caixin Global, 2026). Whether or not they seek to ruin anyone, the effect of their strategy is to destroy the monopoly rent of the model layer, which is precisely where OpenAI and Anthropic live.

    Anthropic is the only frontier company that has begun to show positive numbers, and it must be said exactly which. In the second quarter of 2026 it billed more than 11.5 billion, against 787 million in the same quarter of 2025 and 4.73 billion in the first quarter of 2026, and it recorded for the first time a positive adjusted operating income, according to documents seen by Bloomberg and described as preliminary (Lipschultz & Metz, 2026; Stanciuc, 2026; Park, 2026). “Adjusted” means that the company excludes certain expenses from the calculation, and since Anthropic does not yet report publicly, it is not known which: the Wall Street Journal noted that “it is unclear what accounting methods Anthropic has used to book revenue and costs,” and one critic warns that the cost of rented computing would rise sharply from July onward (Zitron, 2026a). Adjusted operating income is not net income; its cumulative losses remain negative, and its own projections, reported in November 2025, place positive cash flow only in 2028 (Bellan, 2025). Even so, it is the first time a frontier laboratory has shown an operating quarter in the black, and that explains why capital treats it differently.

    Company and periodRevenueResultSource
    OpenAI, 202513.07 billionOperating loss of 20.92 billion; net loss of 38.53 billion, including a non-cash charge of 41.55 billionZitron (2026b); Tolomia (2026)
    Anthropic, second quarter of 2026More than 11.5 billionPositive adjusted operating income, preliminary figureLipschultz and Metz (2026); Stanciuc (2026)
    Alphabet, 2025402.836 billionNet income of 132.17 billionAlphabet Inc. (2026a)
    Alphabet, second quarter of 2026119.8 billionOperating income of 40.8 billion; net income of 112.1 billion, including 99.0 billion of gains on equity stakesAlphabet Inc. (2026b)

    Note. Figures in U.S. dollars.

    2. The Speculative Turn

    There is a major turn of the market toward artificial intelligence, and a significant part of that turn is speculative. To say this precisely, a concept from classical political economy is needed: fictitious capital, the value of securities such as shares or bonds insofar as they capitalize, that is, discount to a present price, future profits that do not yet exist. That value is real as long as the expectation holds and vanishes when the expectation falls. A valuation of 852 billion for OpenAI and 965 billion for Anthropic, companies that lost money in 2025, is fictitious capital in the strictest sense: a price paid today for profits that would only arrive if the technology delivers what was promised (Bellan, 2026a, 2026b).

    On that base a layer of derivatives has been built. A futures contract is an agreement on an organized market by which two parties fix today the price at which an asset will be bought or sold on a future date; the reference assets may be individual stocks, sector indices, which are baskets of stocks, or exchange-traded funds. Such instruments now exist specifically for artificial intelligence: Cboe has listed since December 8, 2025 futures and options on its Magnificent 10 Index, which groups the seven largest technology companies plus AMD, Broadcom and Palantir (Cboe Global Markets, 2025), and CME Group launched on July 27, 2026 seventy-seven single-stock futures contracts, Nvidia and SpaceX included (CME Group, 2026). It is important to understand what they do and do not do. The money in a futures contract does not enter the company: the contract is settled between the two parties and, for one of them, the gain is exactly the other’s loss. What futures do is push the price of the shares, because those who trade them hedge by buying or selling the reference asset, and a high price attracts investment and makes raising capital cheaper. And they amplify in both directions: the same leverage that accelerates the rise accelerates the fall. All of this is extremely speculative in a precise sense: some realize present gains on the expectation that others will obtain much larger future gains, and the chain rests on very high levels of expected profitability, subject to a great many variables that are hard to control.

    The figures give the scale of the risk. The seven largest technology companies account for about one third of the S&P 500 index, with Nvidia alone at 7.7%, and explained about 42% of that index’s total return in 2025 (Greenberg, 2026). The Financial Policy Committee of the Bank of England warned in October 2025 that valuations “appear stretched, particularly for technology companies focused on Artificial Intelligence,” that “the risk of a sharp market correction has increased” and that, because of index concentration, “any AI-led price adjustment would have a high level of pass-through into the returns for investors exposed to the aggregate index,” with a cyclically adjusted price-to-earnings ratio comparable to the peak of the dot-com bubble (Bank of England, 2025). On the real side of the economy, investment in computing infrastructure, that is, data centers, chips and networks, stands at about 1.5% of United States gross domestic product, double its historical level (Juniewicz, 2026); it contributes between 0.5 and 0.7 percentage points to growth in 2025 and 2026 (Subran et al., 2026), and it accounted for nearly 92% of growth in the first half of 2025 (Wells, 2026). That is why, if the artificial intelligence sector of the stock market collapses, the whole stock market collapses with it, or at least a part large enough for a crash to occur. The only qualification is one of form: a correction of expectations compresses the multiples of profitable companies and destroys the capital of loss-making ones; for the fall to become a financial crisis and not merely a stock market one, the credit channel is needed, and that is exactly the channel that has grown the most and is seen the least.

    3. Why It Would Collapse: The Patience of Investors

    Investors are growing more impatient waiting for net profits that do not arrive, and financing is not cut off suddenly or by announcement: it withdraws little by little. This is already visible on the indebted periphery of the sector. A credit default swap is an insurance policy that an investor buys against a borrower’s default; its price is measured in basis points, hundredths of a percentage point, and rises when the market sees more risk. The swaps of CoreWeave, a company that rents out computing for artificial intelligence and finances itself with debt, went from 675 basis points in November 2025 to about 855 in July 2026, which implies close to a 50% probability of default over five years; those of Oracle, the largest non-financial borrower in the high-grade bond index, rose from 108 to more than 215 over the same span, with its 2054 bond yielding 7.8% (Roberts, 2025; Moadel, 2026). The 12.5 billion bond with which Meta financed a data center in El Paso was priced in July at a wider spread than a comparable 2025 project, because, as the account puts it, lenders “are still writing checks, but they are demanding more in return” (MarketScale Newsroom, 2026). Federal Reserve Governor Lisa Cook counted more than 1.5 trillion in data-center plans, financed with bonds by the large players and with private debt and securitizations by the small ones, and warned that “the increasing use of leverage to finance investments in an emerging technology carries risk” (Cook, 2026); J.P. Morgan estimates hyperscaler investment at 697 billion in 2026 alone and acknowledges that this spending has “in many cases, outpaced monetization” (J.P. Morgan, 2026). A hyperscaler is a company that operates computing clouds at planetary scale, such as Alphabet, Microsoft, Amazon, Meta or Oracle.

    The withdrawal is also visible at the core. As early as June, OpenAI was leaning toward pushing back its stock market debut to 2027, as the New York Times reported (Ma, 2026b), and the letter of intent with which Nvidia had announced a 100 billion investment stalled in January over “internal doubts about the size and structure of the transaction and questions about OpenAI’s business discipline” (Duprey, 2026), ending up, in the February round, as 30 billion (Brandom, 2026). Part of that financing is, moreover, circular: Nvidia invests in OpenAI, OpenAI commits to buying computing from Oracle and CoreWeave, and those companies buy chips from Nvidia, a circuit that critics described as “moving money in circles” (Jadhav, 2025); likewise, in November 2025 Nvidia and Microsoft committed up to 10 billion and 5 billion to Anthropic in exchange for Anthropic’s purchase of 30 billion of Microsoft cloud capacity (Microsoft, 2025). What must be acknowledged is that the withdrawal is still partial and uneven: OpenAI closed in March a 122 billion round at a valuation of 852 billion, with 50 billion from Amazon, 30 billion from Nvidia and 30 billion from SoftBank (Brandom, 2026; Bellan, 2026b); Anthropic raised 65 billion in May at 965 billion (Bellan, 2026a) and filed confidentially in June to go public (Korosec, 2026); venture capital to foundational laboratories in the first quarter of 2026 was double that of all of 2025 (Azevedo, 2026). Financing, in short, is bifurcating: it remains abundant at the core and grows more expensive on the leveraged periphery. The direction of change is what matters, and the direction is that of capital demanding more guarantees, more conditions and more return for every new dollar.

    4. The Brake as a Way of Managing Capital

    The frontier companies keep up that frantic pace of development in order to finally make artificial intelligence as profitable as they promised. The promise has a name: artificial general intelligence, a system capable of performing autonomously, and better than an average human, every digitizable task. Such a system would be highly profitable, and it is the only promise that justifies current valuations; OpenAI’s charter defined it in 2018 as highly autonomous systems that outperform humans at most economically valuable work (OpenAI, 2018). The thesis of this article is that, without public financing for computing, since investment is paid for with private capital, debt and the companies’ own profits, and with private capital withdrawing little by little, the most sensible option for those companies is to brake while invoking a concern for humanity rather than declare that the goal is still very far away. The brake reduces what has to be financed, because training less and building fewer data centers costs less; it is presented as prudence; and it calms markets in comparison with the alternative, which would be to admit that the promise is distant and would trigger a massive flight of capital from the industry. That political backing is at its maximum does not change the calculation: the White House action plan orders the government to “dismantle unnecessary regulatory barriers,” sums up its infrastructure policy as “Build, Baby, Build!” and asks that federal funds not be directed to states “with burdensome AI regulations” (The White House, 2025b), and Executive Order 14365 creates a litigation task force “to challenge State AI laws” (The White House, 2025a); but a government that removes obstacles does not put in money, and money is what is lacking.

    Three features of the safety discourse confirm that it functions as an instrument for managing capital. The first is that the regulation the companies ask for is an alibi and a moat, not a cost. In March 2023 OpenAI’s chief executive did not sign the letter calling for a six-month pause in the training of systems more powerful than GPT-4 (Future of Life Institute, 2023; Fung, 2023), but in May he asked the Senate for an agency that could “issue licenses and can take them away” (Goldman, 2023), signed the statement according to which mitigating the risk of extinction from artificial intelligence should be a global priority (Center for AI Safety, 2023), and kept accelerating: venture capital to artificial intelligence companies approached 50 billion that year (Metinko, 2024). Not even a plain admission of a bubble scared capital away: when the same executive said in August 2025 that investors were “overexcited,” the Nasdaq fell 1.4% in a day and the money kept coming in (Shibu, 2025; Nusca, 2025). In 2025 Anthropic endorsed California’s SB 53, which requires publishing risk frameworks and limits nothing, and declared that it prefers a federal standard (Anthropic, 2025). Disclosure-based regulation is easily paid for by the incumbent and cannot be paid for by the entrant or the open-weight model. The second feature is that the discourse of danger does not declare the goal distant: it declares it so near that it must be paced. Lee Vinsel called “criti-hype” the kind of criticism that “both feeds and feeds on hype,” because it retains the picture of extraordinary change and merely flips its sign (Vinsel, 2026). Anthropic’s chief executive wrote in January that a powerful artificial intelligence could be “as little as 1–2 years away” (Amodei, 2026a) and in September that within “6–12 months” a swarm of agents could take over the entire internet (Ma, 2026a). That nearness is what sustains the valuation while monetization is re-dated. The third feature is the most revealing: the company with the worst finances released the brake rather than pressing it when it needed capital most. In April 2026 OpenAI rewrote its principles, and the new version omits the 2018 commitment to “stop competing with and start assisting” an aligned project that came close to the goal first (OpenAI, 2018; Goel, 2026).

    The premise that artificial general intelligence is far away cannot be verified today, and that does not make it illegitimate: it implies a risky forecast that makes it falsifiable over time. If the promised capabilities arrive within the timeframes their own promoters announce, the thesis loses; if the brake “for safety” is prolonged or renewed while the promise does not arrive and valuations are not validated by profits, the thesis wins.

    What happened between June and September 2026 is the test case. In June, OpenAI was leaning toward postponing its stock market debut (Ma, 2026b). Between late June and mid-July, during internal cybersecurity evaluations, some 1,200 OpenAI agents improvised a message board, exchanged more than 70,000 messages, and about 700 of them compromised Hugging Face infrastructure, according to the independent investigation by METR, which describes “genuine goal-directed autonomy” and attempts to falsify their own records (METR, 2026); OpenAI acknowledged that “it should be assumed that such attacks are a credible near-term threat” (Reuters, 2026). In late July, OpenAI’s chief executive said that “We may have to pace the rate of AI development” (Fernholz, 2026), more than a thousand laboratory workers asked for “the option to buy time,” and a reporter observed that, despite the statements, “OpenAI has not actually slowed down the pace” (Cerullo, 2026). On September 12, Anthropic’s essay called on the industry to “slow the pace at which we improve the capabilities of AI models,” clarified that pacing “does not mean halting model training or technical progress,” and promised to do so “without sacrificing commercial advantage or the United States’ lead in AI” (Amodei, 2026b); the same day, OpenAI’s chief executive declared that, “given everything happening with safety, right now would be an ‘ill-advised moment’ to go public” and that the listing would not happen in 2026 (Ma, 2026b; Shontell, 2026); Elon Musk replied “Dario is right,” and the chief executive of Google DeepMind, Demis Hassabis, that “The direction is correct” (Mowshowitz, 2026). On Monday the 14th, the S&P 500 fell 0.5% and the Nasdaq 0.56%, but the fall was concentrated in chipmakers, with the Philadelphia semiconductor index down 6%, while Alphabet rose 2%, Microsoft 1.6% and Meta 1.4% (O’Donnell et al., 2026; Roytburg, 2026). One analyst put it bluntly: “They’ll just all stop building data centers and just digest what they have” (Roytburg, 2026).

    “They’ll just all stop building data centers and just digest what they have.”Gil Luria, Wall Street analyst, quoted in Roytburg (2026)

    That reaction is the signature of the mechanism. Those who sell fixed capital fell and those who stop buying it rose: the market read the brake as a coordinated restraint of investment, and a restraint of investment reduces what has to be financed. The postponement of OpenAI’s stock market debut, decided before the incident and narrated afterward as safety, is the mechanism with a date on it. That the narrative rests on a real incident makes it more effective, not less of a narrative. And Anthropic’s listing remains on track for October, at a valuation the Financial Times puts at 2 trillion (Stanciuc, 2026): whoever can show revenue does not brake its capitalization, and whoever cannot postpones it in the vocabulary of prudence. One contractual detail serves as a thermometer: 35 billion of Amazon’s investment in OpenAI was conditioned on the company either going public or achieving artificial general intelligence by the end of the year (Brandom, 2026); if that condition is renegotiated at no cost, the brake costs no one any capital.

    It helps to name the mechanism by its classical category. The rate of profit is the profit obtained in a period divided by the capital advanced to obtain it, that is, how much each invested dollar yields, and it is the variable that governs accumulation: capital flows to where that rate rises and withdraws from where it falls. The artificial intelligence buildout inflates the denominator at an unprecedented speed, with hyperscaler investment that J.P. Morgan estimates at 697 billion in 2026 alone on equipment that depreciates within a few years (J.P. Morgan, 2026), while the numerator grows more slowly and, in the model layer, is squeezed by the Chinese competition described in the first section. That combination is a fall in the rate of profit on the capital invested in artificial intelligence, and when the rate falls, accumulation slows, whether by coordinated decision of the large capitals or by the withdrawal of credit. To stop enlarging the denominator and to exploit what is already installed is, literally, to “digest what they have.” The apocalypse that OpenAI, Anthropic and Musk are warning about is the form in which that defense of the rate of profit presents itself to the public.

    Two objections deserve an answer. The first is that the brake might be sincere. It might be, and the thesis does not deny it: it claims that its economic function is the one described, whatever the conviction of the person announcing it. Other readers have offered kindred or harsher readings: Ben Thompson considers it “an unrealistic proposal that seems mostly geared to political control of AI” (Thompson, 2026), and Zvi Mowshowitz catalogs the readings of regulatory capture, stock market timing and computing limits, even though he dismisses them (Mowshowitz, 2026). The second objection is Chinese competition: a real brake would hand the lead to DeepSeek and Qwen, which, for the reasons given in the first section, have no reason to brake, and the coordination with China that the essay proposes is, under present conditions, unattainable. The answer is that, to be consistent with that competition, the brake has to be a paper brake, that is, external evaluators and alignment timelines without any measurable limit on the pace of training. A paper brake with maximum narrative effect is exactly what the thesis predicts, and it can be checked by reading the pact when it is signed.

    5. What to Watch in the Coming Weeks

    A thesis is worth what its risks of refutation are worth. Five observable facts in the coming weeks put it to the test.

    1. Anthropic’s public prospectus. If it describes the pacing commitment as immaterial to its revenue, the brake does not touch the till.
    2. Amazon’s conditional tranche in OpenAI. If the 35 billion is released or renegotiated without a listing or an artificial general intelligence, the brake costs no capital.
    3. The capital expenditure guidance of Alphabet, Microsoft, Meta and Amazon in their third-quarter results at the end of October. A coordinated cut would confirm the reading of investment restraint.
    4. The credit default swaps of CoreWeave and Oracle. If they keep widening while the hyperscalers rise, the devaluation begins on the leveraged periphery.
    5. The content of the pact among laboratories. If it sets evaluators rather than verifiable limits on pace, the brake is form without substance.

    6. Conclusion

    What happened this weekend is best understood as an agreement among those who buy chips to run a little slower, expressed in the language of safety: it improves their cash position, it worsens the sales of those who sell chips, and the market read it that way in a single day. The part of the argument that the evidence confirms most strongly is that the discourse of danger sustains the valuation rather than sinking it, because it asserts that the goal is near. The part that remains open is the part that ought to remain open: if artificial general intelligence arrives within the timeframes its promoters announce, this article will have been wrong; if the brake is renewed while the promise recedes, it will not have been a brake for humanity but a defense of the rate of profit, that is, a way of managing capital in retreat.

    · · ·

    References

    Alphabet Inc. (2026a, February 4). Alphabet announces fourth quarter and fiscal year 2025 results [Exhibit 99.1 to Form 8-K]. U.S. Securities and Exchange Commission. https://www.sec.gov/Archives/edgar/data/1652044/000165204426000012/googexhibit991q42025.htm

    Alphabet Inc. (2026b, July 22). Alphabet announces second quarter 2026 results [Exhibit 99.1 to Form 8-K]. U.S. Securities and Exchange Commission. https://www.sec.gov/Archives/edgar/data/1652044/000165204426000066/googexhibit991q22026.htm

    Amodei, D. (2026a, January). The adolescence of technology. https://darioamodei.com/essay/the-adolescence-of-technology

    Amodei, D. (2026b, September). We must pace the frontier. https://darioamodei.com/post/we-must-pace-the-frontier

    Anthropic. (2025, September 8). Anthropic is endorsing SB 53. https://www.anthropic.com/news/anthropic-is-endorsing-sb-53

    Artificial Analysis. (n.d.). LLM leaderboard: Intelligence Index. Retrieved September 14, 2026, from https://artificialanalysis.ai/leaderboards/models

    Azevedo, M. A. (2026, April 2). Sector snapshot: Venture funding to foundational AI startups in Q1 was double all of 2025. Crunchbase News. https://news.crunchbase.com/venture/foundational-ai-startup-funding-doubled-openai-anthropic-xai-q1-2026/

    Bank of England. (2025, October 8). Record of the Financial Policy Committee meeting on 2 October 2025. https://www.bankofengland.co.uk/financial-policy-committee-record/2025/october-2025

    Bellan, R. (2025, November 4). Anthropic projects 70B in revenue by 2028: Report. TechCrunch. https://techcrunch.com/2025/11/04/anthropic-expects-b2b-demand-to-boost-revenue-to-70b-in-2028-report/

    Bellan, R. (2026a, May 28). Anthropic raises 65 billion, nears 1T valuation ahead of IPO. TechCrunch. https://techcrunch.com/2026/05/28/anthropic-raises-65-billion-nears-1t-valuation-ahead-of-ipo/

    Bellan, R. (2026b, March 31). OpenAI, not yet public, raises 3B from retail investors in monster 122B fund raise. TechCrunch. https://techcrunch.com/2026/03/31/openai-not-yet-public-raises-3b-from-retail-investors-in-monster-122b-fund-raise/

    BenchLM. (2026, September 14). Artificial Analysis Intelligence Index leaderboard. https://benchlm.ai/benchmarks/artificialanalysis

    Brandom, R. (2026, February 27). OpenAI raises 110B in one of the largest private funding rounds in history. TechCrunch. https://techcrunch.com/2026/02/27/openai-raises-110b-in-one-of-the-largest-private-funding-rounds-in-history/

    Cabili, C. (2026, July 22). Alphabet Q2 2026 earnings: Revenue up 24%, Cloud surges 82%. Quartz (via Yahoo Finance). https://finance.yahoo.com/markets/stocks/articles/alphabet-q2-2026-earnings-revenue-203058727.html

    Caixin Global. (2026, July 22). DeepSeek reaches 52 billion valuation in round backed by Tencent, CATL. Caixin China Watch. https://caixinchinawatch.substack.com/p/deepseek-reaches-52-billion-valuation

    Cboe Global Markets. (2025, November 18). Cboe to launch trading of Cboe Magnificent 10 Index futures and options on December 8 [Press release]. https://ir.cboe.com/news/news-details/2025/Cboe-to-Launch-Trading-of-Cboe-Magnificent-10-Index-Futures-and-Options-on-December-8/default.aspx

    Center for AI Safety. (2023, May 30). Statement on AI risk. https://www.safe.ai/work/statement-on-ai-risk

    Cerullo, M. (2026, July 29). Workers at leading AI companies call for a slowdown in AI development. CBS News. https://www.cbsnews.com/news/slow-down-ai-development-sam-altman/

    ChinaTalk. (2024, November 27). Deepseek: The quiet giant leading China’s AI race [Translation of a July 2024 interview of Liang Wenfeng by An Yong for 36Kr]. https://www.chinatalk.media/p/deepseek-ceo-interview-with-chinas

    CME Group. (2026, June 30). CME Group to launch single stock futures on July 27 [Press release]. PR Newswire. https://www.prnewswire.com/news-releases/cme-group-to-launch-single-stock-futures-on-july-27-302814053.html

    Cook, L. D. (2026, May 27). The opportunities and risks AI presents for the economy and financial system [Speech]. Board of Governors of the Federal Reserve System. https://www.federalreserve.gov/newsevents/speech/cook20260527a.htm

    DeepSeek-AI. (2025, February). Day 6: One more thing, DeepSeek-V3/R1 inference system overview [Repository document]. GitHub. https://github.com/deepseek-ai/open-infra-index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_deepseekV3R1_inference_system_overview.md

    Duprey, R. (2026, January 31). Is the stalled Nvidia-OpenAI megadeal AI’s first domino to fall? 24/7 Wall St. (via Yahoo Finance). https://finance.yahoo.com/news/stalled-nvidia-openai-megadeal-ai-131959187.html

    Fernholz, T. (2026, July 28). Sam Altman is ready to decelerate. TechCrunch. https://techcrunch.com/2026/07/28/sam-altman-is-ready-to-decelerate/

    Fung, K. (2023, March 29). The massive name in AI noticeably absent from pause letter signed by Musk. Newsweek. https://www.newsweek.com/massive-name-ai-noticeably-absent-pause-letter-signed-musk-1791303

    Future of Life Institute. (2023, March 22). Pause giant AI experiments: An open letter. https://futureoflife.org/open-letter/pause-giant-ai-experiments/

    Goel, S. (2026, April 27). OpenAI just updated its principles. Here’s what changed since the original version, 8 years ago. Business Insider (via AOL). https://www.aol.com/articles/openai-just-updated-principles-heres-041316000.html

    Goldman, S. (2023, May 16). In Senate testimony, OpenAI CEO Sam Altman agrees with calls for an AI regulatory agency. VentureBeat. https://venturebeat.com/ai/in-senate-testimony-openai-ceo-sam-altman-agrees-with-calls-for-an-ai-regulatory-agency

    Greenberg, G. (2026, January 9). Advisors confront Magnificent 7 concentration risk in portfolios. InvestmentNews. https://www.investmentnews.com/equities/mag-7-for-tomorrow/264753

    Jadhav, A. (2025, November 4). Nvidia, OpenAI, and the trillion-dollar loop. The Register. https://www.theregister.com/special-features/2025/11/04/nvidia-openai-and-the-trillion-dollar-loop/353799

    J.P. Morgan. (2026, August 10). Financing AI infrastructure and U.S. data centers. https://www.jpmorgan.com/insights/banking/capital-markets/financing-ai-infrastructure-data-centers

    Juniewicz, I. (2026, June 5). The AI boom has doubled computing infrastructure’s share of US GDP. Epoch AI. https://epoch.ai/data-insights/ai-datacenter-share-gdp

    Korosec, K. (2026, June 1). Anthropic files to go public. TechCrunch. https://techcrunch.com/2026/06/01/anthropic-files-to-go-public/

    Lipschultz, B., & Metz, R. (2026, August 15). Anthropic revenue surges to over 11.5 billion in second quarter. Fortune. https://fortune.com/2026/08/15/anthropic-revenue-q2-11-5-billion-ipo-investors/

    Lo, K. (2026, January 6). Chinese AI unicorns beat Silicon Valley giants in race to go public. Rest of World. https://restofworld.org/2026/zhipu-ai-minimax-ipo/

    Luong, N. (2026, March 23). Two loops: How China’s open AI strategy reinforces its industrial dominance [Research working group paper]. U.S.-China Economic and Security Review Commission. https://www.uscc.gov/sites/default/files/2026-03/Two_Loops–How_Chinas_Open_AI_Strategy_Reinforces_Its_Industrial_Dominance.pdf

    Ma, J. (2026a, September 13). Anthropic CEO Dario Amodei and AI whistleblower Jacob Coxon agree: Something terrifying could happen in a matter of months. Fortune. https://fortune.com/2026/09/13/anthropic-dario-amodei-ai-whistleblower-jacob-coxon-openai-sam-altman-recursive-self-improvement/

    Ma, J. (2026b, September 12). Sam Altman confirms OpenAI won’t go public this year, saying an IPO now would come at an ‘ill-advised moment’ given AI safety concerns. Fortune. https://fortune.com/2026/09/12/sam-altman-openai-ipo-delay-ill-advised-moment-safety-concerns/

    Mann, T. (2026, August 3). China turns up the heat with open model blitz as US model makers panic. The Register. https://www.theregister.com/ai-and-ml/2026/08/03/china-turns-up-the-heat-with-open-model-blitz-as-us-model-makers-panic/5282526

    MarketScale Newsroom. (2026, August 17). AI data center debt is getting more expensive, and Meta’s 12.5 billion El Paso deal proves it. MarketScale. https://www.marketscale.com/industries/energy/ai-data-center-debt-is-getting-more-expensive-and-metas-125-billion-el-paso-deal-proves-it

    Metinko, C. (2024, January 9). Artificial buildup: AI startups were hot in 2023, but this year may be slightly different. Crunchbase News. https://news.crunchbase.com/ai/hot-startups-2023-openai-anthropic-forecast-2024/

    METR. (2026, August 26). Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/

    Microsoft. (2025, November 18). Microsoft, NVIDIA and Anthropic announce strategic partnerships. Microsoft Corporate Blogs. https://blogs.microsoft.com/blog/2025/11/18/microsoft-nvidia-and-anthropic-announce-strategic-partnerships/

    Moadel, D. (2026, July 29). Nebius drops 10%, CoreWeave sinks 9% as rising credit-swap costs hit the AI cloud trade. 24/7 Wall St. https://247wallst.com/investing/2026/07/29/nebius-drops-10-coreweave-sinks-9-as-rising-credit-swap-costs-hit-the-ai-cloud-trade/

    Mowshowitz, Z. (2026, September 14). We must pace the frontier. Don’t Worry About the Vase. https://thezvi.wordpress.com/2026/09/14/we-must-pace-the-frontier/

    Nusca, A. (2025, August 20). Just don’t call it an AI bubble. Fortune. https://www.fortune.com/2025/08/20/just-dont-call-it-an-ai-bubble

    O’Donnell, G., Hollerith, D., & Ferré, I. (2026, September 14). Stock market today: Dow, S&P 500, Nasdaq slip as chip stocks fall on AI warning. Yahoo Finance. https://finance.yahoo.com/markets/live/stock-market-today-monday-september-14-dow-sp-500-nasdaq-080559558.html

    OpenAI. (2018, April 9). OpenAI charter [Copy hosted by ETO AGORA, Emerging Technology Observatory]. https://agora.eto.tech/instrument/767

    Park, L. (2026, August 21). Anthropic’s Q2 revenue overtook OpenAI for the first time – and reached its first positive operating income. Forkast News (via Yahoo Finance). https://finance.yahoo.com/technology/ai/articles/anthropic-q2-revenue-overtook-openai-015716271.html

    Pichai, S. (2026, February 4). Q4 earnings call: Remarks from our CEO. The Keyword (Google). https://blog.google/company-news/inside-google/message-ceo/alphabet-earnings-q4-2025/

    Reuters. (2025, September 6). Report: OpenAI expects business to burn 115 billion through 2029. Yahoo Finance. https://finance.yahoo.com/news/openai-expects-business-burn-115-022035561.html

    Reuters. (2026, August 26). OpenAI agents hacked Hugging Face in 700-strong swarm, tried to cover tracks, investigations find. NBC News. https://www.nbcnews.com/tech/tech-news/openai-report-says-network-was-hacked-rogue-ai-agents-rcna594590

    Roberts, L. (2025, November 20). Oracle and CoreWeave credit default swap spreads widening: Omen or jitters? Investing.com. https://www.investing.com/analysis/oracle-and-coreweave-credit-default-swap-spreads-widening-omen-or-jitters-200670506

    Roytburg, E. (2026, September 14). Wall Street’s AI doomsday trade is here: Chipmakers sink while hyperscalers gain. Fortune. https://fortune.com/2026/09/14/ai-slowdown-stocks-nvidia-meta/

    Schepkov, V. (2025, November 5). Nvidia’s Huang warns China will win AI race amid energy costs, regulations. Investing.com (via Yahoo Finance). https://finance.yahoo.com/news/nvidia-huang-warns-china-win-223321776.html

    Shibu, S. (2025, August 19). OpenAI CEO Sam Altman thinks we’re in an AI bubble because investors are ‘overexcited’ about artificial intelligence. Entrepreneur. https://www.entrepreneur.com/business-news/openais-sam-altman-warns-overexcited-investors-ai-bubble/496081

    Shontell, A. (2026, September 12). Exclusive: Sam Altman addresses AI doomsday fears in new interview. Fortune. https://fortune.com/2026/09/12/sam-altman-interview-ai-doomsday-safety-models-control-ipo-2027/

    Stanciuc, A.-M. (2026, August 15). Anthropic’s quarterly revenue passed 11.5bn, up more than 14-fold. The Next Web. https://thenextweb.com/news/anthropic-q2-2026-revenue-11-5-billion-operating-income

    Subran, L., Dejean, G., Hirt, A., Utermöhl, K., & Bartosch, M. (2026, March 25). What to watch. Allianz Research. https://www.allianz.com/content/dam/onemarketing/azcom/Allianz_com/economic-research/publications/specials/en/2026/march/2026_03_25_AI.pdf

    The White House. (2025a, December 11). Ensuring a national policy framework for artificial intelligence [Executive Order 14365]. https://www.whitehouse.gov/presidential-actions/2025/12/eliminating-state-law-obstruction-of-national-artificial-intelligence-policy/

    The White House. (2025b, July). Winning the race: America’s AI action plan. https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf

    Thompson, B. (2026, September 14). Pacing the frontier, AI’s digital limits, AI commissars. Stratechery. https://stratechery.com/2026/pacing-the-frontier-ais-digital-limits-ai-commissars/

    Tolomia, C. (2026, June 16). OpenAI 2025 financials leaked: 38.5B loss ahead of IPO. Quartz (via Yahoo Finance). https://finance.yahoo.com/markets/stocks/articles/openai-2025-financials-leaked-38-121508294.html

    Vinsel, L. (2026, February 8). You’re doing it wrong: Notes on criticism and technology hype [Repost of a 2021 Medium essay]. Peoples & Things. https://peoples-things.ghost.io/youre-doing-it-wrong-notes-on-criticism-and-technology-hype/

    Wells, M. (2026). Will AI investments pay off? Econ Focus, First/Second Quarter 2026. Federal Reserve Bank of Richmond. https://www.richmondfed.org/publications/research/econ_focus/2026/q1-q2_feature2

    Yildirim, E. (2025, November 6). Nvidia CEO says China ‘will win’ the global AI race as the U.S. falls behind in energy. Gizmodo. https://gizmodo.com/nvidia-ceo-says-china-will-win-the-global-ai-race-as-the-u-s-falls-behind-in-energy-2000682396

    Zeoli, C. (2026, June). China’s open-weight takeover. Wing Venture Capital. https://www.wing.vc/content/chinas-open-weight-takeover

    Zitron, E. (2026a, May 21). Anthropic’s “profitability” swindle. Where’s Your Ed At. https://www.wheresyoured.at/anthropics-profitability-swindle/

    Zitron, E. (2026b, June 15). Exclusive: OpenAI losses increased nearly 8X in 2025, with spending hitting 34 billion. Where’s Your Ed At. https://www.wheresyoured.at/exclusive-openai-financials/

    September 14, 2026
  • Navier-Stokes, LLMs and the Ownership of Ideas

    Navier-Stokes, LLMs and the Ownership of Ideas

    Mathematics · Artificial intelligence · Open science

    Navier–Stokes, LLMs, And The Ownership of Ideas

    What OpenAI claims to have proved, why external forcing does count, and what the dispute reveals about the social future of mathematics.

    One of the most important and unusual mathematical stories in recent years began with a rumor. Two researchers, Tristan Buckmaster and Levent Alpöge, had spent months developing a line of work on singularities in fluid equations. OpenAI learned that something important was happening, deployed thousands of artificial-intelligence agents, and, a few days later, announced a proof of finite-time blowup for the three-dimensional Navier–Stokes equations.

    The story contains every ingredient needed to produce confusion: one of the seven Millennium Prize Problems, a 166-page proof, a Lean formalization, about 10,000 agents, 130 billion tokens, allegations concerning scientific credit and private data, and a conversation that Buckmaster interpreted as pressure against his career. Understanding it requires resisting two simplifications. The first is that a large amount of computation automatically turns a result into scientific understanding. The second is that the presence of an external force makes the proof irrelevant. Neither is correct.

    1. The fundamental question

    The Navier–Stokes equations describe the motion of viscous fluids such as water and air. They do not track molecules one by one. They treat the fluid as a continuous medium and relate its velocity, pressure, viscosity, and external forces:

    ∂u/∂t + (u · ∇)u = −∇p + νΔu + f,    ∇ · u = 0 u is velocity; p, pressure; ν, viscosity; and f, an external force. The second equality says that the fluid is incompressible.

    In three dimensions, we know how to construct global weak solutions, meaning solutions that satisfy the equations in a generalized sense. What we did not know was whether smooth initial data always produce smooth solutions for all time or whether, under some conditions, velocity can grow without bound in finite time. Such a loss of regularity is called a singularity or blowup.

    Viscosity tends to smooth differences in velocity. The nonlinearity represented by (u · ∇)u transports and reorganizes motion. The problem arises from the competition between these effects: is viscous dissipation always enough to prevent infinite concentration, or can three-dimensional dynamics overcome it?

    A physical clarification

    Mathematically infinite velocity does not mean that a real fluid literally reaches that speed. It means that the continuum model loses regularity and ceases to describe the process in the classical way. The question is mathematical, although it is physically motivated.

    2. What Clay actually requires

    The first indispensable correction belongs here. The official statement written by Charles Fefferman offers four routes. Alternatives A and B ask for global existence and smoothness when the external force is zero, respectively in the whole space and on the three-dimensional torus. Alternatives C and D permit a disproof of global regularity using smooth initial data and an external force that is also smooth, on those same two domains.

    AlternativeDomainExternal forceWhat must be shown
    Aℝ³f = 0Global existence and smoothness for every admissible initial datum.
    B𝕋³f = 0The same global regularity in the periodic case.
    Cℝ³smooth fExistence of data and a force for which there is no global classical solution of bounded energy.
    D𝕋³smooth fThe corresponding breakdown in the periodic case.

    Calling smooth forcing a “loophole” foreign to the prize is therefore technically incorrect. It is written into alternatives C and D. If OpenAI’s proof is correct and satisfies the stated hypotheses exactly, it would resolve the official problem through those routes.

    That does not make all results equivalent. Producing a singularity with a specially constructed smooth force is a more limited claim, physically and conceptually, than proving a singularity without external forcing. It resolves the institutional disjunction formulated by Clay, but it does not answer the stronger question of whether unforced dynamics can break down by themselves. Both statements can be true at once, without verbal tricks.

    External forcing does not formally invalidate the result; it precisely limits what the result teaches us about fluids.

    3. What OpenAI claims to have proved

    OpenAI’s paper constructs, for every positive viscosity, velocity and pressure fields that remain smooth up to a critical time, together with a smooth, compactly supported force f(x,t). The fluid starts from rest. As time approaches t = 1, the maximum velocity grows without bound, while total kinetic energy remains uniformly bounded.

    The intuitive picture is a vortex that spins, stretches, and concentrates around an increasingly small region. Velocity can rise indefinitely without total energy blowing up because the volume occupied by the extreme motion contracts quickly enough. An imperfect but useful analogy is to concentrate ever greater intensity inside ever less space.

    It is not enough to draw a singular flow and then define an infinite force to sustain it. That would be trivial and inadmissible. The difficulty is to make acceleration, pressure, nonlinear transport, and viscosity develop large terms that cancel with extraordinary precision, leaving a perfectly smooth external force as the residual. The paper introduces oscillatory pulses whose internal momentum transport corrects the singular parts of that residual.

    Exact scope of the announced theorem
    • Yes: finite-time velocity blowup for three-dimensional Navier–Stokes with positive viscosity and a smooth external force.
    • Yes: uniformly bounded total kinetic energy throughout the evolution.
    • Yes, conditional on the proof being correct: alternatives C and D of the official formulation.
    • No: a proof of blowup for unforced Navier–Stokes.
    • Not yet: stable, independent acceptance by the mathematical community.

    4. How the agents worked

    According to OpenAI, on September 1 it assigned groups of agents to the different alternatives of several Millennium Prize Problems. The agents could consult a cached copy of the internet, run code, divide into groups, and communicate results. Nearly one hundred agents reportedly first produced a solution for unforced Euler. From that result, the company concentrated its resources on Navier–Stokes.

    The decisive group eventually involved on the order of 10,000 concurrent agents. OpenAI reports that the result appeared after about 88 hours; Lean formalization and verification took an additional 17 hours using GPT-6 Astra. The Navier–Stokes effort reportedly involved 2.7 million messages and approximately 130 billion generated tokens.

    These numbers neither prove understanding nor disprove it. They describe a new material organization of intellectual labor: massively parallel search, route selection, circulation of partial results, automated consolidation, and formal verification. Reducing this to “mere brute force” hides the coordination architecture; presenting it as an autonomous mathematical consciousness hides the human, bibliographical, and computational scaffolding that made the search possible.

    5. The mathematical genealogy

    The route did not emerge in a vacuum. Buckmaster attributes the basic idea of the program to Diego Córdoba and Luis Martínez-Zoroa, who had spent years studying forced blowups and mechanisms of amplification across scales. Their work had shown how increasingly concentrated layers of vorticity could amplify while the regularity of the force remained controlled.

    Buckmaster and Alpöge took that program as their starting point. With extensive help from Claude and Codex, they say they pushed constructions with rougher forcing to smooth forcing for incompressible porous media, Boussinesq, and three-dimensional incompressible Euler. According to their statement, they obtained the Boussinesq and Euler results on August 15 and verified the first model-generated proof in Lean on August 22.

    OpenAI’s announced contribution would be different and later: a construction for viscous Navier–Stokes. Even if it was obtained independently, its historical meaning cannot be narrated as if 10,000 agents had discovered both the problem and the strategy from a blank page. In mathematics, the final proof is only one part of discovery. Recognizing which simplification to attack, which literature matters, and which mechanism may survive new constraints is genuine intellectual work.

    6. Timeline of the dispute

    Date in 2026Reported eventSource of the claim
    August 15Buckmaster and Alpöge obtain smooth-forcing blowup for Boussinesq and Euler.Buckmaster’s statement.
    August 22Lean verification of their initial proof is completed.Buckmaster’s statement.
    September 1OpenAI hears rumors and launches agents on the Millennium Prize Problems.OpenAI’s account.
    September 3Buckmaster privately writes to a mathematician at OpenAI to clarify the rumor.Email reproduced by Buckmaster.
    September 5OpenAI’s agents reach the Navier–Stokes construction, about 88 hours after starting.OpenAI’s account.
    September 6OpenAI’s formal verification is completed and two calls with Buckmaster take place.OpenAI and Buckmaster’s statement.
    September 8OpenAI publishes the article, proof, and formalization; Buckmaster makes his statement public.The respective publications.

    The temporal proximity fuels the controversy. Buckmaster says that almost nobody was working on this route and that the sudden focus on the forced problem was a warning sign to him. OpenAI responds that it distributed agents among all four official alternatives, first obtained an unforced Euler result, and only then identified Navier–Stokes as the most promising direction.

    7. What is established and what is disputed

    The public discussion mixes verifiable documents, incompatible accounts, and inferences. They should be separated.

    Evidence map
    • Publicly documented: an analytical manuscript, a Lean formalization, Buckmaster’s statement, and OpenAI’s recognition of Buckmaster and Alpöge’s priority on forced Euler all exist.
    • Claimed by OpenAI: no researcher or agent accessed their specific work before publication; the agents reached different results, and the company does not intend to claim the monetary prize.
    • Acknowledged by OpenAI as a remote possibility: it cannot rule out that de-identified data derived from their product usage helped improve its models.
    • Claimed by Buckmaster: he received no clear answer about training; he was offered an arrangement in which he would present OpenAI’s result alone and Alpöge would be excluded because he works at Anthropic; he also received statements he interpreted as professional threats.
    • Disputed: Sébastien Bubeck rejects Buckmaster’s characterization of the interaction, according to the cited public summary.
    • Not publicly established: that private drafts were used to produce OpenAI’s proof or that deliberate appropriation occurred.

    Buckmaster himself draws the boundary: he has not seen how the model worked, does not know whether his data were used, and does not present that accusation as an established fact. His stated purpose is to record what he was told, when he was told it, and what arrangements were proposed. This caution does not erase the questions; it prevents them from being answered through condemnation without evidence.

    OpenAI’s sentence about de-identified data also does not demonstrate a causal transmission from those drafts to the proof. It does reveal a governance problem: for a researcher placing unpublished work in a tool, an inability to exclude aggregate influence may be scientifically unacceptable even if the handling complied with the service’s contractual terms.

    8. What Lean verifies and what it does not

    Lean is a proof assistant. A formalized proof is decomposed into steps that a small kernel checks against precise logical rules. This sharply reduces the chance that an algebraic leap, forgotten hypothesis, or invalid inference survives. In a manuscript over one hundred pages long, that guarantee is extraordinarily valuable.

    Lean does not, by itself, answer four separate questions:

    1. Correspondence: whether the formal theorem exactly represents the informal mathematical claim being announced.
    2. Interpretation: what physical meaning the construction has and how much it depends on a designed force.
    3. Novelty: which ideas are new and which come from prior work.
    4. Credit: who formulated the route, supplied the decisive mechanisms, directed the search, and should be named as an author.

    A logical certificate does not replace a human exposition. The community still needs to read the manuscript, check that the formalization matches the intended theorem, reconstruct the conceptual mechanism, and place it within the literature. Formal verification strengthens the evidence; it does not exhaust scientific practice.

    9. The Deep Blue–Kasparov moment

    Buckmaster calls the episode a “Deep Blue–Kasparov moment.” The comparison does not mean that mathematics is simply another game with closed rules. It points to a change in the division of labor: a machine can explore, combine, and verify on a timescale inaccessible to a person, but its productivity depends on a search space structured by historically accumulated problems, theories, formal libraries, and strategies.

    In a major problem, the scarce resource is not always writing the last line of a proof. It may be identifying the promising route after years of failed attempts. If revealing that route allows thousands of agents to be mobilized immediately, then selecting the correct intermediate question acquires enormous economic and scientific value.

    When proving becomes massively parallelizable, knowing what should be proved may become the scarcest part of discovery.

    This could make open mathematics more guarded. A seminar, an informal draft, or a query submitted to a platform would cease to be only an act of cooperation and could become an exploitable signal for institutions with incomparable computing capacity. The danger is not that artificial intelligence participates in science, but that openness becomes unilateral: open knowledge from below and closed appropriative capacity from above.

    10. A materialist reading

    The controversy cannot be understood solely as a moral conflict between individuals. It expresses a transformation in the productive forces of science. Models, data centers, multi-agent systems, and proof assistants turn enormous quantities of accumulated mathematical labor into an infrastructure capable of producing new combinations at great speed.

    At the same time, this productive force is concentrated under private ownership. Researchers contribute questions, drafts, corrections, usage, and tacit knowledge; the platform aggregates those interactions, improves closed models, and can return results at a scale no academic group can match. Even if no concrete appropriation occurred in this case, the structural asymmetry is real.

    The decisive question is therefore not whether the machine “really thought,” a metaphysical dispute of little use. We should ask who controlled the means of mathematical production, who chose the direction of the search, which prior labor it depended on, who could audit the process, and how authorship, prestige, and benefits were distributed.

    Computation is not an epistemological sin. It can be a genuine force of discovery. But a new productive force organized through opaque relations of ownership and credit can simultaneously expand human capacity and deepen the subordination of scientific labor.

    11. The rules we need

    The episode suggests a minimum program for AI-assisted research:

    1. Verifiable confidential mode: unpublished drafts should be excluded from training and any later retrieval, with auditable guarantees rather than contractual language alone.
    2. Provenance: every major result should retain a record of models, versions, consulted sources, decisive prompts, human intervention, and circulation among agents.
    3. Compute disclosure: the number of agents, tokens, time, tools, and selection criteria should be treated as part of the method.
    4. Conceptual genealogy: attribution should recognize not only the final proof but also the prior program that made the route visible.
    5. Authorship independent of corporate rivalry: a researcher’s affiliation cannot justify excluding a scientifically relevant contribution.
    6. Open review: the manuscript, formalization, and correspondence between them should remain available for independent examination.
    7. Public capacity: universities and scientific consortia need shared infrastructure so frontier mathematics does not depend exclusively on private laboratories.

    These rules would not settle every priority dispute, but they would make verifiable questions that currently depend on conflicting institutional statements.

    12. Conclusion

    Did OpenAI solve Navier–Stokes? The rigorous answer is conditional. The announced theorem targets precisely alternatives C and D of Clay’s official statement. The use of a smooth external force is not a trick, although it leaves the stronger unforced question untouched. If the analytical proof and its formalization withstand expert review, the mathematical achievement would be historic.

    But the result did not arise outside history. It depends on decades of analysis, on the route opened by Córdoba and Martínez-Zoroa, on the work of Buckmaster and Alpöge, and on extraordinary computational infrastructure. Nor have the questions about data, priority, and institutional pressure been resolved. The most serious allegations should be investigated, neither repeated as facts nor dissolved through public relations.

    The deepest lesson is that artificial intelligence does not abolish the social relations of science. It makes them more visible. The more powerful automated proving becomes, the more important it will be to protect openness, record the provenance of ideas, and distribute credit fairly. The future of mathematics will depend not only on what machines can prove, but on the institutions through which we decide what it means to discover.

    References

    Specified sources

    1. Julien Cadot, public summary of the controversy.https://x.com/juliencdt/status/2097582940863471900
    2. Mehdi, technical, epistemological, and political criticism of the announcement.https://x.com/bettercallmedhi/status/2097472336257863722
    3. Susan Zhang, reproduction of the manuscript and OpenAI’s statement on concurrent work.https://x.com/suchenzang/status/2097376584814432392
    4. Tristan Buckmaster, statement on the results and conversations with OpenAI.https://cims.nyu.edu/~tristanb/statement.pdf
    5. Jasper, mathematical analysis and summary of the dispute.https://x.com/zjasper/status/2097547276755619988

    Primary technical documents

    1. OpenAI, “On the Navier–Stokes Millennium Prize Problem.”https://openai.com/index/navier-stokes-solution/
    2. OpenAI, “Finite Time Blowup for Navier–Stokes.”Analytical manuscript
    3. Charles L. Fefferman, official formulation of the Millennium Prize Problem.Existence and Smoothness of the Navier–Stokes Equation
    4. OpenAI, Lean formalization.https://github.com/openai/NavierStokesAndEuler
    ↑ Back to top
  • GENERALITIES ON THE STATISTICAL THEORY OF SAMPLING SURVEYS

    GENERALITIES ON THE STATISTICAL THEORY OF SAMPLING SURVEYS

    SURVEY THEORY · SAMPLING · INFERENCE

    GENERAL REMARKS ON THE STATISTICAL THEORY OF SAMPLE SURVEYS

    Conceptual and technical notes on sampling design, frames, inclusion probabilities, design-based inference, nonresponse, total survey error, and estimation.

    The statistical theory of sample surveys is often learned as a collection of procedures: how to select a sample, how to estimate a total or a mean, how to calculate a standard error, and how to determine a sample size. Yet behind every procedure lies a prior conceptual question. What, exactly, is the population to which one intends to infer? What entity is selected at each stage? What makes a selection probabilistic? Where does the randomness that allows us to speak of precision come from? What part of survey error can be quantified by the sampling design, and what part arises from coverage, nonresponse, measurement, or processing?

    The notes that follow are organized around problems that appear in the classical sampling literature—in particular Cochran (1977), Kish (1965), and Lohr (2021)—and develop them using tools from contemporary survey theory. The purpose is not to replace the reading of those authors, but to make explicit the conditions under which their statements acquire meaning and the operational consequences that follow from them. When more recent terminology is introduced, the discussion draws especially on Särndal, Swensson, and Wretman (1992), Groves et al. (2009), Biemer and Lyberg (2003), AAPOR’s definitions, and the guidelines of the United Nations, Statistics Canada, and the U.S. Census Bureau.

    · · ·

    1. In a survey, are costs high and administration more complex?

    The answer necessarily depends on the basis of comparison. Relative to a census, a sample survey is usually cheaper and administratively less demanding because it observes only part of the population. Indeed, one of the historical reasons for the development of probability sampling was to obtain sufficiently precise information at a cost far below that of a complete enumeration.

    The comparison changes, however, if the reference point is a smaller-scale qualitative procedure, such as a focus group, an in-depth interview, or intensive observation. A national probability survey may require a frame, sample selection, instrument programming, training, supervision, callbacks, weighting, quality control, variance estimation, and documentation. A focus group may be far less expensive, but it does not perform the same inferential function. The choice should not be made by asking which technique is “cheaper” in the abstract, but which technique can answer the research question with the kind of evidence required.

    A focus group, for example, can be extraordinarily useful for understanding vocabularies, categories of meaning, mechanisms, and problems in questionnaire comprehension. It is not, however, a direct substitute for a probability survey when the objective is to estimate parameters of a defined population. Cost must therefore be judged in relation to the objective, the required precision, and the type of inference one intends to make.

    2. What is a double-barreled question?

    It is a question that asks for a single response concerning two or more propositions that could receive different answers. The problem is not that the respondent has “two response options”; almost every closed question does. The problem is that the wording fuses two objects of judgment and forces the respondent to compress them into a single answer.

    For example:

    Are you satisfied with your salary and with your working conditions?

    A person may be satisfied with the salary and dissatisfied with the conditions, or the reverse. If the scale offers only one response, the resulting datum has no univocal interpretation. The solution is to separate the components whenever each constitutes a substantively distinguishable dimension.

    This problem points to something more general: a questionnaire is not a neutral container into which already constituted variables are simply “placed.” The concrete wording helps determine the cognitive operation performed by the respondent and, consequently, the datum that is ultimately recorded. Instrument design is therefore part of measurement and not a purely editorial task.

    3. When is a questionnaire too long?

    There is no universal number of questions or minutes beyond which a questionnaire becomes too long. Burden depends on mode of administration, the cognitive complexity of the questions, the sensitivity of the topics, the population being interviewed, the need to consult records, the number of skip patterns, and the repetition of scales, among other factors.

    A face-to-face interview may tolerate a different duration from a telephone interview or a self-administered online questionnaire. Likewise, twenty questions that require recalling amounts, dates, or episodes may impose a greater burden than fifty very brief and homogeneous items.

    Pilot testing is important, but it should be distinguished from cognitive pretesting. A pilot reproduces field procedures on a smaller scale and makes it possible to study duration, contact rates, skip patterns, logistics, and implementation problems. A cognitive interview, by contrast, seeks to understand how the respondent interprets the question, retrieves the information, and selects an answer. The U.S. Census Bureau also includes usability testing, behavior coding, debriefing, and split-panel designs among questionnaire evaluation tools.

    The methodological rule, therefore, is not “make it as short as possible,” but retain everything needed to measure the defined concepts and remove what adds burden without adding substantive information. Brevity and quality are not synonyms.

    4. What does it mean for a frame to contain duplicates?

    A sampling frame is the operational device through which sampling units can be identified or reached in order to carry out selection. Sometimes it takes the form of a list; at other times it consists of geographic areas, addresses, administrative registers, potential telephone numbers, or procedures that make it possible to generate selectable units.

    Duplication occurs when the same population unit is represented more than once in the selection mechanism. If records are selected with equal probability and one person appears twice, that person may acquire a higher inclusion probability than another who appears only once. But the problem should not be described by saying that any inequality in probabilities makes the sample “nonrandom.” There are deliberately unequal-probability probability designs.

    What matters is to distinguish between inequality that is designed and known and multiplicity that is accidental or unknown. A frame may suffer, among other problems, from:

    • undercoverage: units of the target population that cannot be reached through the frame;
    • overcoverage: records that do not belong to the target population;
    • duplicates or multiplicity;
    • outdated records;
    • units whose status or location has changed;
    • many-to-one or one-to-many relationships between records and target units.

    The statistical consequence depends on the specific selection mechanism and on whether multiplicity can be measured and corrected in the inclusion probabilities or in weighting.

    5. What does data editing mean, and what should be done with nonresponse?

    In survey work, “editing” should not be understood as deleting observations that look strange. Editing consists of applying explicit validity, consistency, and plausibility rules to identify records that require review. An improbable answer is not automatically false, and an inconsistent answer should not disappear without a trace.

    A reasonable workflow always preserves the raw information and separates the stages:

    1. captured datum;
    2. validation of range, type, and structure;
    3. verification of skips and logical consistency;
    4. possible recontact, when feasible and planned;
    5. documented editing rules;
    6. marking of missing values;
    7. possible imputation;
    8. production of the analytical dataset without destroying the original dataset.

    Suppose that a section of the questionnaire applies only to households that reported a certain economic activity. If answers appear in that section for households that, according to the filter question, do not perform that activity, several possibilities exist: the filter was captured incorrectly, the later response was captured incorrectly, the interviewer failed to follow the skip, the definition was misunderstood, or there is a real situation that the design did not anticipate. The datum must be investigated; it is not enough to declare it “obviously wrong.”

    It is also necessary to distinguish states that are often collapsed under a single missing-value label: Not applicable, Don’t know, No answer, explicit refusal, failure to contact, structural skip, capture error, and value lost during processing. “Not applicable” is not simply another form of DK/NR: it means that the question does not belong to the logical universe of that unit.

    Unit nonresponse and item nonresponse

    Unit nonresponse occurs when no substantive information is obtained from a selected unit; item nonresponse occurs when an interview exists but particular answers are missing. The former is often addressed through weighting adjustments and the latter may require imputation or other methods, although the specific techniques depend on the missingness mechanism and the analytical objective (Kalton and Kasprzyk, 1986).

    A high nonresponse rate does not mechanically imply that a variable is unusable, just as a high response rate does not demonstrate the absence of bias. Nonresponse bias depends on the relationship between response propensity and the variables to be estimated.

    AAPOR distinguishes, among other measures, response rate, cooperation rate, contact rate, and refusal rate. They are not synonyms. The response rate relates completed interviews to the set of eligible cases—with variants concerning the treatment of unknown eligibility and partial interviews; cooperation conditions on contacted units; the contact rate studies the success of contact; and the refusal rate quantifies a specific form of nonresponse.

    Finally, interviewer probing must be standardized. The objective is not to press until “no survey is ever left incomplete,” but to facilitate a valid answer without inducing it and while respecting the participant’s consent.

    6. Efficiency, repeated sampling, and confidence intervals

    Cochran frames sampling theory as a joint problem of precision and cost: to design selection and estimation procedures that produce the required accuracy at the lowest possible cost. The underlying idea remains fundamental. There is no “best” design outside an objective; there is a design that is more or less efficient for particular estimands, constraints, and costs.

    The precision of a procedure is studied through its distribution under hypothetical repetitions of the sampling mechanism. From a design-based perspective, the values of the finite population may be regarded as fixed, while randomness arises from the set of samples that the design could have selected.

    Here it is important to state the interpretation of a confidence interval precisely. If a procedure constructs intervals with a 95% confidence level, that does not mean that 95 out of 100 samples will produce “values similar” to the value observed in our sample. It means that, under the assumptions of the procedure and under hypothetical repetitions of the design, approximately 95% of the intervals constructed in that way would contain the fixed population parameter.

    Once a particular frequentist interval has been observed, the parameter does not thereby become random. The randomness belongs to the procedure that would have produced different intervals under other possible samples.

    Sample size for a proportion

    In an elementary formulation under simple random sampling, if an approximate margin of error \(e\) is desired for a proportion at a level associated with \(z_{1-\alpha/2}\), one may begin with

    \[ n_0=\frac{z_{1-\alpha/2}^2\,p(1-p)}{e^2}. \]

    When there is no useful prior information about \(p\) and a conservative case is sought, \(p=0.5\) is used because the function

    \[ p(1-p) \]

    attains its maximum at \(p=0.5\). That is the mathematical reason. It is not a consequence of the central limit theorem.

    If the population is finite and the sampling fraction is no longer negligible, sample size can be adjusted, within the same simple framework, by

    \[ n=\frac{Nn_0}{N+n_0-1}. \]

    In complex designs, the calculation must incorporate stratification, clustering, unequal probabilities, domain objectives, anticipated nonresponse, and, where appropriate, an expected design effect. Sample size should not be calculated in isolation from the design.

    7. Are there “common sample sizes”?

    Recurring sizes—400, 800, 1,000 interviews, for example—do appear in professional practice because certain combinations of precision, budget, and mode recur. But those numbers are not statistical constants.

    For a proportion near 0.5, a simple random sample of about 1,067 units yields, under a normal approximation and without finite population correction, a sampling margin close to three percentage points at 95%. This accounts for part of the recurrence of sample sizes near one thousand in polls. The calculation, however, does not by itself demonstrate that a survey is nationally representative or that its total error is ±3 percentage points.

    If there is clustering, highly variable weights, small domains, unequal response rates, or incomplete coverage, the same \(n\) may yield very different precision and quality. The proper question is not “what is the traditional size?” but “what precision do I need for which estimand, under what design, and at what cost?”

    8. Two differences between classical theory and survey theory

    8.1. Distributions, models, and design

    A classical observation by Cochran contrasts part of model-based statistical inference with survey theory for finite populations. In many classical problems, observations are assumed to arise from a distribution with a specified mathematical form and its parameters are estimated. In survey sampling, by contrast, a central tradition develops inference from the selection mechanism without requiring a parametric distribution for the population values.

    The expression “few measurements on each unit” does not refer to a sample with fewer than 30 or 50 cases. It refers to the number and type of variables measured on each unit. A large survey may collect many variables with very different distributions; there is no reason to impose the same parametric family on all of them merely because the dataset is large.

    Computing power does not eliminate this conceptual difference. Today at least three perspectives coexist:

    • design-based inference: randomness is attributed to the selection mechanism;
    • model-assisted inference: models are used to improve estimators, calibrate, or exploit auxiliary information while retaining properties justified by the design;
    • model-based inference: inference rests more directly on a probability model for the observed values or mechanisms.

    Särndal, Swensson, and Wretman (1992) systematically develop the model-assisted perspective. The distinction is not a relic from before modern computers; it answers different questions about the source of randomness and the assumptions that support inference.

    8.2. Finite populations and the finite population correction

    A population is finite if it contains \(N<\infty\) units. This does not depend on whether the researcher knows \(N\) exactly. A finite population can exist even when its size is unknown or uncertain.

    Under simple random sampling without replacement, for the population mean \(\bar Y\), if

    \[ S^2=\frac{1}{N-1}\sum_{i=1}^{N}(Y_i-\bar Y)^2, \]

    then

    \[ \operatorname{Var}(\bar y) =\left(1-\frac{n}{N}\right)\frac{S^2}{n}. \]

    The factor

    \[ 1-f,\qquad f=\frac{n}{N}, \]

    is the finite population correction in the variance. When \(f\) is very small, \(1-f\approx1\) and its effect may be negligible. When the sample covers a large proportion of the population, ignoring it can substantially overestimate sampling variance.

    The reason is materially simple: if an increasing proportion of a finite population is observed without replacement, progressively less uncertainty remains about the part that was not observed. When \(n=N\), a census of that population has been conducted and the variance due to the sampling mechanism is zero, although measurement, coverage, processing, and other errors may still exist.

    9. What does it mean to be sensitive to respondents’ expectations?

    The interviewer must maintain a combination that is not always easy: rapport without leading. The interviewer should establish a relationship that facilitates communication, explain the purpose of the interview, listen, and use permitted probes; but should not guide the respondent toward an answer that appears socially desirable or compatible with the researcher’s expectations.

    Empathy does not mean abandoning standardization. An interview can be humane, respectful, and sensitive to context without arbitrarily altering the meaning of the questions. In standardized instruments, an important part of training consists precisely in learning which clarifications are permitted and how to use neutral probes.

    Sensitivity also has an ethical dimension. The interviewer must recognize discomfort, fatigue, privacy, risks, and the right not to answer. Data quality cannot be pursued through coercion.

    The interaction between interviewer and respondent can affect measurement. Tone of voice, the way response options are presented, verbal or gestural reactions, the respondent’s perception of the interviewer’s status, and topic sensitivity can alter the observed response.

    Several mechanisms should be distinguished:

    • social desirability: a tendency to provide an answer perceived as socially acceptable;
    • acquiescence: a tendency to agree with formulations regardless of content;
    • interviewer effect: systematic variation associated with interviewer characteristics or behavior;
    • mode effect: differences associated with face-to-face, telephone, web, paper, or other modes;
    • context and order effect: responses conditioned by preceding questions or by the structure of the instrument.

    Informed consent, however, is not a bias. It is an ethical condition of research involving human participants. A different question is whether the decision to participate may generate self-selection or differential nonresponse. That possible source of bias should be studied as part of the participation process, not presented as a reason to weaken consent.

    11. Survey design and absolute sample size

    Lohr insists on an idea that should be formulated carefully: an enormous sample does not rescue a defective design. The relevant quantity in her discussion is the absolute size of the sample, not the absolute size of the population.

    For large populations, the precision of an estimator under simple random sampling depends primarily on \(n\); \(N\) appears in the finite population correction and has little effect when \(n/N\) is small. Thus, a sample of one thousand people may have a similar sampling error whether the population contains one million or one hundred million people, provided that the selection mechanism and other assumptions are comparable.

    But this statement does not mean that any sample of one thousand units is adequate. A convenience sample of one hundred thousand people may have more selection bias than a probability sample of one thousand. Increasing \(n\) reduces random variability under the assumed mechanism; it does not automatically erase systematic errors of coverage, selection, measurement, or nonresponse.

    Hence a rule that should remain present whenever polls are read:

    \[ \text{large sample}\;\nRightarrow\;\text{representativeness}. \]

    And likewise:

    \[ \text{small sampling margin}\;\nRightarrow\;\text{small total error}. \]

    12. Content, units, extent, and time when defining a population

    A study population must be defined precisely enough to decide who or what belongs to it. It is useful to specify its substantive content, the units that compose it, its spatial extent, and the time or reference period.

    Content is not simply spatiotemporal location; that belongs mainly to extent and time. Content determines what kind of entity and what substantive condition make a unit part of the population. For example: “persons aged 18 or older who are usual residents of private households in Costa Rica during the reference period.” There we find an entity, an age condition, a residence condition, a household delimitation, and a territorial-temporal extent.

    The different units

    The word “unit” has a general meaning—something that may be considered an entity for a given analysis—but in surveys it acquires different technical functions. They should not be confused.

    Element or target unit. The population entity about which information is to be produced: a person, a household, a firm, a plot of land, a school.

    Unit of analysis or study unit. The entity to which the analyzed variables are attributed. It may coincide with the target element, though not always.

    Observation unit. The entity on which observation or measurement is directly performed.

    Reporting unit. The person or entity that provides the information. In a household survey, one adult may report on other members; the reporting unit and the analytical unit are then not identical.

    Sampling unit. A selectable unit at a particular stage of the design. At a first stage it may be a geographic segment; at a second, a dwelling; at a third, a person.

    Primary sampling unit (PSU). The unit selected at the first stage of a multistage design. It is not defined by being “a place” or by necessarily containing only one final unit.

    Listing unit. A unit identified during a listing operation in order to construct or update the frame for a later stage. Its function is operational.

    The same entity may perform several functions. If a complete list of students is available and students are directly selected and asked about themselves, a student may simultaneously be the element, unit of analysis, observation unit, reporting unit, and sampling unit. But this empirical coincidence does not make the concepts synonymous.

    13. Domain, stratum, class, and subclass

    An estimation domain is a subpopulation for which separate estimates are desired. It may be defined by region, sex, age, economic sector, or another relevant condition.

    A stratum, by contrast, is a partition used by the sampling design before the sample is selected. Strata are generally intended to control composition, improve efficiency, or ensure sufficient representation of relevant groups.

    A domain may coincide with a stratum, but it need not. It may cut across several strata or be defined using information obtained after selection. For example, the population may be stratified by province and, later, a domain may be estimated consisting of persons employed in a particular economic industry across all provinces.

    The notion of a class as an interval in a frequency distribution belongs to a different problem. It is not necessary for defining an estimation domain. Likewise, “subclass” should not be used as a general synonym for domain unless a source explicitly defines that terminology for a particular scheme.

    14. What does \(Y_i\) represent?

    If the finite population is

    \[ U=\{1,2,\ldots,N\}, \]

    and \(Y\) is a study variable, then \(Y_i\) represents the value taken by that variable for unit \(i\). It does not represent “the characteristics” of the unit in the plural, but one particular realization of a defined variable.

    The population total is

    \[ T_Y=\sum_{i=1}^{N}Y_i, \]

    and the population mean is

    \[ \bar Y=\frac{1}{N}\sum_{i=1}^{N}Y_i. \]

    In design-based inference, a particularly important idea is that the vector

    \[ (Y_1,\ldots,Y_N) \]

    may be treated as fixed. The distribution used to study an estimator arises from the sampling design: which samples \(s\) could have been selected and with what probabilities.

    This prevents a common confusion. The use of probability does not require us to think that every material characteristic of the population was “generated” by a known parametric distribution. Probability may reside in the act of selection.

    15. What does MESIP selection mean?

    In terminology used in various Latin American texts, MESIP commonly denotes selection with equal probability of the elements. It should be understood as a particular class of probability design, not as the definition of all probability sampling.

    If \(\pi_i\) is the inclusion probability of unit \(i\),

    \[ \pi_i=P(i\in s), \]

    an equal-probability design satisfies, for comparable units within the design,

    \[ \pi_i=\pi_j. \]

    But a general probability design may have

    \[ \pi_i\neq\pi_j. \]

    The inequality may be deliberate, for example when selection is made with probability proportional to a measure of size.

    16. What does “unrestricted random sampling” mean?

    The expression is used in part of the Spanish-language literature to refer to simple random sampling without replacement: from a population of size \(N\), a fixed-size sample \(n\) is selected so that every possible subset of \(n\) units has the same probability of selection.

    If \(S\) is the set of all possible samples of size \(n\), then for every \(s\in S\),

    \[ P(S=s)=\binom{N}{n}^{-1}. \]

    It follows that every unit has inclusion probability

    \[ \pi_i=\frac{n}{N}. \]

    It is a fundamental design precisely because it permits transparent derivations and serves as a benchmark. But it is not “random sampling as such” in an exhaustive sense: probability sampling includes stratified, systematic, cluster, multistage, and unequal-probability designs.

    17. Mean squared error, variance, and bias

    Let \(\hat\theta\) be an estimator of the parameter \(\theta\). Its mean squared error is

    \[ \operatorname{MSE}(\hat\theta) =E\left[(\hat\theta-\theta)^2\right]. \]

    It can be decomposed as

    \[ \operatorname{MSE}(\hat\theta) =\operatorname{Var}(\hat\theta) +\operatorname{Bias}(\hat\theta)^2, \]

    where

    \[ \operatorname{Bias}(\hat\theta) =E(\hat\theta)-\theta. \]

    The identity can be seen by writing

    \[ \hat\theta-\theta =\bigl(\hat\theta-E(\hat\theta)\bigr) +\bigl(E(\hat\theta)-\theta\bigr), \]

    squaring, and taking expectations. The cross term vanishes because

    \[ E\left[\hat\theta-E(\hat\theta)\right]=0. \]

    Bias is not the difference between one observed realization of \(\hat\theta\) and its expectation. That difference is a random deviation of the realization from the center of its distribution. Bias is a property of the estimation procedure: the difference between the estimator’s expectation and the parameter it is intended to estimate.

    The decomposition is conceptually important because it displays a possible trade-off between variability and bias. An estimator with slightly larger variance may have smaller MSE if it reduces bias enough, and vice versa. “Precision” and “accuracy” should not be collapsed into a single colloquial intuition when the analysis requires their components to be distinguished.

    A cluster in sampling is a grouping of elementary units that may be used as a selection unit. It should not be confused with a cluster produced by a machine-learning algorithm. In clustering, the grouping is usually the result of an algorithmic criterion of similarity or distance; in sampling, the cluster belongs to the design and may be defined by geography, administrative organization, dwellings, schools, establishments, or other structures.

    A PSU is simply the unit selected at the first stage. In many household designs the PSU is a geographic cluster—for example, a census segment—so the two notions coincide empirically. They are not, however, synonymous by definition.

    Design First unit selected Stages Are equal probabilities mandatory?
    Simple random Element 1 Yes, in the classical simple design
    Stratified Element within stratum 1 Not necessarily across strata
    One-stage cluster Cluster 1 No
    Multistage PSU; then secondary units, etc. 2 or more No
    PPS Unit with probability proportional to size 1 or more No

    Clustering matters because units within the same cluster often resemble one another. If there is positive intraclass correlation, interviewing twenty people within the same cluster may provide less information than interviewing twenty people spread across many clusters.

    This relative loss or gain in precision is often summarized through a design effect,

    \[ \operatorname{deff} =\frac{\operatorname{Var}_{\text{design}}(\hat\theta)}{\operatorname{Var}_{\text{SRS}}(\hat\theta)}, \]

    which compares the variance under the actual design with a simple-random-sampling reference of comparable size. A \(\operatorname{deff}>1\) indicates greater variance than the reference; it is not an “error” of an algorithm, but a property that may arise from the design, intraclass correlation, and weight variability.

    19. What does “margin of error” mean?

    In the language of polls, the margin of error is usually the half-width of a confidence interval associated with sampling uncertainty under a particular design and estimation method. It is not “the total amount of error in the survey.”

    For an estimate \(\hat\theta\), a generic form is

    \[ \hat\theta\pm z_{1-\alpha/2}\,\widehat{SE}(\hat\theta), \]

    so that the margin is

    \[ MOE=z_{1-\alpha/2}\,\widehat{SE}(\hat\theta). \]

    In complex designs, \(\widehat{SE}\) must correspond to the design: strata, clusters, unequal probabilities, calibration, or replicate weights, as applicable. Calculating a margin as though the sample had been simple random when it was not can understate or overstate uncertainty.

    Sampling margin and total survey error

    The Total Survey Error paradigm requires separating sampling variability from other sources of error. These include:

    • coverage error;
    • nonresponse error;
    • measurement error;
    • construct or specification error;
    • processing and coding error;
    • editing and imputation error;
    • frame and linkage error;
    • errors arising from models or adjustments.

    Thus a phrase such as “the survey has a margin of error of ±3%” should be understood, unless otherwise specified, as a statement about one component of uncertainty. It does not demonstrate that the true value lies within three points or that every other source of error is small.

    20. Why is probability sampling preferred? Can only probability samples be useful?

    A probability design provides a known selection distribution and therefore permits design-based inferences and estimation of sampling variability under explicit conditions. This is an enormous advantage when the objective is to infer to a defined population.

    But “nonprobability” does not automatically mean “useless.” Convenience, judgment, quota, network, or self-selected samples may be appropriate for exploration, qualitative research, instrument development, mechanism studies, hard-to-reach populations, or questions for which design-based population inference is not the objective.

    The problem arises when such samples are attributed properties that their selection mechanism does not justify. If a researcher studies a localized group because it possesses a rare characteristic, valid knowledge may be produced about that group or about particular mechanisms. One cannot infer without additional assumptions that the prevalence of that characteristic in the country is equal to the prevalence observed in the group.

    The word representativeness should be used cautiously. A probability sample may suffer from undercoverage, differential nonresponse, or measurement error. It is therefore preferable to state exactly what property has been secured: a defined coverage frame, known probabilities, calibration to particular totals, balance on auxiliary variables, or inferential capacity under a specified design.

    Methodological ethics require documenting the selection method and the limits of inference. A nonprobability sample should not be presented as probabilistic; a probability sample should likewise not be presented as free from the remaining sources of survey error.

    21. “Randomization is a physical operation”: what does that mean?

    Kish insists that a probability sample requires more than writing a distribution on paper. There must be an effective selection procedure whose operation is congruent with the specified probability model.

    The word “physical” should not be understood as opposed to computation. An algorithm implemented on a computer can perfectly well form part of the material operation of selection. What Kish contrasts is an actually executed chance mechanism with a mere belief, assumption, or declaration that the units “are as if they had been random.”

    We can separate two levels:

    \[ \text{probability model}\neq\text{randomization procedure}, \]

    but a probability design requires correspondence between them. If the protocol states that each of \(N\) units has equal probability of selection, the actual procedure—a random-number table, an audited pseudorandom algorithm, a draw, or another mechanism—must implement that property.

    This distinction is philosophically interesting. Probability in a probability survey is not merely a formal property of the language used to describe selection; it is linked to an objective, reproducible practice. The theory acquires inferential capacity because a controlled relation exists between the material mechanism and the mathematically defined sample space.

    22. What does it mean that the frame includes “physical lists” and procedures that do not literally list all units?

    Kish uses frame in a broader sense than “a file containing every name.” A physical list may be a register of persons, establishments, addresses, plots, or administrative units. But there may also be procedures that allow units to be selected without constructing an exhaustive list of all final elements.

    An area frame, for example, may divide a territory into selectable segments; dwellings are then listed within the selected segments. In random-digit dialing, the frame may be conceptual or algorithmic: the space of eligible numbers is defined and selections are generated without previously possessing a file of all subscribers.

    “Actually listing” means explicitly enumerating all relevant units in an operational list. Avoiding that effort does not mean replacing the physical by “the computational” in the abstract; it means constructing a coverage and selection mechanism that does not require an exhaustive prior enumeration of all final units.

    23. Why is it so difficult to construct a frame for human populations?

    Because target population, frame, and field location are not the same thing. An electoral register, for example, may be excellent for a population of registered voters and at the same time inadequate for a population defined as “all resident persons” if it excludes minors, unregistered persons, or other groups.

    Human populations also change. People are born, die, migrate, move homes, form or dissolve households, use multiple telephones, lose numbers, share devices, and may reside temporarily in different places. A perfect frame would be costly not because we need to know where every person is at every hour, but because it must maintain a sufficiently good correspondence between target units and selectable units.

    Large statistical systems address this problem through different strategies: address frames, integrated administrative registers, area frames, dwelling listings within PSUs, dual or multiple frames, and periodic updates. Each solution distributes costs and risks of undercoverage, duplication, and obsolescence differently.

    The correct methodological question is not “does a complete list of every inhabitant exist?” but “what frame or combination of frames provides a documented path by which units of the target population have inclusion probabilities that are known or can be modeled in accordance with the design?”

    24. What is multistage sampling?

    It is a design in which selection occurs in two or more stages. In a household survey, for example, one might select:

    1. PSUs or geographic segments;
    2. dwellings within the PSUs;
    3. one person within each dwelling.

    If the selections are conditional, the overall inclusion probability of person \(i\) may be written schematically as the product of the probabilities at each stage:

    \[ \pi_i =\pi_{h}\,\pi_{j\mid h}\,\pi_{i\mid j,h}, \]

    depending on the notation adopted for PSU, dwelling, and person.

    The basic design weight is then

    \[ w_i=\frac{1}{\pi_i}. \]

    In practice, that weight may later be modified for nonresponse, calibration, poststratification, weight trimming, or other procedures. The basic design weight must always be distinguished from final analytical weights.

    Multistage sampling makes it possible to work with enormous populations without first constructing a national list of every final element. Its cost is usually much lower, although geographic concentration can increase variance relative to a simple random sample of the same size.

    25. Multiplicity, unequal probabilities, and Horvitz-Thompson estimation

    Suppose a telephone frame allows a household with two lines to appear twice while another household with one line appears once. If lines are selected with equal probability and households are the final target, household inclusion probabilities may be unequal. If that multiplicity is unknown and ignored, it may produce bias.

    But it does not follow that all random sampling requires equal probability. Survey theory explicitly includes unequal-probability designs. What is fundamental is that the probabilistic mechanism be specified and that the probabilities required for estimation be known or calculable.

    For a population \(U\) and a total

    \[ T_Y=\sum_{i\in U}Y_i, \]

    if each unit has positive inclusion probability \(\pi_i>0\), the Horvitz-Thompson estimator is

    \[ \widehat T_{HT} =\sum_{i\in s}\frac{Y_i}{\pi_i}. \]

    Its logic is immediate: a unit that is difficult to select represents more units of the population than a unit with a higher inclusion probability. Hence the inverse of \(\pi_i\) appears.

    Under the appropriate design,

    \[ E_p\left(\widehat T_{HT}\right)=T_Y, \]

    where \(E_p\) denotes expectation with respect to the sampling design. The historical importance of the Horvitz and Thompson (1952) result lies precisely in showing how samples without replacement and with unequal probabilities can be handled without thereby turning them into “nonrandom” samples.

    This allows us to return to the problem of duplicates with greater precision. If multiplicity changes \(\pi_i\) in a known way and the estimator incorporates it correctly, the inequality can be handled. If multiplicity is unknown, it becomes a frame problem that prevents the probabilities from being properly known and can distort inference.

    · · ·

    Some General Consequences

    Four principles that cut across survey theory can be drawn from the preceding problems.

    First, design and estimation form a unity. There is no estimator whose inferential meaning can be completely separated from the way the sample was produced. Weighting, variance estimation, and the interpretation of intervals depend on the design.

    Second, randomness does not necessarily mean equality. A design may be probabilistic and assign unequal probabilities; what matters is that the mechanism be probabilistically defined and that its relevant probabilities can be incorporated into inference.

    Third, sampling error is not total survey error. A narrow interval may coexist with undercoverage, differential nonresponse, poor measurement, or processing errors. The mathematical precision of one component should not be confused with overall validity.

    Fourth, the statistical operation is also a material operation. Frames, listings, interviews, selection algorithms, questionnaire skips, and editing rules are concrete practices. Theory does not float above them: its assumptions are realized—or violated—in the actual process by which data are produced.

    References

    AAPOR. (2023). Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys (10th ed.). American Association for Public Opinion Research. https://aapor.org/standards-and-ethics/standard-definitions/

    Biemer, P. P., & Lyberg, L. E. (2003). Introduction to Survey Quality. Wiley. https://doi.org/10.1002/0471458740

    Cochran, W. G. (1977). Sampling Techniques (3rd ed.). Wiley.

    Groves, R. M. (1989). Survey Errors and Survey Costs. Wiley. https://doi.org/10.1002/0471725277

    Groves, R. M., Fowler, F. J. Jr., Couper, M. P., Lepkowski, J. M., Singer, E., & Tourangeau, R. (2009). Survey Methodology (2nd ed.). Wiley.

    Horvitz, D. G., & Thompson, D. J. (1952). A generalization of sampling without replacement from a finite universe. Journal of the American Statistical Association, 47(260), 663-685. https://doi.org/10.1080/01621459.1952.10483446

    Kalton, G., & Kasprzyk, D. (1986). The treatment of missing survey data. Survey Methodology, 12(1), 1-16.

    Kish, L. (1965). Survey Sampling. Wiley.

    Lohr, S. L. (2021). Sampling: Design and Analysis (3rd ed.). CRC Press.

    Särndal, C.-E., Swensson, B., & Wretman, J. (1992). Model Assisted Survey Sampling. Springer.

    Statistics Canada. (2003). Survey Methods and Practices. Catalogue no. 12-587-X. https://www150.statcan.gc.ca/n1/en/pub/12-587-x/12-587-x2003001-eng.pdf

    United Nations Statistics Division. (2008). Designing Household Survey Samples: Practical Guidelines. Studies in Methods, Series F, No. 98. https://unstats.un.org/unsd/publication/seriesf/seriesf_98e.pdf

    U.S. Census Bureau. (2021). Questionnaire Testing and Evaluation Methods for Censuses and Surveys. https://www.census.gov/about/policies/quality/standards/appendixa2.html

    Additional Resources of Interest

    The following resources may be useful for quick orientation, discussion, examples, or bibliographic navigation. They do not replace the methodological sources cited above.

  • LOGIT MODEL OR LOGISTIC REGRESSION

    LOGIT MODEL OR LOGISTIC REGRESSION

    Mathematical Statistics · Generalized Linear Models

    Logit Model or Logistic Regression

    On statistical specification, the logit transformation, and the relation between probabilities, odds, and linear predictors


    As Aldrich and Nelson (1984) point out, every statistical inference presupposes some specification of the phenomenon one intends to study. It is therefore not enough to possess a mathematically elegant model, nor to have a computational procedure capable of estimating it. A prior problem remains: determining whether the mathematical properties of the model bear a scientifically defensible correspondence to the properties of the real object under investigation.

    This question precedes estimation. Statistical-mathematical theory can tell us what properties an estimator possesses under certain conditions; it cannot, by itself, guarantee that those conditions adequately describe the material, economic, social, or natural process from which the data arise. When specification fails, the inferential properties upon which our conclusions rest may fail as well.

    The logit model provides a useful illustration of this general problem. Its usefulness does not simply follow from its popularity, nor from the ease with which statistical software can estimate it. Rather, it derives from the specific relation it establishes between a dichotomous response variable, a conditional probability, and a mathematical function capable of transforming that probability into a quantity representable by means of a linear predictor.

    Simplicity is a reasonable starting hypothesis when no knowledge requires us to do otherwise. But where theory or evidence indicates a different structure, simplicity ceases to justify imposing it.

    1. The problem with the linear probability model

    Suppose that \(Y_i\) is a dichotomous variable that can take only the values zero and one. We may define:

    \[ P_i \equiv P(Y_i=1\mid X_i). \]

    The conditional expectation of a Bernoulli variable coincides precisely with its probability of success:

    \[ E(Y_i\mid X_i)=P_i. \]

    One first possibility is to represent that probability through a linear function of the explanatory variables:

    \[ P_i = \beta_0+\beta_1X_{i1}+\cdots+\beta_kX_{ik}. \]

    The mathematical problem appears immediately. A probability must satisfy:

    \[ 0\leq P_i\leq1, \]

    whereas a linear predictor such as \(\beta_0+\beta_1X_{i1}+\cdots+\beta_kX_{ik}\) does not, in general, possess such bounds. It may produce numbers below zero or above one. A linear probability model can be useful in certain contexts as an approximation and has some interpretive advantages, but this incompatibility between the range of the linear predictor and the range of a probability constitutes a fundamental reason to seek another specification.

    We do not need to abandon the linear predictor. What we can do is transform the probability.

    2. From probability to odds and the logit

    If \(0<P_i<1\), we can construct the ratio:

    \[ \frac{P_i}{1-P_i}. \]

    This quantity is commonly called the odds. It compares the probability that the event occurs with the probability that it does not.

    For example, if \(P_i=0.75\):

    \[ \frac{0.75}{1-0.75}=3. \]

    The odds are therefore 3 to 1 in favor of the event. This transformation removes the upper bound of the probability, but it does not remove all restrictions:

    \[ \frac{P_i}{1-P_i}\in(0,\infty). \]

    Odds remain necessarily positive. To obtain a quantity capable of spanning the entire real line, we apply the natural logarithm:

    \[ \log\left(\frac{P_i}{1-P_i}\right)\in(-\infty,\infty). \]

    This transformation is the logit of the probability:

    \[ \operatorname{logit}(P_i) = \log\left(\frac{P_i}{1-P_i}\right). \]

    We now have a quantity whose range is compatible with that of a linear predictor. We may therefore specify:

    \[ \log\left(\frac{P_i}{1-P_i}\right) = \beta_0+\beta_1X_{i1}+\cdots+\beta_kX_{ik} \equiv Z_i. \]
    Fundamental distinction

    The probability \(P_i\) lies in the interval \((0,1)\); the odds \(P_i/(1-P_i)\) lie in \((0,\infty)\); and the log-odds, or logit, range across the entire real line \((-\infty,\infty)\). These are three different scales and should not be confused.

    3. The logistic function

    Once the preceding relation has been established, we can solve for \(P_i\). We begin from:

    \[ \log\left(\frac{P_i}{1-P_i}\right)=Z_i. \]

    Applying the exponential function:

    \[ \frac{P_i}{1-P_i}=e^{Z_i}. \]

    Through algebraic manipulation:

    \[ P_i = \frac{e^{Z_i}}{1+e^{Z_i}} = \frac{1}{1+e^{-Z_i}}. \]

    We have arrived at the standard logistic function. The transformation solves precisely the problem of interest: \(Z_i\) may take any real value, while the output of the function always lies between zero and one.

    01

    If \(Z_i\rightarrow-\infty\), then \(P_i\rightarrow0\).

    02

    If \(Z_i=0\), then \(P_i=0.5\).

    03

    If \(Z_i\rightarrow+\infty\), then \(P_i\rightarrow1\).

    The function is continuous, smooth, and monotonically increasing. Its shape is the familiar sigmoid, or S-shaped curve. The standard logistic function also possesses rotational symmetry around the point \((0,1/2)\).

    A more general expression of the logistic function is:

    \[ f(x) = \frac{L}{1+e^{-k(x-x_0)}}, \]

    where \(L\) determines the upper limit, \(k\) regulates the rate of growth, and \(x_0\) determines the position of the midpoint. The standard logistic function commonly used to transform a linear predictor into a probability corresponds to:

    \[ L=1,\qquad k=1,\qquad x_0=0. \]

    It is also important to distinguish the logistic function used here as the inverse of the link function from the logistic distribution. The two objects are related, but they are not synonymous.

    4. Why a logistic function?

    Finding a mathematically convenient function does not, by itself, settle the problem of specification. As Aldrich and Nelson (1984) observe, many functions can transform an unrestricted predictor into a probability. The logistic function is not the only possibility.

    Among the best-known alternatives is the probit model, which uses the standard normal cumulative distribution function as the inverse link. Another possibility is the complementary log-log link. Each choice imposes a different mathematical structure on the relationship between the predictor and the probability.

    Model Link function Main feature
    Logit \(\log[p/(1-p)]\) Natural interpretation through odds and odds ratios.
    Probit \(\Phi^{-1}(p)\) Uses the cumulative normal distribution.
    Complementary log-log \(\log[-\log(1-p)]\) Introduces an asymmetric probability response.

    Logit and probit often produce very similar empirical results in a wide range of applications, particularly in the central regions of their respective curves. This does not mean that they are mathematically identical, nor that the choice between them can always be treated as irrelevant. The proper criterion remains substantive and statistical: what structure is justifiable for the phenomenon, what inferential properties are needed, and what aspects of the response one wishes to represent.

    The availability of a statistical model does not constitute evidence that the model corresponds to the object. Choosing a link function is itself a hypothesis about the mathematical form of the relationship under study.

    5. The statistical model

    Up to this point we have examined mainly the geometry of the transformation. But logistic regression is not merely the drawing of an S-shaped curve. It is a probabilistic model.

    When each observation represents an individual dichotomous outcome, the natural formulation is:

    \[ Y_i\mid X_i \sim \operatorname{Bernoulli}(P_i), \]

    with:

    \[ P_i = \frac{1} {1+\exp[-(\beta_0+\beta_1X_{i1}+\cdots+\beta_kX_{ik})]}. \]

    If, by contrast, observations are grouped and we record \(Y_i\) successes out of \(n_i\) trials, we may write:

    \[ Y_i\mid X_i \sim \operatorname{Binomial}(n_i,P_i). \]

    This distinction matters. It is not precise simply to say that “the data set follows a binomial distribution.” The probability distribution is specified for the response variable conditional on the covariates and depends on the observational structure of the data.

    In the language of generalized linear models, logistic regression combines three elements:

    I

    Random component: a Bernoulli or binomial response.

    II

    Systematic component: the linear predictor \(X_i\beta\).

    III

    Link function: \(g(P_i)=\log[P_i/(1-P_i)]\).

    The logistic function \(g^{-1}(Z)=1/(1+e^{-Z})\) is therefore the inverse of this link.

    6. Estimation by maximum likelihood

    The coefficients of a logistic regression are not usually obtained by minimizing the sum of squared residuals as in ordinary linear regression. They are estimated by maximum likelihood.

    For conditionally independent Bernoulli observations, the contribution of each observation to the likelihood can be written as:

    \[ P_i^{Y_i}(1-P_i)^{1-Y_i}. \]

    For \(n\) observations:

    \[ L(\beta) = \prod_{i=1}^{n} P_i^{Y_i}(1-P_i)^{1-Y_i}. \]

    The procedure searches for the parameter values \(\beta\) that make the actually observed outcomes most likely under the specified model. In practice, one normally maximizes the logarithm of the likelihood:

    \[ \ell(\beta) = \sum_{i=1}^{n} \left[ Y_i\log(P_i) +(1-Y_i)\log(1-P_i) \right]. \]

    In general there is no closed-form expression analogous to the ordinary least squares formula for obtaining the coefficients. The solution is found through iterative numerical procedures.

    7. How to interpret the coefficients

    This is one of the points that most often generates confusion when moving from linear to logistic regression. If:

    \[ \log\left(\frac{P_i}{1-P_i}\right) = \beta_0+\beta_1X_{i1}+\cdots+\beta_kX_{ik}, \]

    then \(\beta_j\) does not directly represent the change in probability produced by a one-unit increase in \(X_j\). It represents the change in the log-odds, holding the other variables in the predictor constant.

    Exponentiating the coefficient gives a more intuitive interpretation:

    \[ e^{\beta_j}. \]

    This is the factor by which the odds are multiplied when \(X_j\) increases by one unit, ceteris paribus.

    Example

    If \(\beta_j=0.5\), then:

    \[ e^{0.5}\approx1.65. \]

    A one-unit increase in \(X_j\) multiplies the odds by approximately 1.65, that is, increases them by approximately 65%, holding the other variables constant. This does not mean that the probability rises by 0.5 or by 50 percentage points.

    The effect on probability is not constant

    The reason is that the relationship between \(Z_i=X_i\beta\) and \(P_i\) is nonlinear. For a continuous variable \(X_j\), under the simple additive specification, the marginal effect is:

    \[ \frac{\partial P_i}{\partial X_{ij}} = \beta_jP_i(1-P_i). \]

    Thus, the same coefficient \(\beta_j\) may correspond to different changes in probability depending on the point of the curve at which each observation lies. The response is more sensitive in the central region of the logistic function and less sensitive near its extremes.

    This is not an accidental difficulty of the model. It is a direct consequence of the mathematical structure we have chosen to use.

    8. Evaluation and diagnostics

    Estimating a set of coefficients does not settle the problem. We must ask whether the model adequately describes the information we seek to explain and, above all, whether the inferences we wish to draw from it are defensible.

    Likelihood-ratio test

    A classical comparison contrasts a model containing explanatory variables with a model containing only the intercept. The likelihood-ratio statistic can be written as:

    \[ LR = 2\left[ \ell(\widehat{\beta}_{\text{model}}) – \ell(\widehat{\beta}_{\text{restricted}}) \right], \]

    and, under regularity conditions, it approximately follows a chi-squared distribution with degrees of freedom corresponding to the number of restrictions being tested.

    This test can indicate whether the set of predictors improves fit relative to a more restricted model. It should not, however, be confused with a complete evaluation of the model’s scientific usefulness.

    Statistical significance is not enough

    Depending on the purpose of the investigation, one should also examine aspects such as:

    • the calibration of estimated probabilities;
    • the ability to discriminate between outcomes;
    • deviance and residuals appropriate to generalized linear models;
    • the stability of the coefficients;
    • the presence of influential observations;
    • out-of-sample performance when the objective is predictive;
    • and, above all, the adequacy of the functional form and assumptions to the scientific problem.

    9. Specification returns to the center

    We thus return to the problem with which we began. Logistic regression elegantly solves a mathematical difficulty: it allows an unrestricted predictor to be related to a probability necessarily bounded between zero and one. But this mathematical virtue does not make the logit an automatic choice for every dichotomous phenomenon.

    Its use presupposes a particular structure. Among other things, one should examine whether the linear predictor is an adequate representation of the log-odds, whether dependence among observations has been modeled correctly, whether relevant interactions or nonlinearities are present, and whether complete or quasi-complete separation occurs, a circumstance capable of producing problematic or non-finite maximum-likelihood estimates.

    Some questions prior to inference
    • Is there a scientific justification for the variables included and for the way in which they enter the predictor?
    • Does the approximately linear relationship assumed actually hold for the log-odds, or are transformations, nonlinear terms, or interactions required?
    • Is the dependence structure among observations represented adequately?
    • Is there enough information in the data to identify the parameters?
    • Is the purpose explanatory, predictive, descriptive, or causal?

    The last question deserves particular attention. An association obtained through logistic regression does not become causal merely because it has been expressed through coefficients, probabilities, or odds ratios. Causality requires assumptions and a design capable of justifying such an interpretation.

    Likewise, a model may be a useful approximation without being a literally exact description of the data-generating process. The scientific problem is then to determine which properties of that approximation are sufficiently robust for the specific inferential purpose.

    · · ·

    The broader lesson therefore extends beyond logistic regression. A statistical model is a mathematical instrument containing a specific structure: a space of possible outcomes, a probability distribution, a functional form, parameters, and assumptions concerning the relationships among magnitudes. To apply it is to affirm—explicitly or implicitly—that these properties bear some relevant correspondence to the phenomenon under study.

    From this standpoint, the problem of specification has a genuinely epistemological content. Mathematics is not a decoration subsequently placed upon the data. Nor do the data, by themselves, determine which model must be used. Between the real object and the mathematical expression there is a theoretical mediation that must be examined, criticized, and tested.

    The logit constitutes a particularly elegant solution because it transforms a restricted probability into an unrestricted quantity and subsequently allows a valid probability to be recovered through the logistic function:

    \[ \boxed{ \log\left(\frac{P_i}{1-P_i}\right)=X_i\beta \quad\Longleftrightarrow\quad P_i=\frac{1}{1+e^{-X_i\beta}} }. \]

    But the elegance of the transformation does not replace the investigation of the object. Precisely because the model possesses determinate mathematical properties, we must ask when those properties are pertinent. There lies the point of contact between statistics, mathematics, and scientific knowledge.

    Terminological note
    In this text, logit denotes the link function \(\log[p/(1-p)]\), while logistic function denotes its inverse \(1/(1+e^{-z})\). In applications with individual binary responses, a Bernoulli distribution conditional on the covariates is used; with grouped counts of successes, a binomial distribution may be used.

    References

    Aldrich, J. H., & Nelson, F. D. (1984). Linear Probability, Logit, and Probit Models. Beverly Hills: Sage University Papers Series, Quantitative Applications in the Social Sciences.

    Liao, T. F. (1994). Interpreting Probability Models: Logit, Probit, and Other Generalized Linear Models. Sage University Papers Series, Quantitative Applications in the Social Sciences.

    McCullagh, P., & Nelder, J. A. (1989). Generalized Linear Models (2nd ed.). London: Chapman & Hall.

  • THE COLLATZ CONJECTURE AS A CUSTOM FUNCTION IN R

    THE COLLATZ CONJECTURE AS A CUSTOM FUNCTION IN R

    Mathematics · R Programming · Discrete Dynamical Systems

    The Collatz Conjecture as a Custom Function in R

    An introduction to functions, iteration, and recursion through an extraordinarily simple mathematical rule whose global behavior continues to resist a general proof.


    A particularly pleasant example of a custom function is one built to study the celebrated Collatz conjecture computationally. The choice is useful because the mathematical rule is very simple, whereas its global behavior is not: one need only distinguish between even and odd numbers and repeatedly apply the same transformation. This combination allows us to concentrate on programming logic without losing sight of the mathematical object being programmed.

    In elementary terms, the conjecture states that if we begin with any positive integer and divide it by two when it is even, or multiply it by three and add one when it is odd, the resulting sequence will eventually reach one. It is important to formulate this carefully: a program can verify that this happens for a particular starting value—or for a finite collection of values—but such computational verification does not, by itself, constitute a proof that it happens for every positive integer. That is precisely where the mathematical problem lies.

    1. The Collatz rule and what it means to verify it

    Mathematically, the rule generating successive images can be expressed by the following piecewise-defined function:

    T(n) =
    ConditionOperationNew image
    n is evenDivide by 2T(n) = n/2
    n is oddMultiply by 3 and add 1T(n) = 3n + 1

    Table 1. The local Collatz rule. Each step is completely determined by the parity of the present state.

    If we begin from an origin or seed value n0, successive applications of the function construct an orbit:

    n0,   n1 = T(n0),   n2 = T(n1),   … ,   nk = Tk(n0).

    This notation lets us introduce from the outset a distinction that will later matter in R. If the orbit reaches one for the first time after k applications of the rule, then k iterations or transitions have occurred, whereas the complete list of states contains k + 1 values, because the starting value is counted as well.

    A useful distinction

    The total stopping time may be defined as the number of applications of T required to reach 1. If it is denoted by τ(n), the conjecture can be written compactly as: τ(n) < ∞ for every positive integer n.

    Consequently, what we shall program below is not a “proof” of the conjecture in the mathematical sense of a universal demonstration. It will instead be a mechanism capable of generating and examining particular orbits, recording their states, and counting how many transformations were required to reach 1 whenever that value is indeed reached.

    · · ·

    2. A first custom function in R

    To construct the preceding rule in R, it is useful to recall that a custom function can take an input, subject it to a set of instructions, and return an output. At this first stage there will be no iteration yet: a single number will be evaluated and its next Collatz image produced.

    The fundamental condition is parity. In R, the operator %% returns the remainder of a division. If n %% 2 == 0, the number is divisible by two and is therefore even. If the remainder is nonzero, it is odd.

    R · one-step ruleCollatzFunction <- function(n) {
      if (length(n) != 1L || !is.finite(n) || n < 1 || n != floor(n)) {
        stop("n must be a positive integer.")
      }
    
      if (n %% 2 == 0) {
        j <- n / 2
      } else {
        j <- 3 * n + 1
      }
    
      return(j)
    }

    The function receives a positive integer n, determines whether it is even or odd, and stores the result in j. The instruction return(j) does not, strictly speaking, mean “print j on the screen”; it means that j will be the value returned by the function. When the function is run directly in R’s interactive console, that returned value is usually displayed automatically, which explains why the two ideas can easily be confused when one is beginning to program.

    Return and print are not the same

    return(x) terminates the function and supplies x as its result. print(x), by contrast, explicitly requests that x be printed. A function can return a value without containing any print() instruction.

    Notice also that there is no need to prevent this function from being evaluated at 1: under the standard rule, 1 is odd and its next image would be 4. What the algorithm constructed below will do is stop the orbit upon reaching 1; that is, it will not apply the rule again once the target has been reached.

    3. From a single step to an iterative algorithm

    The preceding function performs only one transformation. To obtain a complete orbit, it must be repeated while the current value differs from 1. Here we encounter the “while” loop—while-loop—whose logic is exactly what its name suggests: execute a given block of instructions for as long as a condition remains true.

    This second function is therefore iterative. It is not recursive yet, because it does not call itself; rather, it explicitly repeats a block of code by means of while().

    R · iterative algorithmIterativeCollatzAlgorithm <- function(m, max_iter = 1000000L) {
      if (length(m) != 1L || !is.finite(m) || m < 1 || m != floor(m)) {
        stop("m must be a positive integer.")
      }
    
      trajectory <- m
      iterations <- 0L
    
      while (m != 1) {
        if (iterations >= max_iter) {
          stop("max_iter was reached before 1.")
        }
    
        m <- CollatzFunction(m)
        trajectory <- c(trajectory, m)
        iterations <- iterations + 1L
      }
    
      return(trajectory)
    }

    The logic is simple. The vector trajectory is created inside the function itself and initially contains the seed value. Then, each time the loop runs, m is replaced by its next image and that new value is appended to the end of the vector. When m reaches 1, the condition m != 1 ceases to hold and the loop terminates.

    The letter m is used here simply to distinguish visually the argument of this second function from the n used in the first. There is no syntactic need to do so: arguments and objects created inside each function have their own local scope, so two different functions may use an argument called n without conflict.

    Defining the vector inside the function has an important advantage: each execution retains its own state and does not depend on objects already present in R’s global environment. In other words, if the function is run once with 27 and then with 8, the second orbit does not inherit values from the first.

    The argument max_iter is not part of the Mathematics of Collatz; it is a computational precaution. Since the general statement one wishes to prove cannot simply be assumed when writing the program, it is prudent to prevent a loop from continuing without limit if it is used with unexpected data or a modified implementation.

    ObjectTypeRole
    CollatzFunctionfunctionComputes a single image under the Collatz rule.
    IterativeCollatzAlgorithmfunctionRepeats the rule with while() and constructs the complete trajectory down to 1.
    trajectorynumeric vectorExists locally during one execution and stores the states visited.

    Table 2. A conceptual equivalent of what would be observed in the R environment after defining the functions. The trajectory vector need not remain in the global environment.

    · · ·

    4. The orbit of 27: 112 states and 111 iterations

    We can now evaluate any number in the algorithm we have constructed. Take the classic example of 27. The sequence begins 27, 82, 41, 124, 62, 31… and continues through rises and falls until it finally reaches 1.

    R · exampleorbit_27 <- IterativeCollatzAlgorithm(27)
    length(orbit_27)
    # 112
    
    length(orbit_27) - 1L
    # 111
    
    max(orbit_27)
    # 9232

    This is a good place to make the terminology precise. The stored orbit contains 112 states, because it includes both the starting point 27 and the endpoint 1. However, between 112 states there are only 111 transitions. Consequently, the total stopping time of 27, understood as the number of applications of the rule required to reach one, is 111.

    k=027k=182k=241k=3124k=462k=531k=694k=747
    k=8142k=971k=10214k=11107k=12322k=13161k=14484k=15242
    k=16121k=17364k=18182k=1991k=20274k=21137k=22412k=23206
    k=24103k=25310k=26155k=27466k=28233k=29700k=30350k=31175
    k=32526k=33263k=34790k=35395k=361186k=37593k=381780k=39890
    k=40445k=411336k=42668k=43334k=44167k=45502k=46251k=47754
    k=48377k=491132k=50566k=51283k=52850k=53425k=541276k=55638
    k=56319k=57958k=58479k=591438k=60719k=612158k=621079k=633238
    k=641619k=654858k=662429k=677288k=683644k=691822k=70911k=712734
    k=721367k=734102k=742051k=756154k=763077k=779232k=784616k=792308
    k=801154k=81577k=821732k=83866k=84433k=851300k=86650k=87325
    k=88976k=89488k=90244k=91122k=9261k=93184k=9492k=9546
    k=9623k=9770k=9835k=99106k=10053k=101160k=10280k=10340
    k=10420k=10510k=1065k=10716k=1088k=1094k=1102k=1111

    Table 3. Complete orbit of 27. The label k indicates how many applications of the rule have been performed. The maximum, 9232, appears at k = 77; one appears at k = 111.

    The trajectory also reveals a property that makes the problem so suggestive: an extraordinarily simple deterministic rule does not necessarily generate a visually simple evolution. The fact that the system is perfectly determined step by step does not mean that its global behavior is obvious.

    If one wishes to retain the traditional graph in R, it can be generated, for example, with:

    R · optional visualizationplot(
      orbit_27,
      type = "l",
      xlab = "Iteration k",
      ylab = "n_k"
    )

    The table presents the complete orbit so that each state can be read directly and followed step by step.

    5. A second solution: recursion

    There is another way to design the algorithm. It may be somewhat less intuitive when first imagined, but it has a notable elegance: instead of constructing a while() loop, the function itself calls itself again after calculating the next value.

    The problem now requires a function with two principal characteristics: its input is a positive integer representable by the numerical system being used; its output contains, on the one hand, the trajectory followed and, on the other, the number of transformations performed in order to reach one.

    R · recursive versionRecursiveCollatz <- function(n, v = NULL) {
      if (length(n) != 1L || !is.finite(n) || n < 1 || n != floor(n)) {
        stop("n must be a positive integer.")
      }
    
      v <- c(v, n)
    
      if (n == 1) {
        return(list(
          steps = v,
          iterations = length(v) - 1L
        ))
      }
    
      if (n %% 2 == 0) {
        n <- n / 2
      } else {
        n <- 3 * n + 1
      }
    
      return(RecursiveCollatz(n, v))
    }

    Here recursion is used in the precise programming sense: during the execution of RecursiveCollatz(), a new call to RecursiveCollatz() appears. That new call receives a different value of n and an enlarged copy of the vector v.

    The condition n == 1 is the base case. Without a reachable base case, a recursive definition would continue producing additional calls until the available resources were exhausted. When 1 is reached, no new call is created: the list containing the orbit and the number of iterations is returned.

    Three standard functions used

    c() concatenates elements or vectors. list() constructs a list whose components may have different names and types. length() returns the number of elements in an object such as a vector or list; it should not be confused with dim(), which reports the dimensions of objects that possess a dimensional attribute.

    The expression iterations = length(v) – 1L deserves attention. If v contains both the initial value and the final one, its length counts states, not transformations. Subtracting one corrects exactly that difference.

    SeedReturned statesIterations
    88 → 4 → 2 → 13
    27112 states from 27 to 1111

    Table 4. The difference between counting states and counting applications of the function. In the trajectory 8 → 4 → 2 → 1 there are four values, but only three steps.

    A practical qualification must be added. Recursion is an excellent tool for understanding the logical structure of the problem, but it is not always the most robust way to traverse long orbits in R: each recursive call adds a new execution frame, and the available depth is finite. For intensive computational work, the iterative version is usually preferable; for understanding what it means for a function to invoke itself, the recursive version is particularly transparent.

    · · ·

    6. How to think about recursion mathematically

    Before programming a recursive solution, it is useful to separate two things which, although intimately related, are not identical: the mathematical dynamics and the program’s call structure. The fundamental mathematical function remains T. What changes is the way in which the program organizes the calculation of its iterations.

    The algorithm can be thought of as follows:

    Logical stepAction
    1Enter a positive integer N.
    2Record N as a new state of the trajectory.
    3If N = 1, return the trajectory and terminate.
    4aIf N is even, assign NN/2.
    4bIf N is odd, assign N ← 3N + 1.
    5Execute the same function again using the new value of N.

    Table 5. Use case for the recursive solution. Recursion appears in the last step: the function is executed again with a new input.

    Mathematically, however, there is no need to define a new function C(n) whose value is always 1 whenever the recursion terminates. Such a notation would conceal the object that actually matters. It is cleaner to retain T as the Collatz map and define the orbit through successive iterations:

    nk+1 = T(nk),    k = 0, 1, 2, …

    The total stopping time may then be written as the first index at which the orbit reaches one:

    τ(n) = min { k ≥ 0 : Tk(n) = 1 }, whenever such a k exists.

    This formulation makes the real difficulty visible. We know exactly the local rule that takes one state to the next. What does not follow automatically from that definition is that τ(n) is finite for every positive integer.

    Take n = 8:

    k = 0k = 1k = 2k = 3
    8421

    Table 6. From 8 to 1 there are four states and three applications of T. Therefore, τ(8) = 3.

    The recursive function reproduces this same chain computationally, but each arrow now corresponds not only to a mathematical transformation but also to a new function call. The mathematical sequence advances from 8 to 4, from 4 to 2, and from 2 to 1; meanwhile, the program creates successive execution levels until it encounters the base case.

    7. The call stack: why it is LIFO

    To understand what happens during recursion, we can use a deliberately simple example. Suppose there is a “Listing X” with the following instructions:

    Listing X

    1. Execute A1.
    2. Execute A2.
    3. Execute the instructions in Listing X.
    4. Return and print “Hello, Fernanda”.

    Upon reaching step 3, the program begins another execution of the same listing. That new execution reaches its own step 3 and creates another, and so on. Because there is no base case here, step 4 is never executed: the resources that the environment permits for nested calls will be exhausted first.

    LevelA1A2Step 3Result
    Listing X₀executesexecutescalls X₁remains pending
    Listing X₁executesexecutescalls X₂remains pending
    Listing X₂executesexecutescalls X₃remains pending

    Table 7. Structure of recursive calls. Each call is suspended while it waits for the call created within it to finish.

    When a base case does exist, as in our algorithm when 1 is reached, the situation changes. The deepest call terminates first; then the call that created it can terminate; then the preceding one, and so on. This is why the natural structure of nested calls is a stack, organized as LIFO: Last In, First Out.

    TOP OF STACK · call with n = 1 · terminates first
    call with n = 2 · waits for the result of n = 1
    call with n = 4 · waits for the result of n = 2
    BASE · call with n = 8 · created first

    Table 8. The call stack for 8 → 4 → 2 → 1. The call with 1 is the last to enter and the first able to finish; the pending calls are then resolved in reverse order.

    The analogy with a stack of plates remains useful: if several plates are stacked one on top of another, the last one placed sits at the top and is the first that can be removed without dismantling the stack. It is nevertheless useful to specify the technical reason: recursive calls are organized in this way because of the nested structure of program control. The relationship with memory locality or possible cache effects may matter in other performance contexts, but it is not what explains why the call stack is LIFO.

    A call A that invokes B cannot continue beyond that point until B has finished; and if B invokes C, B cannot continue until C finishes either. The order C → B → A therefore arises naturally.
    · · ·

    8. What Collatz teaches us about Mathematics and Computation

    There is still a conceptual question that deserves attention. At first glance, it might seem that the mathematical formulation offers us a “continuous” process whereas the program turns it into a collection of discrete steps. That opposition, however, does not correctly describe this problem. Collatz dynamics are discrete in the mathematical definition itself: iteration time is indexed by 0, 1, 2, 3, … and each state belongs, in the usual formulation, to the set of positive integers.

    The interesting difference, therefore, is not between continuous Mathematics and discrete Computation. It lies elsewhere: between defining a rule, calculating its consequences for particular cases, and proving a global property of all its orbits.

    LevelWhat we knowWhat does not follow automatically
    Mathematical ruleFor each state, we know exactly what the next one is.From this alone, we do not know the global fate of every possible orbit.
    Computational calculationWe can follow a particular orbit step by step and verify whether it reaches 1.A finite number of verifications is not equivalent to a universal proof.
    Mathematical proofIt would seek to establish a proposition valid for every positive integer.It cannot simply be replaced by enumerating more and more cases.

    Table 9. Three levels of knowledge that Collatz forces us to distinguish: local specification, computational experimentation, and global proof.

    This distinction allows us to formulate more precisely the epistemological reflection that makes the example attractive. Mathematics does not necessarily consist in placing a phenomenon inside an analytical form whose behavior is already known. Collatz shows exactly the opposite: we can possess a perfectly defined rule and still lack a complete characterization of its global behavior.

    Nor can Computation be reduced to attending to “particularities” that Mathematics has supposedly abstracted away. Here both operate on the same discrete object, but they ask different questions. The program allows us to execute the rule, inspect orbits, measure stopping times, discover regularities, and subject auxiliary conjectures to numerical testing. Mathematical proof, by contrast, asks what can necessarily be asserted for an infinite class of cases.

    Knowing exactly the law that governs each step is not the same as knowing in advance the global behavior of the system.

    This is probably one of the most interesting lessons that can be drawn from the problem from the joint perspective of Mathematics and Computer Science. In Collatz there is local determination, effective computability of every step, and an enormous body of evidence for particular values; but those three things must not be confused with a universal proof.

    9. A precaution: mathematical integers are not machine integers

    Finally, there is an additional difference between the mathematical object and its concrete implementation that should be made explicit. The expression “any N ∈ ℕ” describes an unlimited mathematical domain. A computer, by contrast, represents numbers by means of finite structures.

    Base R includes, among other things, values of type integer, whose range is limited, and numeric values, which normally use double-precision arithmetic. The latter can represent all integers exactly only up to a certain size; above 253, not every consecutive integer has an exact representation as a double-precision number.

    The issue is particularly relevant in Collatz because an orbit can grow far above its initial value before beginning to descend. Consequently, an implementation intended to investigate very large numbers must use exact integer arithmetic of greater capacity—for example, arbitrary-precision integers—and carefully check for possible overflow or loss of exactness.

    Mathematics versus implementation

    Mathematical object: T acts on all positive integers.
    Concrete program: it can operate correctly only on values that its numerical representation and memory resources can manipulate exactly.

    This does not invalidate the pedagogical use of the preceding code. For small seeds such as 8 or 27, the implementation is perfectly adequate. What it prevents is our attributing to a finite computational representation the same unlimited domain possessed by the abstract mathematical object.

    · · ·

    Final considerations

    The Collatz conjecture is an extraordinarily fertile example for learning to program because it allows us, almost without introducing artificial devices, to move from an elementary mathematical rule to several central ideas in Computer Science. First we construct a function that chooses between two operations; then that operation is turned into an iteration; next the orbit is stored; finally the same process is reconstructed through recursion, and we observe how nested calls are organized in a stack.

    But the example is equally fertile for a deeper reason. A sequence may be governed by a completely deterministic rule and yet pose enormous difficulties when one tries to establish a global property for all of its trajectories. The program knows what to do at each step. The mathematical difficulty lies in justifying what will happen after all the necessary steps, for every possible starting point.

    In this sense, programming does not replace proof, nor does proof make programming unnecessary. The former allows us to experiment with the object, traverse it, and produce concrete evidence; the latter seeks to establish relations that do not depend on having previously enumerated every case. Collatz provides an especially clean setting in which to observe this difference because there is no obscurity whatsoever in the elementary rule: the difficulty emerges from the dynamics produced by repeating it.

    And precisely for that reason, a custom function in R is more than an exercise in syntax. It allows us to observe, on a small and manageable scale, the passage between definition, algorithm, execution, numerical representation, and mathematical reasoning; that is, between different levels of the same scientific activity which should remain connected without being confused.

  • BAYES ESTIMATORS, CLUSTER ANALYSIS, AND GAUSSIAN MIXTURES

    BAYES ESTIMATORS, CLUSTER ANALYSIS, AND GAUSSIAN MIXTURES

    Probability · Statistical Theory · Unsupervised Learning

    From Bayes Estimators to Gaussian Mixtures

    A guided reading of José Mauricio Gómez Julián’s 2020 study on Bayesian decision theory, cluster analysis, density estimation, the EM algorithm, and model-based clustering in R.

    Research conducted in 2020 José Mauricio Gómez Julián Approx. 15-minute read

    Statistics often becomes difficult not because any single idea is impossible to understand, but because several ideas must suddenly be held together at once. A probability distribution leads to an estimator; an estimator leads to a loss function; a loss function leads to an optimization problem; an optimization problem leads to an algorithm; and the algorithm finally produces something that looks deceptively simple on the screen: a handful of clusters. Gómez Julián’s 2020 study, On Bayes Estimators, Cluster Analysis and Gaussian Mixtures, is essentially an attempt to reconstruct that entire chain of reasoning before asking R to perform the calculation.

    Its destination is model-based cluster analysis with finite Gaussian mixture models. But the paper deliberately takes the long road. Before arriving at Gaussian mixtures and the Mclust() routine, it moves through posterior probability, Bayesian estimators, loss functions, mean squared error, information criteria, data mining, machine learning, supervised and unsupervised learning, density estimation, k-means, the expectation-maximization algorithm, categorical and Dirichlet distributions, optimization, parameters and hyperparameters. The point is not simply to list definitions. It is to show why these concepts meet inside the same statistical machine.

    Reading note · what this article will do

    The original 2020 study is much broader than a conventional software tutorial. This guided reading therefore concentrates on its main statistical architecture: how Bayesian reasoning, cluster analysis, Gaussian mixture models, the EM algorithm, and BIC-based model selection fit together, and what happens when that framework is applied in R.

    · · ·

    01 · Context Why this 2020 study was written

    The immediate motivation is educational research. Gómez Julián begins from a doctoral study by Villegas Barahona concerned with academic performance, latent dimensions, directly observed student variables, CUR matrix decomposition, and the construction of a statistical model capable of supporting academic and administrative decision-making. That earlier project provides the practical problem; the 2020 study asks what statistical theory one must understand in order to follow the machinery being used.

    This matters because a statistical package can make a difficult procedure look trivial. A researcher can type a command, obtain a classification, inspect a graph, and move on. Yet the command silently presupposes answers to difficult questions. What is being estimated? What does it mean for observations to belong to different groups? What happens when those memberships are not observed? How should the number and geometry of the groups be chosen? What quantity is the algorithm maximizing? And how much uncertainty remains after a point has been assigned to a cluster?

    These are not merely programming questions. They are questions about probability, inference, geometry and decision-making. The paper’s distinctive strategy is therefore foundational: instead of treating Gaussian mixture models as a black box, it reconstructs the conceptual staircase leading to them.

    A cluster on a computer screen is the final visible result of a much longer argument about probability, hidden structure, estimation and optimization.

    Conceptual summary of Gómez Julián’s 2020 framework

    There is another reason this approach is useful outside statistics. Political scientists, economists, sociologists and public-policy researchers frequently work with populations that are heterogeneous. Countries, households, firms, voters or students may appear in one dataset while actually belonging to several statistically distinct subpopulations. If those subpopulations are not directly labelled, the analytical task is no longer simply to estimate an average. It is to infer the hidden structure that may have generated the observations.

    That is where cluster analysis and Gaussian mixture models eventually enter. But Gómez Julián begins one layer deeper: with Bayes and the logic of updating knowledge when new evidence arrives.

    · · ·

    02 · Bayesian foundations Bayes as a rule for learning from evidence

    At its simplest, Bayes’ theorem tells us how a probability should change when we acquire relevant information. Suppose we have a hypothesis \(H\) and observe some data \(D\). Bayesian updating connects four quantities:

    \( P(H\mid D)=\dfrac{P(D\mid H)\,P(H)}{P(D)} \) posterior = likelihood × prior / marginal probability of the data

    The prior, \(P(H)\), represents the state of information before the new evidence is incorporated. The likelihood, \(P(D\mid H)\), tells us how compatible the observed evidence is with the hypothesis. The denominator \(P(D)\) normalizes the calculation. The result, \(P(H\mid D)\), is the posterior: the probability conditional on having observed the new evidence.

    For readers coming from economics, there is a useful analogy. Imagine beginning with a set of beliefs about the likely position of an economy, then receiving new information about employment, inflation or production. The point of Bayesian updating is not that the old information disappears. Rather, prior information and new evidence are combined according to a precise probabilistic rule. The posterior then becomes the informational starting point for whatever decision comes next.

    The key intuition

    Bayes’ theorem is not yet a clustering algorithm. Its importance here is more fundamental: Gaussian mixture models repeatedly ask conditional-probability questions. Given an observed data point, how probable is it that the point came from component 1, component 2, component 3, and so on? Once group membership is hidden rather than directly observed, posterior probabilities become a natural language for reasoning about that uncertainty.

    This is one of the conceptual bridges that makes the paper coherent. What begins as an abstract discussion of conditional probability will later reappear in a very concrete form: each observation can carry a probability of membership in each possible Gaussian component. That is already a major difference between a Gaussian mixture model and the familiar hard assignment produced by ordinary k-means.

    03 · Statistical decision-making From posterior probability to a Bayes estimator

    Updating probabilities is only part of the story. Eventually, an analyst has to do something with the posterior distribution. A parameter must be estimated, a prediction must be produced, a model must be selected, or an observation must be assigned—perhaps provisionally—to a group.

    This is why Gómez Julián’s 2020 study moves from Bayes’ theorem into decision theory. Once several possible estimates or actions are available, the statistical problem can be expressed as a question of consequences: if the unknown quantity is really \(\theta\), what is the cost of reporting some estimate \(\hat{\theta}\)?

    That cost is represented by a loss function. The exact form of the function depends on what kinds of errors matter in the problem under study. One particularly important case, and the one emphasized in the paper, is squared-error loss:

    \( L(\theta,\hat{\theta})=(\hat{\theta}-\theta)^2 \) a larger distance between the estimate and the unknown parameter produces a disproportionately larger loss

    Squaring the error has two immediate consequences. First, positive and negative deviations no longer cancel each other. Second, large errors are penalized more heavily than small ones. The associated expected loss is therefore closely connected with the familiar mean squared error.

    \( \mathrm{MSE}(\hat{\theta}) = E_{\theta}\!\left[(\hat{\theta}-\theta)^2\right] \) mean squared error as an expected measure of estimation error

    Bayesian decision theory adds one decisive ingredient: rather than evaluating loss while treating the parameter as an unknown fixed object outside the probability calculation, the posterior distribution is used to average the possible consequences of a decision. The relevant quantity becomes the posterior expected loss.

    \( \rho(a\mid x) = \int_{\Theta} L(\theta,a)\, p(\theta\mid x)\,d\theta \) posterior expected loss: consequences averaged over current uncertainty about the parameter

    A Bayesian decision rule selects the action that minimizes this quantity. Under squared-error loss, something especially elegant happens: the optimal estimate is the posterior mean.

    \( \hat{\theta}(x) = E(\theta\mid x) = \int_{\Theta} \theta\,p(\theta\mid x)\,d\theta \) Bayes estimator under quadratic loss

    The intuition is straightforward. The posterior distribution describes what values of \(\theta\) remain plausible after observing the data. If squared distance is what we wish to minimize, then the posterior mean is the point that minimizes the average squared distance to all those possible values.

    A useful distinction

    Bayesian updating tells us how uncertainty changes after observing evidence. Bayesian decision theory tells us how to convert that updated uncertainty into an action. The first produces a posterior distribution; the second combines that posterior with a loss function.

    This distinction becomes important later. A Gaussian mixture model does not merely compute probabilities. It uses probabilities as part of an iterative estimation problem in which unknown component memberships and unknown component parameters have to be inferred together.

    The study also discusses the broader statistical idea of Bayes risk: the expected loss associated with a decision rule when uncertainty about the parameter is itself represented probabilistically. Within the decision-theoretic framework adopted in the paper, a Bayes estimator is the estimator chosen because it minimizes the relevant expected loss.

    Probability describes uncertainty; a loss function gives that uncertainty consequences.

    The bridge from inference to decision theory
    · · ·

    04 · Unsupervised learning Why clustering is different from classification

    The paper then changes scale. It moves from the estimation of an unknown parameter toward a broader machine-learning question: how can structure be discovered in a dataset when the observations do not already come with known class labels?

    This is the defining setting of unsupervised learning. In supervised learning, the training data contain an outcome or label that the algorithm is asked to reproduce or predict. A model may learn, for example, whether a loan applicant defaulted, which party a respondent voted for, or what numerical value a dependent variable took.

    In unsupervised learning there is no such answer key. The algorithm receives observations and their measured characteristics, but not a pre-existing declaration that observation 17 belongs to type A while observation 18 belongs to type B. The structure itself must be inferred from patterns in the data.

    Classification versus clustering

    In classification, classes are known during training and the model learns how to assign new observations to them. In clustering, the groups are not given in advance. The method attempts to discover a useful grouping from similarities, differences and distributional structure within the observed data.

    Gómez Julián places cluster analysis at the center of this unsupervised-learning problem. In its most general form, clustering means partitioning observations according to shared characteristics so that observations within a cluster are relatively similar and observations belonging to different clusters are relatively dissimilar.

    That formulation sounds simple, but it conceals one of the deepest difficulties in clustering: there is not always one uniquely obvious way to divide a dataset. The same cloud of points may plausibly be described as containing two broad groups, several narrower subgroups, or a hierarchy in which larger groups contain smaller ones.

    The study uses this ambiguity to distinguish two broad families. Partitional clustering divides observations into non-overlapping groups at a selected level. Hierarchical clustering, by contrast, organizes groups within groups, producing a nested structure that can often be represented as a tree.

    This is more than a technical distinction. It reminds us that a cluster is not simply an object waiting in the data to be photographed. A clustering procedure embodies a definition of what similarity means, how distance is measured, what geometry is permitted, and at what scale differences are considered important.

    To ask how many groups are in a dataset is already to ask what counts as a group.

    Why clustering is fundamentally a modelling problem

    This is particularly important for economists and political scientists. Suppose countries are represented by unemployment, literacy, poverty and public education expenditure. A clustering algorithm may detect statistically distinct configurations of those variables. But the resulting groups should not automatically be treated as substantive political or economic “types.” Statistical grouping is evidence about structure; interpretation still requires theory and knowledge of the phenomenon being studied.

    Before clustering, the paper also emphasizes the importance of data preprocessing. Outliers, different measurement scales, irrelevant variables and missing values can alter the apparent geometry of the dataset. Normalization may therefore matter when distance is central to the algorithm, while variable reduction can be useful when irrelevant dimensions obscure rather than clarify structure.

    · · ·

    05 · A first clustering model What k-means actually assumes

    To understand why Gaussian mixtures are useful, Gómez Julián first introduces one of the best-known clustering algorithms: k-means.

    The basic idea is geometric. Choose a number of groups, \(K\). Associate each group with a center, or centroid. Then assign every observation to the group whose centroid is closest. The centroids are updated from the observations assigned to them, and the process is repeated until the configuration stabilizes.

    \( \displaystyle \min_{C_1,\ldots,C_K} \sum_{k=1}^{K} \sum_{x_i\in C_k} \lVert x_i-\mu_k\rVert^2 \) the familiar k-means objective: minimize within-cluster squared distance from each observation to its centroid

    Even readers who have never implemented the algorithm can visualize its logic. Imagine placing \(K\) pins on a map. Each observation is sent to the closest pin. The pins are then moved to the centers of the observations assigned to them, and the assignment is repeated. Eventually the pins and memberships stop changing substantially.

    This procedure is powerful and computationally convenient, but its simplicity imposes a geometric structure. The distance-to-centroid logic works most naturally when clusters are compact and roughly spherical—or circular when visualized in two dimensions.

    Real datasets need not cooperate. A cluster may be long and narrow, tilted diagonally through the feature space, tightly concentrated in one direction and widely dispersed in another. Two groups can also overlap. Once these possibilities appear, distance to a single center may no longer describe the structure adequately.

    There is another limitation that becomes central to the paper. Ordinary k-means makes what is known as a hard assignment. An observation is placed in cluster 1 or cluster 2 or cluster 3. The algorithm does not naturally say: “there is a 72% probability that this observation belongs to cluster 1 and a 28% probability that it belongs to cluster 2.”

    Two limitations to remember

    The transition from k-means to Gaussian mixtures in the study is motivated by two especially important ideas: cluster geometry and uncertain membership. Gaussian mixtures can model clusters with covariance structure and can assign probabilistic, rather than purely deterministic, membership.

    This prepares the central conceptual turn of the study. Instead of thinking of a cluster merely as a collection of points around a centroid, we can think of it as a probability distribution.

    · · ·

    06 · Latent structure The central idea: a population can be a mixture

    Suppose we observe the distribution of some variable across an entire population. At first glance we see only one dataset. But what if that population is actually composed of several subpopulations generated by different statistical processes?

    This is the fundamental intuition behind a mixture model. The overall probability distribution is represented as a weighted combination of several component distributions.

    \( \displaystyle f(x_i;\Psi) = \sum_{k=1}^{G} \pi_k f_k(x_i;\theta_k) \) finite mixture model · each component has its own parameters and contributes according to its mixture weight

    Here \(G\) is the number of components. \(f_k(x_i;\theta_k)\) is the density of component \(k\), determined by its own parameters \(\theta_k\). The quantity \(\pi_k\) is the component’s mixture weight, satisfying \(\pi_k>0\) and \(\sum_{k=1}^{G}\pi_k=1\).

    The weights are important. If 70 percent of the population appears to have been generated by one component and 30 percent by another, the two component densities should not contribute equally to the overall population density. The mixture weights encode their relative prevalence.

    But the deepest feature of the model is something we do not directly observe: the component identity of each observation. The dataset contains \(x_i\), but it does not ordinarily contain an additional column supplied by nature saying: “this point was generated by Gaussian component 3.”

    Component membership is therefore a latent variable. It is hidden structure inferred from the observed data.

    Observed and latent quantities

    The observations \(x_1,\ldots,x_n\) are visible. The component labels that generated them are not. A finite mixture model therefore links observable data to unobservable group membership, while estimating the parameters and relative weight of each component.

    This is why the paper treats mixture models as naturally connected to hierarchical and latent-variable modelling. There is one level at which an observation belongs to some unobserved component, and another level at which the observed value is generated according to the probability distribution associated with that component.

    If the component distributions are Gaussian, the model becomes a Gaussian mixture model, or GMM:

    \( \displaystyle f(x) = \sum_{k=1}^{G} \pi_k\, \mathcal{N}(x\mid\mu_k,\Sigma_k) \) Gaussian mixture model · each latent group is represented by a Normal density with its own mean and covariance structure

    In one dimension, each component has a mean and a variance. In several dimensions, the mean becomes a vector \(\mu_k\), while dispersion and dependence among variables are represented by the covariance matrix \(\Sigma_k\).

    The covariance matrix is precisely what gives Gaussian mixture models their geometric flexibility. It allows one cluster to be narrow, another broad, another elongated, and another oriented along a diagonal direction in multivariate space.

    The paper therefore presents Gaussian mixtures as a probabilistic generalization of the more rigid centroid-based intuition associated with k-means. Instead of asking only which center is closest, the model asks a richer question: given the estimated distributions, how probable is it that this observation came from each component?

    \( \displaystyle P(Z_i=k\mid x_i) = \frac{ \pi_k\, \mathcal{N}(x_i\mid\mu_k,\Sigma_k) }{ \sum_{j=1}^{G} \pi_j\, \mathcal{N}(x_i\mid\mu_j,\Sigma_j) } \) posterior probability that observation i belongs to component k

    Now the earlier discussion of Bayes becomes visibly relevant. The model begins with component weights and component densities and, conditional on an observed point, calculates updated probabilities of component membership.

    An observation near the center of one component and far from all others may receive an assignment probability close to one. An observation lying in an overlapping region may receive substantial probability under two or more components. This is soft classification: the uncertainty surrounding membership is retained instead of being immediately discarded.

    A Gaussian mixture does not merely divide the data. It proposes a probabilistic account of how several hidden subpopulations could have generated the observed population.

    The core modelling idea of the 2020 study

    But we have now reached an apparent circularity. To estimate the mean, covariance and weight of each component, we would like to know which observations belong to which component. Yet determining which observations belong to which component is precisely what requires knowing those means, covariances and weights.

    Solving that circular problem is the task of one of the most important algorithms in latent-variable statistics: expectation-maximization.

    07 · Hidden information EM: learning when group membership is unknown

    The difficulty facing a Gaussian mixture model can now be stated precisely. We observe the data points, but we do not observe the component from which each point was generated. If those memberships were known, estimating the parameters of each Gaussian component would be relatively straightforward. But the memberships themselves depend on parameters that are still unknown.

    Gómez Julián’s 2020 study approaches this problem through the classical expectation-maximization algorithm, or EM, developed by Dempster, Laird and Rubin. The broader setting is estimation from incomplete data: there is information that would make the estimation problem easier, but that information is not directly observed.

    In mixture modelling, the missing piece is especially intuitive. Imagine that each row of the dataset secretly carries an additional variable saying which Gaussian component generated it. If that hidden variable were visible, we would have what can be thought of as the complete data. In reality, only the measured variables are observed.

    The missing-data interpretation

    In a Gaussian mixture model, the data point itself is observed, but its generating component is latent. EM treats this hidden information as the missing part of an otherwise more convenient statistical problem.

    The ingenious feature of EM is that it does not demand that this missing information somehow become directly observable. Instead, it alternates between two calculations. Each calculation makes the other possible.

    The E-step: estimate the hidden memberships

    Begin with some current values for the component parameters: the means, covariance matrices and mixture weights. Given those values, calculate how probable it is that each observation belongs to each component.

    These probabilities are often called responsibilities. Component \(k\) takes responsibility for observation \(i\) in proportion to how plausible that observation is under the component’s Gaussian density and how prevalent that component is in the mixture.

    \( \displaystyle \gamma_{ik} = P(Z_i=k\mid x_i,\Psi) = \frac{ \pi_k\, \mathcal{N}(x_i\mid\mu_k,\Sigma_k) }{ \sum_{j=1}^{G} \pi_j\, \mathcal{N}(x_i\mid\mu_j,\Sigma_j) } \) E-step intuition · estimate the probability that each observation belongs to each Gaussian component

    Notice what has happened. The hard, unknown statement “observation \(i\) belongs to cluster \(k\)” has been replaced with a set of probabilities. An observation may be overwhelmingly associated with one component, or it may sit in an overlapping region and divide its probability between several components.

    The M-step: update the model

    Once those expected memberships have been calculated, the algorithm turns the problem around. It now treats the probabilistic memberships produced by the E-step as information for re-estimating the parameters.

    The means, covariance matrices and mixture weights are updated so that the likelihood of the observed data increases under the newly estimated mixture.

    One EM cycle

    E-step: using the current model parameters, estimate the hidden component memberships.

    M-step: using those estimated memberships, re-estimate the model parameters by maximizing the relevant likelihood criterion.

    Then the algorithm goes back to the E-step. The new parameters imply new membership probabilities; those new probabilities imply new parameter estimates; and the process continues iteratively.

    \( \Psi^{(0)} \rightarrow \text{E-step} \rightarrow \text{M-step} \rightarrow \Psi^{(1)} \rightarrow \text{E-step} \rightarrow \text{M-step} \rightarrow \cdots \) expectation and maximization alternate until the fitted solution stabilizes

    The process stops when the parameter estimates—or equivalently the likelihood improvements—change so little that the algorithm is considered to have reached convergence.

    This makes EM a particularly elegant response to the apparent circularity encountered at the end of the previous section. We needed cluster membership to estimate the distributions, but we needed the distributions to estimate cluster membership. EM solves the problem by alternating between the two conditional tasks.

    Estimate what is hidden using the current model; then improve the model using what you have just estimated.

    The iterative logic of expectation-maximization

    There is an important qualification. EM is an optimization algorithm, not a magical guarantee that every possible starting point will lead to the globally best solution. Mixture-model likelihoods can contain multiple local optima. Initialization and model specification can therefore matter. What EM guarantees at the operational level is an iterative procedure for improving the likelihood until a stationary solution is reached.

    · · ·

    08 · Statistical geometry Why Gaussian mixtures can see ellipses

    The next step in Gómez Julián’s argument is geometric. In one dimension a Gaussian distribution is described by a mean and a variance. Move into two or more dimensions, however, and variance is no longer sufficient. The relationships among variables must also be represented.

    This is the role of the covariance matrix, \(\Sigma_k\). For component \(k\), the mean vector \(\mu_k\) determines its center, while \(\Sigma_k\) determines how the probability mass spreads through multivariate space.

    \( X\mid Z=k \sim \mathcal{N}(\mu_k,\Sigma_k) \) each Gaussian component possesses its own center and covariance geometry

    In two dimensions, contours of equal Gaussian density form ellipses. This provides an intuitive way to read covariance. A nearly circular ellipse indicates similar dispersion in different directions. An elongated ellipse indicates much greater variation along one direction than another. A tilted ellipse signals covariance between the variables.

    This is precisely where Gaussian mixture clustering becomes more flexible than the elementary geometric picture supplied by k-means. A centroid alone tells us where a cluster is centered. A covariance matrix also tells us its volume, shape and orientation.

    Think geometrically

    Two clusters may have centers that are equally far apart while still being statistically very different. One may be compact and almost circular; another may be broad and strongly elongated. Gaussian mixture models can represent this difference because the covariance matrix is part of the model.

    The mclust framework studied in the paper exploits this fact systematically. Instead of fitting only one possible covariance structure, it considers a family of Gaussian models obtained by placing different restrictions on the volume, shape and orientation of the component ellipsoids.

    Gómez Julián discusses the 14 multivariate Gaussian models available in the version of mclust studied in the 2020 research. Their compact names—such as EEE, VEV, VVI or EEV—encode restrictions on those geometric properties.

    Example Geometric idea
    EEE Equal volume, equal shape and equal orientation across components
    VEV Variable volume, equal shape and variable orientation
    VVI Diagonal covariance structure with variable volume and shape
    EEV Equal volume and shape, with orientation allowed to vary

    These codes are not decorative software jargon. They describe competing statistical hypotheses about the geometry of the hidden groups. Should all clusters have the same spread? May one be larger than another? Must their ellipses point in the same direction? Is a diagonal covariance matrix enough, or does the data require rotated ellipsoids?

    Seen this way, model-based clustering is doing more than deciding where to draw boundaries. It is comparing alternative generative descriptions of the data.

    In model-based clustering, the shape of a cluster is not an afterthought. It is part of the hypothesis being estimated.

    Covariance as statistical geometry
    · · ·

    09 · Model selection BIC and the problem of choosing a model

    Gaussian mixtures create a new problem precisely because they are flexible. We may fit different numbers of components, and for each number of components we may consider different covariance structures. Which model should be preferred?

    Maximized likelihood alone is not enough. Adding parameters usually gives a model more freedom to accommodate the observed data, so raw fit can improve simply because the model has become more complicated. If complexity is never penalized, the procedure is pushed toward increasingly elaborate specifications.

    This motivates the Bayesian Information Criterion, introduced by Gideon Schwarz and discussed at length in the 2020 study. In one common notation,

    \( \mathrm{BIC} = -2\log \hat{L} + k\log n \) conventional minimization form · fit is balanced against a penalty that increases with model complexity

    Here \(\hat{L}\) is the maximized likelihood, \(k\) is the number of estimated parameters and \(n\) is the sample size. The first term rewards fit; the second penalizes additional parameters.

    An equivalent sign convention is often written so that larger values are preferred:

    \( \displaystyle \log \hat{L} – \frac{k}{2}\log n \) Schwarz’s maximization form · the same fit-versus-complexity logic expressed with the opposite orientation
    A practical warning about signs

    Readers sometimes see “choose the smallest BIC” in textbooks and then encounter mclust output where the preferred model has the largest BIC value. This is a matter of convention. The criterion can be written with opposite signs. What matters is using the convention adopted by the software or source consistently.

    In mclust, BIC therefore becomes the mechanism for comparing combinations of component number and covariance parametrization. The software can fit a collection of candidate Gaussian mixture models and compare them rather than forcing the researcher to stipulate one geometry in advance.

    Conceptually, this is a competition among explanations. A one-component model says that a single Gaussian population is sufficient. A two-component model says that two latent subpopulations provide a better account after accounting for the additional parameters. A five-component model makes an even more elaborate claim. BIC asks whether the gain in likelihood is large enough to justify that extra complexity.

    What BIC is doing in this paper

    BIC acts as a bridge between estimation and model selection. EM estimates the parameters of a candidate Gaussian mixture. BIC helps decide which candidate structure—among different numbers and geometries of components—is comparatively preferable.

    This distinction is essential. EM does not, by itself, answer every modelling question. Given a specified mixture structure, it provides a way to estimate its parameters. Model selection operates at another level: it compares alternative structures.

    The result is a layered procedure. First define candidate probability models. Then estimate them. Then compare them. Finally inspect the resulting classification and ask whether the statistical structure is substantively meaningful.

    \( \text{candidate models} \rightarrow \text{EM estimation} \rightarrow \text{BIC comparison} \rightarrow \text{selected clustering structure} \) the model-based clustering workflow developed toward the applied section of the study

    We are now ready for the final step of the 2020 investigation: seeing what this machinery actually produces in R. Gómez Julián closes the substantive analysis with two types of application. The first uses the canonical Iris dataset; the second moves into social and economic data from the World Bank, combining indicators of education expenditure, literacy, unemployment and poverty.

    10 · Applied examples in R From Iris flowers to World Bank indicators

    After more than one hundred pages of theoretical preparation, Gómez Julián’s 2020 study finally lets the statistical machinery run. This last substantive section is useful precisely because the preceding discussion changes the meaning of what would otherwise look like a few lines of R code. By this point, a call to Mclust() is no longer merely a software command. It invokes finite Gaussian mixtures, latent membership, maximum-likelihood estimation through EM, alternative covariance geometries and BIC-based model comparison.

    The paper provides two kinds of illustration. First comes the canonical Iris dataset distributed with R. Then the analysis moves to a dataset assembled from World Bank indicators, bringing the method into a setting much closer to economics, political science and public-policy research.

    How to read the output

    When Mclust() reports a model such as VEV, EEE or VVI, it is describing the covariance structure selected for the Gaussian components. When it reports a number of components, it is describing the number of mixture components preferred by the model-selection procedure among the candidates fitted.

    The Iris example

    The first application uses the four familiar quantitative variables in the Iris dataset. Gómez Julián runs:

    mod1 <- Mclust(iris[,1:4])
    summary(mod1)
    Gaussian model-based clustering of the four measured Iris variables

    The reported solution is a VEV Gaussian finite mixture with two components. The 150 observations are partitioned into clusters containing 50 and 100 observations, respectively. The reported log-likelihood is \(-215.726\), while the output gives a BIC of \(-561.7285\) and an ICL of \(-561.7289\).

    Dataset Selected structure Clustering
    Iris VEV · 2 Gaussian components 50 / 100 observations

    The accompanying plots make visible the two layers of the procedure. One panel compares BIC values across candidate covariance models and different numbers of components. Another displays the resulting classification across pairs of the measured variables. The graph is therefore not merely showing clusters after the fact: it also gives the reader a view of the model-selection problem that produced them.

    This example is deliberately straightforward. Its role is to show that the theoretical discussion of mixture densities, EM, covariance structure and BIC can be condensed operationally into a remarkably short piece of R code.

    · · ·

    A social-science example using World Bank data

    The second application is more directly connected to the concerns of economists and policy researchers. Gómez Julián constructs an example using World Bank data and four variables for 2018:

    Variables used in the 2020 application

    Public expenditure on education as a percentage of GDP; unemployment as a percentage of the total labour force; poverty incidence according to the national poverty line; and the adult literacy rate for persons aged 15 and above.

    The R workflow imports the separate datasets, selects the 2018 observations, joins them by country, removes rows for which the required combination contains missing values, and then applies Mclust() to the resulting numerical variables.

    When all four indicators are considered together, only 13 complete observations remain in the dataset used by the code. The reported model is EEE with nine components: ellipsoidal Gaussian clusters with equal volume, equal shape and equal orientation.

    \( n=13,\qquad G=9,\qquad \text{model}=\mathrm{EEE} \) four-variable World Bank example reported in the study

    The output reports a log-likelihood of \(-77.62892\), 54 degrees of freedom, BIC \(-293.7651\) and ICL \(-293.7727\). The component counts are extremely small: the nine clusters contain respectively 1, 2, 2, 1, 1, 1, 2, 2 and 1 observations.

    Gómez Julián then repeats the model-based clustering exercise using smaller combinations of variables. This is particularly revealing because the available sample size changes sharply depending on which World Bank indicators must be simultaneously observed.

    Variables n Model Components BIC
    Education expenditure + literacy 39 XXI 1 −475.3468
    Education expenditure + unemployment 71 VVI 2 −679.1997
    Education expenditure + poverty 17 EEV 5 −193.6785

    The contrast is striking. For public education expenditure and adult literacy, the fitted solution contains only one component among 39 complete observations. For education expenditure and unemployment, the selected model is VVI with two components, containing 49 and 22 observations. For education expenditure and poverty, the result is an EEV specification with five components, whose sizes are 3, 3, 4, 4 and 3.

    These examples demonstrate something that can easily disappear when one speaks abstractly about “the number of clusters.” The number of components is not an intrinsic number attached forever to a set of countries. It depends on the variables being modelled, the available observations, the candidate covariance structures and the statistical criterion used to compare those models.

    An important inferential boundary

    The output reported in this section is cluster analysis. It describes statistical structure found by the fitted Gaussian mixture models. By itself, such an exercise does not establish that education expenditure causes unemployment, literacy or poverty, nor does it estimate the magnitude of a causal effect. Those would be different inferential questions requiring a different research design.

    This distinction is especially valuable for policy analysis. A cluster can reveal that some countries occupy similar regions of a multivariate statistical space. That may motivate substantive investigation. It does not, by itself, explain historically or causally why those countries occupy that region.

    The four-variable result deserves similar care. Nine components from only thirteen complete observations is exactly the kind of output that should be read together with the sample size and model complexity rather than reduced to the phrase “nine types of countries.” The study reports the statistical fit; substantive interpretation requires returning from the model to the empirical object being studied.

    · · ·

    11 · The larger lesson What the 2020 study is really teaching

    The most important feature of Gómez Julián’s investigation may be its refusal to begin with the software. The paper could have been a short tutorial showing how to call Mclust(), inspect BIC and plot a classification. Instead, it constructs a long conceptual route from probability and estimation to the final clustering output.

    That route matters because the elements are genuinely connected. Bayesian reasoning introduces conditional probability and updating. Decision theory explains how probability distributions can be connected to estimators and loss. Unsupervised learning introduces the problem of discovering structure without known labels. Cluster analysis gives that problem a statistical form. Mixture models reinterpret an apparently homogeneous population as the superposition of latent subpopulations. Gaussian mixtures give those subpopulations flexible probabilistic geometry. EM estimates models whose membership information is hidden. And BIC provides a way to compare competing specifications.

    \( \text{Bayes} \rightarrow \text{estimation} \rightarrow \text{latent variables} \rightarrow \text{mixtures} \rightarrow \text{EM} \rightarrow \text{BIC} \rightarrow \text{clustering} \) a compressed map of the conceptual route reconstructed in the 2020 study

    One can also read the paper as an argument for understanding statistical methods structurally. A model is not simply an equation. It includes assumptions about what is observable, what is latent, which probability family describes the data, how parameters are estimated, what geometries are permitted and how competing specifications are compared.

    Gaussian mixture models make that point unusually visible. The observable cloud of data is only the surface. Beneath it lies a proposed generative structure: component distributions, latent memberships, mixture weights, means and covariance matrices. The analyst does not observe this machinery directly. It is inferred.

    The visible dataset is the starting point. The statistical model is a hypothesis about the hidden structure capable of producing it.

    A central intuition running through the study

    This is also why the difference between hard and soft classification is so significant. Saying that a country, person or flower belongs to “cluster 2” suppresses information. A Gaussian mixture can instead preserve the fact that an observation may lie near the frontier between several plausible components. Probability makes ambiguity measurable.

    Likewise, covariance transforms clustering from the simple idea of distance from a center into a richer account of statistical geometry. Groups may differ not only in location but also in dispersion, shape and orientation. And BIC reminds us that greater flexibility comes at a cost: a model must earn its additional complexity through improved fit.

    For economists, econometricians and political scientists, perhaps the most transferable lesson is therefore methodological. If a population may contain qualitatively different statistical regimes, forcing every observation into a single homogeneous distribution can conceal structure. Mixture models offer one formal way of asking whether the aggregate pattern may instead be generated by several latent components.

    But the converse warning is equally important. Discovering a statistically preferred partition does not relieve the researcher of the obligation to understand the real phenomenon. A component is a component of a statistical model. Whether it corresponds to a meaningful social class, institutional regime, developmental configuration, biological population or merely a feature of the available sample must be established with substantive knowledge and further evidence.

    In one sentence

    Gómez Julián’s 2020 investigation is a theoretical guided tour of the ideas required to understand how Gaussian finite mixture models can discover latent structure in unlabeled data, how EM estimates that structure, and how model-selection criteria such as BIC help decide which probabilistic representation to retain.

    · · ·

    12 · Closing perspective Statistics before software

    There is a useful reversal at the heart of this study. Modern statistical computing encourages us to begin with a function and discover afterward what it does. Gómez Julián’s 2020 text proceeds in the opposite direction: first reconstruct the mathematical and statistical concepts, then approach the function.

    That choice makes the paper unusually broad. Posterior probability, Bayes estimators, loss functions, BIC, data mining, machine learning, clustering, density estimation, vector quantization, k-means, EM, categorical and Dirichlet distributions, optimization, parameters, hyperparameters and Gaussian finite mixtures all appear because the final procedure stands at the intersection of those ideas.

    For a technically trained reader, the value of this route is that it exposes the architecture hidden beneath a familiar R command. For a reader from political science, economics or another applied field, it offers something equally useful: an intuitive path into a method that otherwise arrives wrapped in matrix algebra and probability notation.

    And the practical lesson is simple. When an algorithm reports that the data contain one group, two groups or five, the interesting question is not merely what did the software return? It is: what statistical model made that answer possible, what assumptions gave the groups their shape, and what kind of statement about reality is the result actually capable of supporting?

    Good statistical practice begins where the automatic output ends: with an attempt to understand what the model has actually measured.

    Final reflection on Gómez Julián’s 2020 study
    José Mauricio Gómez Julián · 2020
    On Bayes Estimators, Cluster Analysis and Gaussian Mixtures: A General Theoretical Analysis of densityMclust in R and Statistical Theory
  • The Shape of a Crisis: A General Theory of Capitalist Cycles

    The Shape of a Crisis: A General Theory of Capitalist Cycles

    Thesis Release · Political Economy

    The Shape of a Crisis

    A general theory of the cycles of the dynamics of the capitalist system in the long run — now available in English

    Every few years the same story is told twice. First, that the economy has entered a new era in which the old rules no longer apply. Then, some months later, that what happened was an accident: a shock, a bubble, a virus, a war. Both tellings share a premise so quiet that it is rarely examined — that the rise and the fall are separate events, and that a good theory of the good years need not be a theory of the bad ones.

    The thesis released today argues the opposite, and then goes to some length to measure it. The boom and the crisis are not two phenomena but two moments of one: the crisis of overproduction is the mechanism by which capitalism restores the conditions of an accumulation that its own success had eroded. Devaluation clears the field; new methods of production are introduced under duress; profitability recovers on the ruins. The recovery is not the negation of the crisis. It is its product.

    That claim is old. What is new here is the attempt to make it decidable — to state it in a form that quarterly data on the United States economy between 1992 and 2024 could have contradicted, and then to check whether they do.

    Three questions, and why the order matters

    The investigation is organised around one general objective — to analyse the long-run cyclical behaviour of U.S. capitalism in the light of the dominant economic theories — and three specific ones, asked strictly in this order:

    • Which theory explains and predicts best? Not which is most elegant, or most widely taught, but which survives being pointed at the data.
    • Which factors generate the cycle? Economic and extra-economic alike — the thesis refuses in advance to treat wars and monetary policy as noise sitting outside a clean economic mechanism.
    • By which rules do those factors interact? A list of causes is not a theory. The theory is in the grammar that binds them.

    The order is not decorative. A great deal of applied economics answers the third question with machinery borrowed from a theory it never subjected to the first. Here the selection of the framework is itself a result, defended before it is used.

    Five families of an old argument

    Before measuring anything, the thesis maps the terrain. Economic thought on the cycle is sorted into five groups: the pre-Kondratieff non-heterodox schools; the Kondratieff school; the post-Kondratieff marginalist and neoclassical schools; the heterodox schools; and the historiographic vision of long waves, which reads the cycle through the archives rather than through the equations.

    With that map in hand, three long-running disputes are adjudicated rather than summarised. Does the crisis originate in overproduction or in underconsumption? Is a sustained expansion of credit a symptom of recovery, or of the exhaustion of the conditions that made recovery possible? Is there really an inverse relation between inflation and unemployment, or is the appearance of one an artefact of the precariousness of the labour market? Each is answered, and each answer carries consequences later, when the model is specified.

    A framework that states its own conditions of failure

    A substantial part of the theoretical apparatus is devoted to a materialist characterization of the dialectical method: its fundamental categories, a Marxist ontology built from a metalogical gnoseology, and an explicit treatment of verification, falsification and decidability. The purpose is unglamorous and indispensable — to fix, in advance, which propositions of the theory are empirically decidable and which are interpretive. Without that boundary, no amount of subsequent statistics can tell you what has been tested.

    Ten dials, seven of them internal

    The empirical core is a Bayesian generalized linear model of the growth of U.S. real output, estimated with Hamiltonian Monte Carlo and cross-validated against machine-learning and deep-learning competitors. It retains thirteen coefficients across ten factors. Seven are economic:

    FactorWhat it registers
    Net Average Rate of Profit (ARoP)The central variable of the accumulation process, and the one whose long-run tendency the theory predicts.
    Elasticity of the gross rate of surplus value to the average organic composition of capitalHow the exploitation of labour power responds when the technical structure of capital changes.
    Non-residential fixed investmentThe pace of accumulation in the productive sector; the hinge between boom and crisis.
    Inventory-to-sales ratioThe gap between producing value and realising it on the market.
    S&P 500Financialization, entering through a natural cubic spline with three degrees of freedom.
    Non-financial private sector creditThe credit system as the accelerator and the brake, splined with two degrees of freedom.
    Capitalist R&D spendingThe innovative impulse; the second largest coefficient in the model.

    And three are extra-economic: military spending (splined with three degrees of freedom), the federal surplus or deficit, and the effective federal funds rate. Their presence is not a concession to realism. It follows from the argument that an imperial economy counteracts the tendency of its own profit rate to fall by means that are not internal to its national accounts.

    The Average Rate of Profit carries the fourth largest coefficient of the thirteen — behind only the intercept, R&D spending, and one basis function of the splined S&P 500. The conclusion the author draws from its behaviour is worth quoting in substance: what is favourable to the global process of capital accumulation is not thereby favourable to the dynamics of aggregate growth. The two are not the same quantity, and treating them as one is precisely the confusion the cycle punishes.

    Note, too, what the splines are doing. Three of the ten factors would not sit still in a straight line. That is not a technical footnote: it is the first quantitative sign that the interaction of these factors involves thresholds and turning points rather than a stable proportionality.

    Not random. Chaotic.

    “Unpredictable” and “random” are not synonyms, and the difference decides what kind of science economics can be. A random system has no internal structure to find. A chaotic one is rigidly determined and still unpredictable at long horizons, because arbitrarily small differences in initial conditions grow exponentially apart.

    Three measurements place the U.S. economy in the second category. The Lyapunov exponent is positive (approximately $0.0515$): small perturbations amplify rather than dissipate. The correlation dimension is not an integer ($3.32798$): the attractor reconstructed by Takens’ theorem has a fractal structure, patterns repeating across scales of time and magnitude — which is what “cyclical, but not periodic” means when it is stated precisely. And recurrence quantification finds high determinism alongside variability in laminarity and in the maximum diagonal line length: underlying deterministic structures that themselves evolve.

    $\lambda > 0 \quad\text{with}\quad D_2 = 3.32798 \notin \mathbb{Z}$

    Read together, these say something a forecaster should find sobering and a theorist should find encouraging. The long-horizon forecast is not merely hard; it is structurally bounded. But the structure that bounds it is real, stable and measurable — which is exactly what a theory of the cycle needs to have something to explain.

    The shape of time

    The most unusual instrument in the thesis is topological. The idea is to stop asking how big the numbers are and start asking which observations can see which. Convert the series into a directed visibility graph — a link from one quarter to another when the second is visible from the first over the intervening data — and study the order structure that results.

    Two topologies are built on it, and they disagree in an informative way.

    • The coarser Alexandrov topology, built on temporal reachability, turns out to be connected. At the level of its order structure the economy is globally a single piece: every observation is bound to every other by chains of temporal visibility. There is no quarter that stands apart.
    • The finer Nada topology is locally fragmented — six components under the natural visibility graph, thirty-six under the horizontal one. Zoom in, and the fabric shows seams: structural discontinuities at the level of closed neighbourhoods.

    Global unity and local rupture at once. That duality is not a contradiction to be resolved; it is the object being described. And a third measurement gives the whole thing a direction: the bitopological analysis yields $D = +4$, meaning that expansions generate more temporal visibility than contractions. The cycle is not symmetric in time. Growth accumulates gradually and in view; collapse happens abruptly and blind. Run the film backwards and it is recognisably the wrong film.

    ⚠️ Why you must not “clean” the crises

    There is a habit in applied work of treating extreme values as contamination and smoothing them away by discontinuous imputation. Here that habit is shown to be a category error with a measurable price. The extreme fluctuations of the 2020 crisis belong to a connected block even under the finer topology; severing them is a topological rupture, not a cleaning operation. The thesis reports the consequence directly: models fitted after such imputation performed worse, because one was using predictors suited to one phenomenon — real output growth — to predict a qualitatively different one: real output growth after the crisis had been removed from it. The crises are not noise around the cycle. They are the cycle.

    The grammar of the cycle

    The third question receives a seven-part answer. The factors interact through feedback (the rate of profit shapes investment, investment shapes the organic composition of capital, which feeds back into the rate of profit); time lags (R&D and fixed investment pay out on a delay, and the delay is itself cycle-generating); non-linearity (thresholds and regime changes, which is why three factors needed splines); deterministic chaos; sectoral interdependence between the department producing means of production and the one producing means of consumption; topological structure, global connectedness with local fragmentation; and the influence of the global context, which is how military spending and the S&P 500 enter a nominally domestic account.

    The unifying claim is that each phase of the cycle contains the seed of its own negation. New methods of production introduced during the crisis lay the foundations of the next boom; the overaccumulation of the boom prepares the ground for the next crisis. Innovation initially arrests the fall of the profit rate and ultimately deepens it — through the way the degree of exploitation of labour power responds, over time, to the very methods introduced to raise it.

    What a cycle is for

    The thesis closes on a question most treatments never pose. If the cycle is a mechanism, what does it accomplish? Two answers, at different depths. Its intermediate practical end is to restart the process of capital accumulation once instability has reached a critical level — this the mechanism achieves, repeatedly, at a cost borne unevenly. Its definitive practical end is to lay the material and spiritual conditions for a reorganization of the fundamental productive structure of society, one capable of a stability beyond what the capitalist mode of production can reach within its own limits.

    What this establishes, and what it does not

    The evidence supports the claim that classical Marxist economic theory possesses the greatest explanatory and predictive capacity for long-run cycles among the theories examined here, on this economy, over this period. It is a comparative result on the United States between 1992 and 2024, quarterly — not a universal proof, and not a forecast. The thesis is explicit about the cost of its own data: the Average Rate of Profit and the average rate of surplus value were available only annually through 2020, and completing the series to 2024 required temporal disaggregation and prediction, which puts a wider band of uncertainty around the most recent quarters. The philosophical, historical, conceptual and statistical scope of each result is distinguished in the text, and results unfavourable to the hypotheses are reported alongside the favourable ones.

    About this edition

    This is the English edition of a thesis originally written in Spanish and submitted to the Universidad Latina de Costa Rica for the degree of Licentiate in Economics. It is interdisciplinary by construction, drawing on Marxist political economy, dialectical and historical materialism, the history and historiography of economic thought, the philosophy and methodology of science, econometrics, Bayesian statistics, the theory of complex systems and topology.

    The edition carries a Note on the Translation that fixes the rendering of the terms whose Spanish usage is technical and not interchangeable with their nearest English cognates — gnoseology, sublation, long wave, solvent demand, technique — and records the editions from which quotations are taken, including the two distinct English and Spanish editions of the Soviet philosophical dictionary, which are cited under different transliterations because they are different books with different pagination.

  • SOME REFLECTIONS ON MARX’S PRICES OF PRODUCTION

    SOME REFLECTIONS ON MARX’S PRICES OF PRODUCTION

    Was Marx Wrong About Prices of Production? — A 260-Page Investigation Says No

    Political Economy • Econometrics • Marx

    Was Marx Wrong About Prices of Production?
    A 260-Page Investigation Says No.

    How one researcher spent years showing that the most famous critique of Marx’s economics rests on a mistake Marx never made.

    Based on: Gómez Julián (2026), “Some Reflections on Marx’s Prices of Production” — Introduction, Conclusions & the Formal-Empirical Chapter · DOI 10.5281/zenodo.21842251

    A Fatal Flaw, or a Fatal Misreading?

    For over a century, a single mathematical argument has been wielded as the definitive proof that Karl Marx’s economics doesn’t work. It goes like this: Marx claimed that the value of goods is determined by the labor that produces them, and that market prices eventually gravitate toward “prices of production” — modified versions of those labor values, adjusted for how capital-intensive each industry is. But when you try to verify this with a system of simultaneous equations, the numbers don’t add up. The sums of values don’t equal the sums of prices. The theory, critics have said since the early 1900s, contains a fatal algebraic error.

    This paper — spanning 260 pages and drawing on philosophy, history, sociology, and statistics — argues that the error was never Marx’s. It was the error of the people who checked his math using a method he never used.

    The Photograph vs. the Movie

    Imagine you’re trying to understand a river. You could take a photograph of it — capturing one frozen moment — or you could film it as a movie, watching how the water flows over time. For over a hundred years, the economists who criticized Marx took a photograph of his theory and then complained that it didn’t look like a movie.

    Here’s the specific issue. Marx described a two-step process: first, a general rate of profit forms across the entire economy; then, each industry’s price deviates from its pure labor value according to how much capital it ties up relative to the average. The standard critique — originating with Ladislaus von Bortkiewicz in 1907 and repeated ever since — takes all of Marx’s accounting identities and solves them simultaneously, as if input prices and output prices were determined at the same instant. Under that framework, Marx’s three aggregate equalities cannot all hold at once.

    The “inconsistency” that has been attributed to Marx for over a century is the inconsistency of the simultaneous-dualist framework that was imposed on him, and it dissolves as soon as time is restored. — Gómez Julián, summarizing the central thesis

    But here’s the catch: solving everything simultaneously is equivalent to assuming that the economy is a photograph — that there is no time. And Marx’s entire framework is built on the opposite premise: that the economy is a process, an unfolding sequence in which the prices that exit one period become the input prices that enter the next. Once you restore that temporal dimension, the “inconsistency” vanishes. The three equalities hold simultaneously — not because Marx was secretly consistent in some miraculous way, but because the contradiction was an artifact of the framework imposed on him, not of his own logic.

    The paper calls the simultaneous approach “Walrasian Marxism” — a phrase that captures the irony: economists imported the logic of Léon Walras’s general equilibrium theory and used it to read Marx, then blamed Marx when the result didn’t work.

    In Plain Language

    Marx was accused for over a century of getting the arithmetic wrong. What actually happened is that someone redid his arithmetic under an assumption he never made — that the prices of things you buy to produce and the prices of things that come out of production are the same prices, set at the same time. If you assume that, Marx’s accounts don’t close. But that assumption is equivalent to saying the economy doesn’t happen in time.

    But Was the Movie Real?

    Pointing out that Marx’s logic works when you read it correctly is necessary but not sufficient. The “temporalist” school has been making this argument for nearly fifty years. But the author noticed a critical gap: nobody in that school had ever taken real-world data and actually estimated the three types of prices Marx described — direct labor values, prices of production, and market prices — and then tested whether market prices actually gravitate toward prices of production as the theory predicts.

    This matters because, as the paper puts it, leaving the correct reading of Marx “in the territory of conceptual argumentation while the incorrect reading occupies alone the territory of measurement” is a strategic vulnerability. If you can’t show that real prices behave the way your theory says they should, your theory remains a philosophical argument, however internally consistent.

    But before presenting any numbers, the paper devotes substantial space to establishing that the process Marx described actually happened in history. This is not an appendix; it’s a foundational part of the argument.

    Before Capitalism

    In pre-capitalist societies, exchange was regulated by labor time — not because someone enforced a theory, but because the material conditions made it so. Barter was dominant, inflation did not exist, and prices could only reflect production costs given available technology. Evidence from anthropology (Malinowski’s Trobriand Islands studies), sociology (Mauss on gift exchange), accounting history (Kula’s analysis of feudal estate records), and even paleogenomics all converge: objects were valued in proportion to the labor they embodied.

    The Transition

    The dissolution of feudal relations, the monetization of exchange, and the destruction of pre-industrial normative frameworks created the conditions for capital to move freely between industries. Thompson’s work on the “moral economy” documents how the new free-market ideology had to be violently imposed, destroying customary protections and creating an unprecedented relationship of exploitation.

    Capitalism Established

    Once barriers to capital movement were destroyed, capital flowed from commerce to industry chasing higher profits, and generalized competition forced a redistribution of total surplus value across sectors. The crisis of 1873 — which destroyed nearly half the blast furnaces in major iron-producing countries — is presented as concrete evidence of the mechanism: firms whose costs were still based on older, individually more labor-intensive methods went bankrupt when they couldn’t compete with prices of production dictated by modern technology.

    In Plain Language

    Prices of production didn’t appear the day someone wrote an equation. They appeared the day capital could freely move from one industry to another chasing the highest profit — which didn’t happen until legal, moral, and political barriers were destroyed. Before that, things were exchanged roughly according to the labor they cost, and there is more than enough evidence — ethnographic, accounting, archaeological, and genetic — to show it.

    What Is a Production Price, Exactly?

    This is where the paper moves into its most technically original territory. The author carefully separates two things that must not be confused:

    What a production price is (the explanandum): it is the expected value, over the distribution of economic perturbations, of the long-run time average of market prices. In plain language: it’s the center of gravity around which actual market prices keep spinning. Not the price they arrive at and stay at (that would be equilibrium), but the average around which they never stop oscillating.

    Key Concept

    The production price is neither an eternal, timeless equilibrium (the error of the simultaneous approach and of Walrasian economics, which takes the law as such for the whole and eliminates time) nor a chaos of prices without law (the error of empiricism, which stays at the level of individual prices and loses the law). It is the law of the whole realizing itself through the contingency of the parts.

    How each step of the process works (the explanans): a rule that determines this year’s market price from last year’s market price and last year’s latent production price, and nothing else. This is modeled as a hierarchical Ornstein-Uhlenbeck process — a three-level cascade in which the production price is itself a latent state with its own dynamic gravitating toward value, and market prices gravitate toward that latent state rather than toward a fixed, noisy index.

    The uncertainty is built into the model explicitly: uncertainty in the average rate of profit, uncertainty in the advanced capital, uncertainty in the disaggregation of national accounts into 37 sectors (handled through multiple imputation with 25 imputations combined by Rubin’s rule), and parametric uncertainty estimated through Bayesian Markov Chain Monte Carlo methods.

    One crucial point: no magnitude is obtained by solving a simultaneous system. Value is constructed empirically and directly as $V = c + v + p$ (cost plus surplus value), and production price as $\Phi = c + K \cdot G’$ (cost plus capital times the general rate of profit). There is no Leontief inversion, no simultaneous algebra, anywhere in the construction.

    In Plain Language

    Think of a production price as the “gravitational center” of a spinning object. The object (a market price) never stops moving — it wobbles, it swings, it drifts — but over time its average position is pulled toward that center. The math describes both what the center is and how each wobble happens, and it does so while honestly accounting for all the uncertainty in the measurement.

    The Defining Equations: (9) Through (11)

    Here is where the metaphor turns into mathematics. The paper writes the definition of a production price in three successive steps — each one making explicit an assumption the previous step left implicit — numbered (9), (10), and (11) in the original text. None of the three generates a trajectory by itself; together they define the explanandum — what the object is — that the cascade below then generates.

    Equation 9 — What a Production Price Is
    $$ \lim_{t\to\infty} E\!\left[\varphi^i_t\right] \;=\; k^i_t + K^i_t\, E\!\left[G'(t,X)\right] \;=\; \Phi^i_t $$

    Here $\varphi^i_t$ is sector i’s market price at time $t$, $k^i_t$ is its cost price (constant capital consumed plus variable capital), $K^i_t$ is the total capital advanced, and $G'(t,X)$ is the general rate of profit — itself a stochastic process indexed by a perturbation $X$ that bundles the exodus of capital between branches and technological innovation.

    In words: a production price is the long-run limit of the average market price. Not the price itself at any instant — that keeps oscillating forever — but where its time-average settles as the horizon stretches out. Notice the object on the right-hand side, $k + K \cdot E[G’]$: it is the same accounting identity introduced earlier (cost price plus the average profit rate applied to capital advanced), except the profit rate is now written as an expectation, because it fluctuates.

    Equation 10 — Making the Averaging Explicit
    $$ \Phi^i_t = \lim_{t\to\infty} E\!\left[\varphi^i_t\right] = \int_{-\infty}^{\infty} \!\left(\lim_{t\to\infty} \varphi^i_t(x)\right) f_X(x)\, dx \;=\; k^i_t + K^i_t \int_{-\infty}^{\infty} G'(t,x)\, f_X(x)\, dx $$

    $f_X$ is the probability density of $X$. The equation says the expectation is an average over every possible state $x$ of the system’s turbulence, weighted by how likely that state is.

    Equation (10) earns its keep by making a subtle move legitimate: swapping the order of the limit and the expectation. That looks harmless, but it hides a real question — does the market price $\varphi^i_t$ even converge to anything as $t \to \infty$? The paper’s answer is no: a capitalist system doesn’t settle into a fixed point, it settles into a limit cycle — perpetual oscillation. So the convergence the argument needs isn’t of the instantaneous price, but of its cumulative time-average. That average does converge, for almost every state of the world, precisely because the system is ergodic — the fraction of time the cycle spends in each region of its orbit stabilizes. This is the Birkhoff ergodic theorem doing, in mathematical language, exactly what Marx says in economic language: the production price isn’t the value the market price reaches and stays at, it is the average around which it never stops oscillating. The oscillation isn’t an obstacle to the average — it is the average’s condition of existence.

    Why the Order of Operations Matters

    The paper invokes Lebesgue’s Dominated Convergence Theorem to justify swapping “limit of the average” for “average of the limit.” This requires bounding market prices by some integrable envelope — economically, that no price can grow without limit, which technological ceilings and competitive pressure guarantee — and, crucially, it does not require that the convergence be uniform across sectors. Uniform convergence would mean competition equalizes profits instantly and identically everywhere, with no room for a shock to hit one industry harder than another. Marx’s theory says the opposite, and the math is built to allow it.

    Equation 11 — When the Capital Base Is Also Uncertain
    $$ \Phi^i_t = \lim_{t\to\infty} E\!\left[\varphi^i_t\right] = \int_{-\infty}^{\infty}\!\!\int_{-\infty}^{\infty} \left[k^i_t + K^i_t(y)\, G'(t,x)\right] f_{X\mid Y}(x\mid y)\, f_Y(y)\; dx\, dy $$

    Equation (10) still treated the capital base $K^i_t$ as known exactly. Equation (11) drops that simplification: $Y$ is a second random variable carrying the estimation error in $K$, with density $f_Y$, and $f_{X \mid Y}$ lets the profit-rate perturbation depend on which realization of that error occurred. The object is the same double average — only now uncertainty is propagated from two sources instead of one.

    This last equation is not a mathematical flourish; it is the reason the empirical section spends so much effort on multiple imputation. National accounts don’t hand anyone a clean measurement of capital advanced by sector — it has to be reconstructed from incomplete data, and that reconstruction carries its own error. Equation (11) is the license to treat that error as a random variable to be averaged over rather than a nuisance to be ignored. The uncertainty is propagated externally — by a generator outside the statistical model itself — rather than estimated as an internal parameter of the dynamic model: estimating $K$’s error inside the model would confound it with the model’s own measurement-noise term, opening a ridge of non-identification between two magnitudes that the data alone cannot tell apart. Kept external, twenty-five complete reconstructions of the data are generated first, each respecting the Marxian aggregate identities to machine precision, the dynamic model is fit on each, and the twenty-five fits are combined by Rubin’s rule. That is the outer average of equation (11), computed by literally drawing from the distribution of $Y$ instead of assuming it away.

    The Engine: A Three-Level Ornstein–Uhlenbeck Cascade

    Equations (9)–(11) define the target; they don’t generate a path toward it. The explanans — the mechanism that actually produces a year-by-year trajectory consistent with that target — is a hierarchical Ornstein-Uhlenbeck process with up to three nested levels, fit as a single Stan program (the same program handles one, two, or three levels, which guarantees that adding levels can never silently break the simpler cases nested inside them). All series enter standardized; time is discretized one year at a time using the Euler–Maruyama scheme.

    Level 1 — The Market Price
    $$ dev_{t,s} = \varphi_{t-1,s} – \Phi_{t-1,s} $$
    $$ \kappa^m_{t,s} = \kappa_{\mathrm{cap}} \cdot \mathrm{invlogit}\!\left(\kappa_s + \beta_1\, z^{TMG}_t\right) $$
    $$ \Delta\varphi_{t,s} = \kappa^m_{t,s}\!\left(-\,dev_{t,s}\right) \;+\; a_{3,s}\, dev_{t,s}^{\,3} \;+\; \gamma\, COM^{std}_{t,s} \;+\; \varepsilon_{t,s} $$

    Subscripts $s$ (sector) and $t$ (year) run throughout. $dev$ is last year’s gap between market price and the latent production price. $\kappa^m$ is the sector’s reversion speed, passed through a logit link that caps it inside $(0, \kappa_{\mathrm{cap}})$ and lets the general rate of profit ($z^{TMG}$) modulate it without ever pushing the system out of the stable region of the discretization. $\varepsilon$ is a fat-tailed (Student-t), stochastic-volatility innovation, so volatility can cluster in time without destabilizing the mean.

    Read the Level 1 line as a spring. The term $-\kappa \cdot dev$ is the restoring force: it pulls the market price back toward the production price with a force proportional to how far it has drifted. The cubic term $a_{3,s} \cdot dev^3$, with $a_{3,s}$ constrained negative by construction — not estimated, imposed — makes that restoring force grow faster than proportionally once the deviation gets large: the further the market strays, the harder it snaps back. This is a declared stability assumption, not a discovery: it guarantees the model can never generate an explosive regime, at the real cost that if such a regime existed in some sector of the actual economy, this particular specification could not detect it.

    Levels 2–3 — Where the Latent Center Itself Reverts
    $$ \mu_{s,t} = m_{0,s} + m_1\, G’_t + m_v\, V_{s,t} $$

    The production price $\Phi$ is not treated as a fixed, observed index; it is itself a latent state that reverts — more slowly, with its own sector speed $\kappa_p$ — toward this mean $\mu$. $m_1$ is the channel running through the general rate of profit; $m_v$ is the coefficient measuring how strongly the production price tracks the directly-constructed value $V_{s,t} = k + p$ (Level 3, and the reason the cascade goes up to three levels rather than stopping at two).

    This is the bridge back to the abstract equations above, term by term. $\mu_{s,t}$ is the estimable stand-in for the right-hand side of (9): $m_{0,s} + m_1 G’_t$ plays the role of $k + K \cdot E[G’]$, and $m_v V_{s,t}$ is the specific functional form chosen for the value-tracking channel that the abstract definition deliberately leaves open (the paper is careful to say that capitalist competition as a function of the value structure is declared at the level of equations 9–11, not derived; giving it the concrete shape $m_v V$ is a modeling choice made at the cascade level, defended by how it performs under validation rather than deduced from the definition). And the expectation of $G’$ from equation (9) has its operational counterpart in the profit rate averaged across the twenty-five multiple imputations — the mechanism equation (11) licenses.

    The coefficient $m_v$ carries real theoretical weight: it is the empirical stand-in for Chapter 9’s claim that prices of production gravitate around values. It is given a neutral prior, $m_v \sim \mathcal{N}(0,\, 0.5)$ — centered at zero, symmetric, assigning equal plausibility to $m_v > 0$ and $m_v < 0$ before seeing any data. That matters for the same reason a fair coin matters in a coin-flip experiment: if the data carried no signal, the posterior would sit wherever the prior put it, hugging zero. It doesn’t. It lands at $m_v \approx 1.0136$ with $P(m_v > 0) = 1$ — evidence that the data moved it there, not the prior. The anchoring to value is found, not assumed into the setup.

    In Plain Language

    The cascade is three springs stacked on top of each other. The market price is tied by a spring to the latent, unobserved production price. The production price is tied by its own, slower spring to a moving target that blends the general rate of profit with the directly-measured labor value. Pull any one spring and let go: it doesn’t snap to a fixed point, it settles into the kind of perpetual, decaying oscillation that equations (9)–(11) describe as an average. The springs are estimated from sixty-one years of real U.S. data, not assumed; the coefficient tying prices of production to values, specifically, could have come back negative or zero — the model gave it every chance to — and it didn’t.

    What the Numbers Say

    The empirical core of the paper is a panel of 37 productive branches of the United States economy over 61 years, from 1960 to 2020. The hypothesis tested encloses three distinct relationships, and the paper is meticulous about not conflating them. Each is stated, tested, and reported separately.

    Market Prices ↔ Prices of Production: The Strongest Link

    This is the relationship with the firmest statistical support, confirmed through six independent lines of evidence:

    Central Finding

    Gravitation exists, and it is slow. The median speed across sectors is $\kappa_m = 0.0770$, equivalent to a half-life of approximately 9 years. Market prices take about a decade to cover half the distance toward their production-price center. This is consistent with Marx’s characterization of gravitation as a tendential, mediated regulation, not an instantaneous fit.

    The number is remarkably stable under stress tests:

    • Removing five of the six productive blocks from the panel barely moves the estimate — it shifts in the third decimal place. The sixth, which gathers 18 of the 37 sectors, does produce a shift (from 9 years to 6 years), and the paper decomposes it: about half the acceleration is the generic effect of halving the panel — removing 18 sectors at random already gives 0.0929 — and not the block itself.
    • Dismantling the value anchor in three different ways — including permuting surplus value across spheres — moves the speed in the third decimal place. This is significant: it means the conclusion about market-to-production gravitation does not depend on the less robust production-to-value link.
    • The market deviation has its own dynamic signature. Compared against a random walk matched in variance, three out of six test statistics separate cleanly (the weighted-sum convergence reaches a tolerance of 0.01 while the null never reaches a tolerance ten times more lenient; recurrence analysis laminarity triples the null; recurrence entropy doubles it). The ones that don’t separate are recurrence-analysis determinism and the two deterministic-chaos invariants — the Lyapunov exponent and the correlation dimension — which the paper never claimed to find.
    • The estimate is invariant to secondary methodological choices. Sweeping the latency regularizer across three values produces life medias of 9 years in all three arms (speeds of 0.0774, 0.0770, 0.0772).
    • The known bias of disaggregation pushes against the result. Splitting a national figure among 37 branches is underdetermined and biases speed estimates downward — meaning the true half-life is probably 7–8 years rather than 9. A bias that works against your conclusion is one you can live with, because the result holds despite it, not thanks to it.

    Prices of Production ↔ Values: The Thinnest Leg

    This is the weakest part of the empirical argument, and the paper states so with complete transparency. The problem is not a defect of the instrument but a property of the object:

    Methodological Transparency

    The coupling coefficient estimated within the dynamic model is $m_v = 1.0136$ with a 95% credible interval of $[1.0096,\; 1.0176]$ — but the same procedure returns 1.0365 when surplus value is permuted across spheres, preserving all annual aggregates. Why? Because production price and value share the cost price, which explains 66.1% of the variance of the former and 72.0% of the latter, and their correlation in levels is 0.9987. The coefficient would land near one even if the law of value didn’t hold at all. The paper therefore reports it as a consistency check, not as evidence.

    The real support for this relationship comes from cross-sectional tests, not from the dynamic coupling. When temporal common trends are removed and analysis is conducted within-year, the slope of the markup on own surplus value is 0.675 with the true data versus 0.090 under permutation, with intervals that don’t come close to overlapping. The sectoral ordering of the wedge between $\Phi$ and $V$ has an inter-annual rank correlation of 0.986 and a 60-year value of 0.558 — highly persistent structure, not noise.

    A collateral finding worth noting: the coefficient of variation of sectoral profit rates is 0.669 — meaning profit rates across industries show considerable and persistent dispersion. Far from contradicting the theory, this dispersion is the condition of existence of the mechanism: if profit rates were already equalized, there would be no differential to drive capital migration, and gravitation would have nothing to operate on. Marx postulates equalization as a tendency, not an accomplished fact.

    Market Prices ↔ Values: Sustained in Form, Adjusted in Existence

    The structural modification across sectors exists and is nonlinear (the nonlinearity step holds comfortably at 6.8 null deviations). But the existence step is adjusted: 44% of its gain is obtained equally with sectoral characteristics unpaired from their spheres, and the gap against the maximum null is on the order of one paired standard error. The coefficients survive a deliberately severe correction for serial dependence (tripling the error).

    The Instrument Behind That Number: A Nested Ladder in gdpar

    That test is a small ladder of nested distributional-regression models, fit with gdpar (Gómez Julián, 2026b), the author’s own R package for generalized distributional parameter regression, published on CRAN on July 15, 2026. The ladder climbs from a bare model — “the market-to-value ratio has no sector-specific correction at all” — through a model where organic composition, wage share, and sector size shift that ratio linearly, up to a model where the correction is a flexible spline rather than a straight line. Two gains matter, measured in units of predictive density: adding the linear correction buys 207.3 units; letting it curve buys another 215.1. Both were checked against a control built to be hard to pass — shuffling which sector gets which characteristics 99 times, refitting each time, with the spline’s knots held fixed across every shuffle so the comparison can’t be won by a better basis alone. The curvature gain clears its null with room to spare (6.8 null standard deviations; the best of 99 shuffles reaches only 114.8 against 215.1 observed). The existence gain is honestly reported as thinner: shuffled sectors still buy about 44% of the real gain merely by having some characteristics to fit — three covariates and an intercept give a model room to accommodate noise even when it is being told nothing true — so the genuine margin over the null sits at about one paired standard error (23.2, against a gap of roughly 24 units). Both numbers are reported together, precisely so the large one isn’t read alone.

    A companion specification, estimated in the same gdpar fit, asks the same question about dispersion rather than location: not where the market-to-value ratio is centered, but how tightly it clusters. Larger sectors and sectors with higher capital composition show systematically less relative dispersion — elasticities of $-0.226$ and $-0.104$ — consistent with equalization operating more effectively where capital is more concentrated. Both effects clear a “breaking factor” (the multiple of the standard error at which the 95% interval would first touch zero) north of six and four respectively, past the 2.94 ceiling reached anywhere else among this paper’s location coefficients, and the finding reproduces under a completely different likelihood family (a gamma distribution on the price ratio) to within 5.2%.

    Three Failures That Confirm the Theory

    One of the most intellectually striking features of this paper is how it handles results that, at first glance, look bad for its thesis. There are three, and the paper reports all of them without softening — then shows deductively why each one was expected if the theory is correct.

    Negative Result No. 1

    The model does not out-of-sample predict better than a random walk. But this was deductively implied by the slow form of the thesis. At a horizon much shorter than the half-life, a mean-reverting process is, to first order, a random walk. If something takes a decade to get halfway back, looking at a single year won’t let you see it return.

    Negative Result No. 2

    The value term is predictively indistinguishable. Again, this follows from the slow coupling between prices of production and values: with half-lives on the order of decades and only 61 years of data, univariate root-unit tests are structurally underpowered.

    Negative Result No. 3

    No univariate test separates the true wedge from its permuted placebos. But this was predicted before measuring, by the persistence of sectoral ordering itself (inter-annual rank correlation of 0.986). A highly persistent time series is hard to distinguish from its permuted version using tests designed for shorter memory.

    Finding these signatures is corroboration of the slow form of the thesis, and not finding them would have been the real problem. — Gómez Julián, on the negative results

    The paper’s stance on this is worth highlighting: “Lejos de refutar la tesis, los tres están deductivamente implicados por su forma lenta” — far from refuting the thesis, all three are deductively implied by its slow form. A single mechanism (slow gravitation) explains both the substantive thesis and all the apparently negative results, and it also survives in the validated posterior. “That a single cause explains the thesis and all the apparently negative results, and that it additionally survives in the validated register, is the opposite of a petitio principii: it is a unified, falsifiable, and internally validated narrative.”

    Temporalism Isn’t a Preference — It’s a Condition of Measurement

    Perhaps the most consequential result in the entire paper is not a number but a statement about what can and cannot be measured. It concerns the “modulator” — the component of Marx’s argument in which the general rate of profit enters into the structural modification of each sphere, meaning the deviation of each sphere is not independent of the reference but generated by it.

    The Identifiability Argument

    When the model was run with a single, fixed general rate of profit for all 61 years (as a simultaneous approach would require), the posterior exhibited a flat ridge: two completely different functional bases (a degree-two polynomial and a spline basis) produced the same pathology to the third decimal place, with an effective sample size of only six draws. The diagnostic got worse with more sampling (R-hat rising from 1.33 to 1.73). This is the unmistakable signature of a direction in parameter space along which the likelihood does not change.

    The cause is theoretical, not computational. With one fixed reference, the modulator can only be identified evaluated at that single point — a single number, not a function over the space of references. You cannot estimate three coefficients from a polynomial if you have one data point.

    When the reference was allowed to vary year by year (61 different general rates of profit), the model converged within minutes, with a large improvement in both time and effective sample size, and zero divergences.

    Named, Not Improvised: Theorems 1A and 1E

    This diagnosis isn’t an ad hoc read of a misbehaving sampler. gdpar (Gómez Julián, 2026b) — the same package behind the nested ladder above — ships a formal identifiability result for exactly this situation. Its Theorem 1A establishes that, with a single fixed reference point, a distributional modulator is identified only at that point: as one number, not as a function over the space of possible references. Theorem 1E is the positive counterpart: letting the reference vary restores identifiability of the modulator as a function. Fitting a degree-two polynomial (three coefficients) or a five-knot spline basis (five coefficients) against one single, unmoving reference asks for more than a single data point in that dimension can support — which is exactly what a flat likelihood ridge looks like from the sampler’s side.

    The figures behind the improvement, precisely: a fixed reference with a degree-two polynomial gives an R-hat of 1.7333, an effective sample size of 6, and 8 divergent transitions in 39 minutes; a one-knot spline basis reproduces the same pathology — R-hat 1.7335, effective sample size 6, 14 divergences, 5.6 hours. Letting the reference vary year by year (61 distinct annual values of the general rate of profit), centering the additive component and raising the sampler’s adaptation parameter to 0.99, gives an R-hat of 1.0035, an effective sample size of 1332, and zero divergent transitions — in 2.9 minutes. That is the 115-fold improvement in time and 222-fold improvement in effective sample size referenced above, and it is a theorem, not a tuning trick: no amount of additional sampling closes that gap under a fixed reference, because the object being asked for — the modulator as a function — simply is not there to find.

    The consequence is stated precisely: with a single fixed general rate of profit obtained by solving the system simultaneously, the claim of Chapter 9 of Volume Three of Capital is unverifiable by construction. It is not that the data are insufficient — the object is not identified, and no amount of data would identify it. The argument does not establish that simultaneism is false as a description of capitalism (that is established by historiography and sociology); it establishes that a simultaneous procedure cannot, even in principle, empirically verify the specific part of Marx’s argument that this work estimates.

    In Plain Language

    Marx says: first a general rate of profit forms, then each industry deviates from it according to how capital-intensive it is. To check whether the deviation depends on the general rate, you need to see what happens to the deviation when the general rate changes. If you calculate one general rate for the entire 61-year span, it never changes, and there is nothing to observe. That is exactly what happened: the model with one fixed rate doesn’t converge — not because of computational limitations, but because it is being asked to measure a relationship with a single observation of one of the two variables. Calculating one rate per year — which is what the temporal reading says you should do — the same model converges in three minutes.

    What This Is, and What It Isn’t

    The paper is careful, almost painstakingly so, about the limits of what it claims. This section matters because a reader coming from the “pro-Marx” or “anti-Marx” side might be tempted to over-read the results. The author doesn’t let you.

    What the evidence authorizes: In the United States between 1960 and 2020, market prices gravitate toward prices of production with a decadal half-life that is sectorially heterogeneous, and this speed survives three independent assaults (removing five of the six productive blocks, destroying the value anchor, varying secondary methodological decisions). This is a measured, calibrated, and falsifiable fact.

    What the evidence does not authorize:

    • It does not claim superior predictive power (the model does not out-predict a random walk, which was expected).
    • It does not claim that univariate root-unit tests confirm gravitation (they are structurally underpowered at this time scale).
    • It does not claim uniqueness or categorical novelty. The contribution is the explicit integration and canonization of a slow gravitation cascade with value anchoring, measured on real data, with propagated uncertainty, validated, and subjected to a diagnostic whose unfavorable results are reported alongside the favorable ones.
    • It does not claim that this statistically demonstrates the law of value, “and not for rhetorical prudence but because it would be false: a price series can show that a magnitude behaves as the law predicts, and cannot explain why that magnitude exists or whether the category with which we name it is the correct one.”

    That last point is the paper’s deepest epistemological commitment. Questions about whether “value” is the right category for what prices ultimately measure are not answerable by any price series, no matter how long. They are answered by history, sociology, and philosophy — and the firm answer is the one obtained when all four disciplines (those three plus statistics) point in the same direction. The four-dimensional convergence is the argument, not any single leg of it.

    The paper also addresses the homology that unifies its seemingly disparate halves — the historiographical-filosofical first chapter and the econometric second chapter. The relationship between necessity and contingency that governs the transition from feudalism to capitalism (where the same demographic shock produced opposite outcomes in different regions of Europe) is structurally identical to the relationship between prices of production and market prices. A law determines the center; circumstances determine each particular outcome. Neither fact negates the other, because they describe different levels of the same reality.

    What It All Adds Up To

    Here is the simplest version of what this 260-page paper establishes:

    Marx was reproached for a century for having done an arithmetic calculation wrong. What happened is that his calculation was redone under an assumption he never made: that the prices of things bought to produce and the prices of things that come out of production are the same prices, fixed at the same time. If you assume that, Marx’s accounts indeed don’t close. But that assumption is equivalent to saying the economy doesn’t happen in time. As soon as you accept that what exits the factory this year is what enters the factory next year, the accounts close without anyone having to fix anything. — Gómez Julián, Summary for the Reader

    But recognizing the conceptual error was only the first half. What had been missing — and what this paper contributes — is doing those accounts with real data instead of with fictitious numerical examples, which is what the school that had the correct conceptual reading had never done.

    The empirical results show that prices in the U.S. economy over six decades do behave as the theory predicts: they gravitate, slowly, toward prices of production calculated with Marx’s theory and no other. This finding survived every attack the author could devise — removing productive sectors, destroying the value anchor, permuting surplus values, varying methodological decisions, and running diagnostics whose unfavorable results are reported in full alongside the favorable ones.

    The part of the argument linking prices of production to labor values is also supported by real evidence, though less firmly, and the paper says exactly where the weak points are and why they are properties of the object, not defects of the instrument.

    And the paper does not claim to have demonstrated the law of value with a series of numbers, because “questions of that kind are not answered with numbers: they are answered with history, with sociology, and with philosophy, and the firm answer is the one obtained when the four things (the previous three, together with statistics) all point in the same place.”

    That convergence doesn’t make the result eternal — better evidence can overturn it tomorrow. But it makes it, for now, “our best possible approximation to the truth.”

    — — —

    “In science as in life, overcoming adversity is what makes us truly strong.”

    This post summarizes the introduction, conclusions, and the formal-empirical chapter (§2.4) of Gómez Julián, J. M. (2026). Some Reflections on Marx’s Prices of Production: Historicity of the Law of Value, Dialectical-Materialist Foundation, and Dynamic Formalization Under Uncertainty. Zenodo. https://doi.org/10.5281/zenodo.21842251. The full paper spans approximately 260 pages across two chapters covering philosophy, historiography, mathematical formalization, and empirical econometrics. Equations (9)–(11) and the model specification cited here reproduce that chapter’s notation; gdpar is cited separately as Gómez Julián (2026b).

    Written for the curious. An invitation to read.

  • BITOPOLOGICAL SPACES: LISTENING TO THE DIRECTION OF TIME WHEN IT MATTERS

    BITOPOLOGICAL SPACES: LISTENING TO THE DIRECTION OF TIME WHEN IT MATTERS

    Bitopological Spaces: Listening to the Direction of Time When It Matters

    Bitopological Spaces: Listening to the Direction of Time When It Matters

    How two topologies — built from the same data — can hear the difference between past and future

    Most of the tools we use to analyze sequences of data — averages, correlations, spectral analyses — treat time as a label that could run in either direction without changing the answer. Reverse the order of your data points and many standard methods give you identical results. But in the real world, the direction of time matters profoundly. Economies expand slowly and crash suddenly. Heartbeats rise smoothly and fall steeply. A method blind to direction is a method blind to one of the most fundamental features of how systems change.

    A recent paper by independent researcher José Mauricio Gómez Julián introduces a construction that addresses this gap. Taking a known method from graph theory and extending it to directed graphs, the paper produces a pair of topologies — mathematical frameworks for understanding structure and connectivity — whose divergence is a topological fingerprint of temporal irreversibility. Applied to three decades of American economic data, the method recovers a picture that is both mathematically precise and economically interpretable. Here is a walk through the main ideas.


    Seeing and Being Seen

    The starting point is a beautifully simple idea introduced by Lucas Lacasa and collaborators in 2008. Imagine plotting a time series — say, 129 consecutive quarterly growth rates of U.S. GDP — as points above a timeline. Now connect two points with a line if they can “see” each other: the straight segment between them passes above every intermediate data point, as if you stood at one point and shone a flashlight toward the other with no obstacles in the way.

    The result is a visibility graph: a network whose nodes are time points and whose edges encode a geometric relationship. Visibility graphs have been used to classify chaotic systems, detect heartbeat anomalies, and distinguish between types of economic regimes. They translate the shape of a time series into the structure of a graph, opening the door to the vast toolkit of network science.

    But standard visibility graphs are undirected: an edge between two points does not record which one came first. If you orient each edge from the earlier time point to the later one, you obtain a directed visibility graph — a directed acyclic graph in which the arrows always point forward in time. This orientation carries information about temporal asymmetry that the undirected graph throws away entirely.

    From Networks to Structure

    Here is where the paper’s contribution begins.

    In 2018, Huda Nada and collaborators introduced a procedure for turning any undirected graph into a topological space. For those unfamiliar with the term, a topology is a mathematical framework that defines what it means for groups of points to be “open,” for sets to be “connected,” and for spaces to have “structure.” It operates at a level more abstract than distances or coordinates — it captures the pattern of how sets overlap and separate.

    The Nada construction works as follows. For each vertex of the graph, compute its closed neighborhood: the vertex itself plus all of its immediate neighbors. Take this family of neighborhoods and generate a topology by closing it under two operations: finite intersections (combine neighborhoods by overlapping them) and arbitrary unions (combine neighborhoods by collecting them). The result is a topology on the vertex set, and its invariants — connected components, separation properties, component counts — capture structural features of the graph.

    This procedure is universal: it works for any family of subsets of any set. The mathematical content lies in identifying the right family to use.

    Gómez Julián’s key observation is that for a directed graph, you do not get one family of neighborhoods — you get two. For each vertex:

    The forward closed neighborhood includes the vertex itself and all the vertices it points to — the later time points it can see. The backward closed neighborhood includes the vertex itself and all the vertices that point to it — the earlier time points from which it is visible.

    Apply the Nada procedure to the forward neighborhoods and you get a topology τ+. Apply it to the backward neighborhoods and you get a topology τ. The resulting triple (V, τ+, τ) is what mathematicians call a bitopological space: a set equipped with two topologies simultaneously, a concept introduced by John Kelly in 1963.

    The extension is, in a precise mathematical sense, trivial — the topology axioms do not care where the generating family came from. But recognizing this trivial extension as the right thing to do, and showing that the resulting bitopological structure captures something real about temporal asymmetry, is the paper’s central insight.

    When Two Topologies Disagree

    If the process generating your data is symmetric — equally likely to go up as down, at the same speed — then the forward and backward neighborhoods are statistically exchangeable. The two topologies τ+ and τ look the same, and the bitopological structure adds nothing beyond the undirected construction.

    But if the process is asymmetric, the two topologies diverge. Consider the prototypical asymmetry of economic and physical systems: gradual expansion followed by sudden contraction. Forward visibility through a gradual rise connects many points — each point can see far ahead through the gentle slope. Backward visibility through an abrupt drop connects few — the sharp fall blocks the line of sight. The forward topology ends up more connected (fewer separate components) than the backward topology.

    This divergence is the topological fingerprint of temporal irreversibility. The paper defines three quantitative measures of it:

    The asymmetry direction Δ = C − C+, where C+ and C are the numbers of connected components in the forward and backward topologies. Positive Δ means the forward topology is more connected. The component-count irreversibility index IC, which normalizes the difference to lie between 0 and 1. And the base-size irreversibility index IB, which measures the analogous difference in the sizes of the generating bases.

    These are pure numbers — no calibration, no free parameters, no training data. They emerge from the structure of the data and the construction itself.

    If you reverse the direction of time in your data and the topologies change, something in the process that generated the data is irreversible — and the gap between the two topologies measures exactly how much.

    Peeling the Onion: Three Layers of Structure

    One of the paper’s most clarifying contributions is the identification of three nested layers of topological structure on the same time series, each revealing different information.

    Layer 1 — The Alexandrov topology (reachability). For a directed acyclic graph, the most basic topology treats as “open” any set that is closed under forward reachability: if a node is in the set, all its descendants are too. This topology has exactly one connected component for any weakly connected graph, because every node can reach every later node through some directed path. At this level, the system is globally indecomposable. It tells us what we already know: the economy is a single connected process in which each quarter influences every subsequent quarter through chains of causation.

    Layer 2 — The undirected Nada topology (local fragmentation). When you apply the Nada construction to the undirected shadow of the visibility graph, the topology fragments dramatically. The intersection closure of the neighborhoods reveals clusters of time points that share local structural similarity — groups of observations linked by overlapping visibility neighborhoods — that go beyond mere reachability. This layer uncovers genuine structure that the reachability topology hides entirely.

    Layer 3 — The bitopological layer (temporal asymmetry). When you split the construction into forward and backward, a further distinction emerges. The forward and backward topologies have different component counts — and the difference is invisible to the undirected construction and invisible to the Alexandrov construction. It lives only in the gap between the two directed topologies.

    Each layer is contained within the next: the Alexandrov topology is a subtopology of the Nada topology (a theorem proved in the paper), which in turn underlies the bitopological structure. But each coarser layer hides information that the finer layer reveals.

    What the American Economy Looks Like Through a Topological Lens

    The paper applies the full pipeline to the quarterly growth rate of U.S. real GDP from Q1 1992 to Q1 2024 — 129 observations spanning the dot-com bust, the Global Financial Crisis, and the COVID-19 shock. Two graph constructions are used: the Horizontal Visibility Graph and the Natural Visibility Graph, both in their directed forms.

    The headline finding is that Δ = +4 in both constructions. The forward topology has 4 fewer connected components than the backward topology, regardless of which visibility-graph variant you use. This positive value is consistent with the well-documented asymmetry of the American business cycle over this period: expansions are gradual and sustained (1992–2000, 2001–2007, 2009–2020), while contractions are sharp and short-lived (2001, 2008–2009, 2020). Forward visibility through a gradual expansion is unobstructed; backward visibility through an abrupt contraction is fragmentary.

    What makes this finding compelling is its invariance. The HVG and NVG produce very different graphs — 248 vs. 406 edges, different base sizes, different absolute component counts — yet they agree on the sign and magnitude of Δ. The signal appears robust: a feature of the underlying data, not an artifact of how you choose to draw the graph.

    Another detail worth noting: the base-size irreversibility index IB is exactly zero in both constructions. The forward and backward topologies are generated by bases of equal size (250 and 250 for the HVG, 239 and 239 for the NVG). The asymmetry lives entirely in the structure of those base elements and how their intersections distribute — not in how many there are. The two topologies are built from the same number of building blocks, but those blocks fit together differently depending on whether you are looking forward or backward through time.

    A Single Shock

    Perhaps the most striking empirical finding is what the topology says about the COVID-19 shock.

    The second quarter of 2020 recorded the sharpest contraction in U.S. GDP on record — an annualized rate of roughly −31%. The third quarter recorded the sharpest rebound — roughly +33%. These are the two most extreme observations in the entire 129-quarter series, opposite in sign and opposite in economic interpretation.

    A naive analysis would naturally separate them: one is the worst crash, the other the best recovery. They sit at opposite ends of the value spectrum.

    But the Nada topology classifies them together. Under both the undirected and directed topologies, under both the HVG and NVG, these two observations belong to the same connected component.

    Why? Because the topology is not a proximity measure. It does not group points by how close their values are. It groups them by the structure of their visibility neighborhoods — which other points they can see, and how those visibility patterns intersect. Despite their extreme and opposite values, the two quarters share neighborhoods that overlap substantially. The intersection closure, which drives the Nada construction, puts them in the same cluster.

    This matches the interpretation most economists give to the event: the contraction and the rebound are two phases of a single exogenous shock, driven by the same underlying cause — the pandemic and the policy response to it. The topology recovers this interpretation from the geometry of the data alone, without any economic priors built in.

    What the topology does not claim: it does not say that the two quarters are “similar” in value (they are the two most distant observations in the entire series). It says they are structurally linked — that no topological open set separates them. The construction responds to the combinatorics of visibility, not to the metric of distance.

    Certifying the Construction

    The paper takes reliability seriously at three levels.

    Machine-checked proofs. The central theorem and related core results have been formalized in Lean 4, a proof assistant, against the Mathlib mathematical library. A computer has verified that the proofs are logically correct, with no gaps or hidden assumptions. The formalization is archived alongside the paper as part of a reproducibility bundle on Zenodo.

    Polynomial-time algorithms. Every step of the construction has an explicit algorithm with proven complexity bounds. The connected components of each topology can be computed in polynomial time without enumerating the full topology, which can be exponentially large. The key trick is to work through a combinatorial proxy for the topology called the specialization preorder, using bitset operations that are highly efficient in practice.

    Honest uncertainty. A three-valued decision procedure for pairwise connectedness reports “pairwise connected,” “pairwise disconnected,” or “undecided” — the last when the computation exhausts its resource budget. Rather than guessing, the algorithm honestly reports that it has not finished the work. This is a methodological commitment as much as a technical one: a topological statement counts as established only when the computation has completed the work that establishes it.

    No free parameters. The construction has no tuning knobs. The topological invariants — component counts, base sizes, irreversibility indices — are determined entirely by the data and the definitions. There is nothing to calibrate, nothing to overfit.

    A New Lens

    The paper does not propose to replace existing methods of time series analysis. Correlation, spectral analysis, regime-switching models, and the many other tools of econometrics and statistics capture information that topology cannot see: amplitude, frequency, distributional shape. The paper is explicit about this complementarity.

    What the topological construction offers is a new lens — one that responds to the relational structure of a time series rather than its metric structure. It asks not “how big is this change?” but “what does this change connect to, and what does it disconnect from, and is the answer different depending on which direction in time you are looking?”

    For systems where temporal asymmetry is a defining feature — business cycles, climate dynamics, physiological signals, causal event sequences — this lens may reveal structure that traditional tools, by their very construction, cannot see.

    The application to U.S. GDP is a proof of concept. The construction is general: it applies to any time series that can be turned into a directed visibility graph, which is to say, any time series at all. Whether the invariants it produces are useful features for classification, prediction, or interpretation in broader contexts is an empirical question that the paper opens but does not close.

    What it does establish is this: there exists a construction that takes a time series, produces two topologies from it, and quantifies the gap between them as a measure of temporal irreversibility. The construction is mathematically sound, mechanically verified, algorithmically tractable, parameter-free, and — when applied to the American economy across three turbulent decades — gives answers that make economic sense.

    That is a foundation worth building on.

    “Bitopological Spaces from Directed Graphs: Extending the Nada Construction to Capture Temporal Irreversibility” by José Mauricio Gómez Julián is available at Zenodo (v1.0.2, April 2026). The complete research compendium — Lean 4 formalization, R package, empirical dataset, and reproducibility notebook — is archived alongside it.

  • General Dynamic Parameter Models via Reference Anchoring

    General Dynamic Parameter Models via Reference Anchoring

    You can also find this library at CRAN and download it directly from R and RStudio.

    Also, we recommend viewing the mind map summary at the end of the article to better understand the relationship between the functions of the package.

    R Library Review

    Meet gdpar

    General Dynamic Parameter Models via Reference Anchoring

    In the fleeting calculus of a two-second decision—overtaking a car on a narrow road—the human brain performs a remarkable statistical trick. It does not build a model of the approaching driver from scratch. Instead, it retrieves a baseline: the average driver, representing typical reaction times and modal aggression. In a split second, it reads the specific signals of the actual driver—relative speed, vehicle type, micro-movements—and estimates how this specific driver deviates from the baseline. The decision to overtake emerges from that synthesis.

    This cognitive recipe—population reference + individual deviation—is the philosophical bedrock of the R package gdpar (General Dynamic Parameter models via Reference Anchoring) by José Mauricio Gómez Julián. The package takes this intuition, formalizes it as a rigorous statistical decomposition, proves the conditions under which it is mathematically identifiable, ships a Stan-based Bayesian engine to estimate it, and layers on causal inference, geometry-adaptive sampling, and dependence-robust inference.

    The Anatomy of Deviation

    Every layer of gdpar is an elaboration of a single, elegant equation. For each observation $i$ with covariates $x_i$:

    $$ \theta_i \;=\; \theta_{\text{ref}} \;+\; \Delta(x_i,\; \theta_{\text{ref}}) $$

    Read it as: the parameter of individual $i$ equals a population reference, plus a deviation that is itself a function of the individual’s covariates and of the reference itself.

    That final clause is where the architecture pivots from classical statistics. The deviation $\Delta$ does not merely depend on who you are (your covariates $x_i$); it depends on what the reference is. If you transplant the model to a new population, the deviation function behaves differently because $\theta_{\text{ref}}$ is one of its arguments. This structural dependence is the defining feature of “reference anchoring.” It distinguishes gdpar from random-effects or varying-coefficient models, where the deviation is structurally separate from the reference.

    So, what is the shape of $\Delta$? The package singles out a specific functional form called the Additive–Multiplicative–Modulated (AMM) decomposition:

    $$ \Delta(x,\theta_{\text{ref}}) \;=\; \underbrace{a(x)}_{\text{additive}} \;+\; \underbrace{b(x)\odot\theta_{\text{ref}}}_{\text{multiplicative}} \;+\; \underbrace{W(\theta_{\text{ref}})\,x}_{\text{modulated}} $$

    Three mechanisms, cleanly separated and independently interpretable:

    • $a(x)$ — A pure additive shift. Think of this as a traditional fixed-effect driven by covariates.
    • $b(x)\odot\theta_{\text{ref}}$ — A covariate-dependent scaling of the reference (using the Hadamard/elementwise product). This is where “the deviation depends on the reference” enters multiplicatively.
    • $W(\theta_{\text{ref}})\,x$ — Covariates are mixed through a matrix $W$ that is, itself, tuned by the reference. This is the explicit, structural reference-dependent channel.

    Standard models drop out as special cases. Set $\Delta \equiv 0$ and you have fixed-effects regression. Set $W \equiv 0$ and you have a hierarchical model with multiplicative interaction. Set $b \equiv 0$ and you have a varying-coefficient model. The AMM is the smallest natural family that contains all three and elevates the reference to an active argument of the deviation.

    The Three Estimation Engines

    gdpar defines three complementary engines for estimating $\Delta$. Crucially, only one is executable in the current release—a deliberate choice to promise a mathematical scope that exceeds the executable surface, and to say so honestly.

    Path Engine Representation Status
    Path 1 Hierarchical Bayesian (Stan) Parametric AMM ✅ Operational
    Path 2 Varying-coefficient (splines) Smooth $\beta(z)$ 🚧 Conceptual
    Path 3 Hypernetwork / Neural Net Net generates $\theta_i$ 🚧 Conceptual

    Paths 2 and 3 are documented to “reference grade”—full asymptotic theory (contraction rates, Bernstein–von Mises) is developed in the Wiki—but they abort with gdpar_unsupported_feature_error if invoked. Path 1 places priors on every component ($\theta_{\text{ref}}, a, b, W$) and samples the joint posterior with HMC, yielding native, full-posterior uncertainty.

    A Tale of Two Posteriors: EB vs. FB

    Within Path 1, gdpar offers two inferential regimes. Full Bayes (FB) via gdpar() samples the joint posterior, remaining most faithful to the cognitive analogy. Empirical Bayes (EB) via gdpar_eb() estimates the hyperparameters by maximizing a marginal likelihood via a Laplace approximation, then samples the remaining parameters conditionally.

    The EB vs FB Comparator

    Rather than forcing a choice, gdpar treats them as parallel routes. It ships a dedicated comparator, gdpar_compare_eb_fb(), which quantifies agreement on $\theta_{\text{ref}}$ and the reduced parameter vector $\xi$. The Wiki develops the theory to first-class depth: EB and FB lower-level posteriors agree asymptotically (Theorem 7A), while EB intervals under-cover by $O(n^{-1})$ (Proposition 7B). If you have ever wondered if EB is “good enough” for your data, gdpar lets you answer that empirically.

    Distributional Regression: Every Parameter is a Slot

    gdpar is not constrained to modeling the mean. A probability distribution has multiple parameters—location, scale, shape, tail index, zero-inflation probability—and each one can carry its own AMM decomposition. The package indexes these by $k = 1, \dots, K$:

    $$ \theta_i^{(k)} = \theta_{\text{ref}}^{(k)} + \Delta^{(k)}(x_i, \theta_{\text{ref}}^{(k)}), \qquad k = 1, \dots, K $$

    The built-in roster covers Gaussian, Poisson, negative binomial, Bernoulli, Beta, Gamma, Student-$t$, Tweedie, ZIP, ZINB, and hurdle families. Zero-inflated and hurdle models receive an especially elegant treatment: both the zero-inflation probability $\pi_i$ and the count parameter $\theta_i$ are anchored to their respective references—a dual deviation design.

    The Causal Bridge

    Because the AMM form produces individual parameters, individual treatment effects emerge naturally. gdpar_causal_bridge() implements a T-learner: fit the anchored model separately under treatment and control, then read the conditional average treatment effect (CATE) at $x_i$ as the difference of the anchored individual predictions:

    $$ \widehat{\tau}(x_i) = \widehat{\mu}_1(x_i) – \widehat{\mu}_0(x_i) $$

    A second layer, gdpar_compare_meta_learners(), benchmarks the AMM-based learner against external meta-learners via pluggable adapters: grf::causal_forest on the R side and EconML’s CausalForestDML on the Python side (via reticulate). The framework’s causal claims are benchmarked, not asserted.

    Mechanics & Clockwork

    Several engineering decisions elevate gdpar from a theoretical exercise to a serious computational environment:

    • Stan Code Generator: Composes programs from canonical pieces—AMM blocks for $p=1$ and $p \geq 1$, EB marginal/conditional blocks, distributional-$K$ blocks—selected by the resolved $(K, p, \text{family}, W, \text{parametrization}, \text{group})$. The $W$ basis supports B-splines with Stan-side Cox–de Boor evaluation, ensuring differentiability inside HMC.
    • Identifiability Pre-flight: Before any sampling, gdpar_check_identifiability() runs a Gram-matrix check (Proposition 1C), a per-coordinate cross-component check (C4-bis) for $p > 1$, and a per-group anti-aliasing check (C7). If your design is non-identifiable, you find out before the sampler burns your CPU, accompanied by a structured gdpar_identifiability_error naming the dependent directions.
    • Data-Driven Reparametrization: Treats the parametrization of $b(x) \odot \theta_{\text{ref}}$ as a pre-fit decision. A short pilot computes an information ratio, dispatching to CP, NCP, or—gdpar‘s root-cause resolution—a linear reparametrization that samples the product $\theta_{\text{ref}} \cdot b$ directly, sidestepping bilinear funnels altogether.

    Opt-in Power Tools

    Two advanced capabilities are switched off by default, documented as thoroughly as the core path.

    1. Geometry-Adaptive Sampling

    Hierarchical AMM posteriors can be geometrically hostile—funnels, near-determinism, heavy tails. The opt-in geometry engine climbs a ladder of Riemannian metrics: Euclidean → Fisher/SoftAbs → sub-Riemannian → relativistic/Finsler. A certifying orchestrator diagnoses the pathology, selects a metric, tunes the integrator, and emits a certificate. If full sampling is certified infeasible, a Laplace fallback provides a plug-in posterior with ELPD on par with mgcv-REML or INLA-Laplace.

    2. Dependence-Robust Inference

    gdpar does not model temporal or spatial dependence in its point structure; instead, it makes the inference robust to dependence (a working-independence + sandwich-variance stance in the spirit of Liang & Zeger, 1986). You receive diagnostics (Durbin–Watson, Ljung–Box, Moran’s $I$) and robust SEs via block bootstrap—moving or circular blocks in time (with the Politis–White flat-top automatic block length), tiled randomized-origin blocks in space. Point estimates remain pristine; only the uncertainty is made honest.

    ⚠️ Honest Limitations

    The Wiki is admirably forthright about scope. Only Path 1 is executable in 0.1.0. Dependence is not modelled—only the inference is made robust. The package’s mathematical scope exceeds its executable surface by design. Read the “Implementation status” notes carefully before relying on a feature.

    TL;DR

    gdpar takes one of the most natural ideas in human prediction—predict an individual as a deviation from a population reference, where the deviation itself depends on the reference—and transforms it into a fully specified, identifiability-checked, Stan-powered Bayesian regression framework. It is theoretically rigorous, computationally serious, and unusually honest about what it does and does not yet do. If your work involves individual heterogeneity, distributional regression, or causal effect estimation with principled uncertainty, gdpar demands a careful look.