Relocated from: @fiat_lux@lemmy.world ⛓️‍💥(04-2026)

  • 0 Posts
  • 11 Comments
Joined 5 months ago
cake
Cake day: April 24th, 2026

help-circle

  • If a country is an enemy, it is condemned as undemocratic and sanctioned in various ways; but if it is an ally, the fact that it lacks freedom of expression, human rights or democracy is ignored …

    The European Union, in fact, imposes economic sanctions on one country, and sends financial aid and weapons to another; yet fails to do the same in the face of other, even more serious invasions with even more brutal consequences for entire populations …

    These contradictions … suggest that, in practice, concerns boil down to the political and economic interests of different regions of the globe …

    There is no longer a real and stable framework of truth and values.

    - Cardinal Víctor Manuel Fernández, prefect of the Dicastery for the Doctrine of the Faith, present day.

    Hmm…

    The decision by a sovereign instrumentality to give funds to a foreign anti-Communist political movement rather than to a Communist regime, at the time where the Cold War was beginning in earnest in Europe, is not a “commercial” act; it is jure imperii, a deeply sovereign act

    - The Vatican, ~2005, arguing in Alperin v. Vatican Bank why they shouldn’t have to hand over documents about, or, compensate victims of the holocaust, for laundering the gold the fascists brought them in WW2.

    I look forward to seeing the documents and restitution now that the Vatican has reversed it’s position. I expect this is the height of pontification though.



  • lemmy.zip was the nearest similar comm? This could have been put in an instance that is involved with the dispute, many of them have Fediverse communities.

    Putting aside the irony of cross-posting a thread about the potentially ideologically inspired muting of smaller instances by Piefed to draw attention to lemmy.ml’s potentially heavy-handed censorship and bias harming the growth of the Lemmy-verse - it looks less like “vitilizing” and more like fragmenting discussion.

    More ironic still is that if I see one of your posts now it means that I’ll probably go look at ml to see the actual discussion and hear more from the OP. Perhaps posting original content might go further to achieve your goals?





  • This list is weird, aside from the length. They must be using a very greedy regexp for this many instances to have their names partially censored.

    The text “buds” has been censored, all the instances using the TLD “university” have had “univer” removed, and the word “hangout” is also gone. “Shitpisscum” made it through, so it can’t just be about slightly naughty words. Also annihilation.social is listed 3 times for some reason.

    Are these slurs in a culture I’m not familiar with? Does piefed do this everywhere?


  • We can see that it’s solved by the fact that AI models continue to get better despite an increasing amount of AI-generated data being present in the world that training data is being drawn from.

    Even if it logically followed that model improvement means model collapse is a solved problem, which it absolutely doesn’t, even the premise that models are improving to a significant degree is up for debate.

    MMLU pro benchmark over time line graph showing plateauing values Massive Multitask Language Understanding (MMLU) benchmark vs time 07-2023 to 01-2026

    A lot of people really want to believe that AI is going to just “go away” somehow, and this notion of model collapse is a convenient way to support that belief

    Model collapse may for some people be an argument used to support a hope that AI will go away, but the reality of that hope does not alter the validity of the model collapse problem.

    You can tell it’s not a solved problem because researchers are still trying to quantify the risk and severity of collapse - as you can see even just from the abstracts in the links I provided.

    Some choice excerpts from the abstracts, for those who don’t want to click the links:

    Our results show that even the smallest fraction of synthetic data (e.g., as little as 1% of the total training dataset) can still lead to model collapse

    …we establish … that collapse can be avoided even as the fraction of real data vanishes. On the other hand, we prove that some assumptions … are indeed necessary: Without them, model collapse can occur arbitrarily quickly, even when the original data is still present in the training set.


  • It can’t only be from data from previous generations, even if the initial demonstration used that, because that would mean a single piece of human-generated text is sufficient to avoid collapse.

    The loss of data from generation to generation is one way model collapse can occur, but it’s only one way. The actual issues that cause collapse are replication of errors and increasing data homogeneity. In a world where an unknown quantity of new data is AI generated, it is not possible to ensure only a certain quantity is used as future training data.

    Additionally, as new human generated content is based on the information provided by AI, even if not used intentionally in the construction of the text itself, the error replication and data diversity issues cross over from being only an AI-generated content problem to an all content problem. You can see examples of this happening now in the media where a journalist relies on AI output to fact check, and then the article with the error gets republished by other media outlets.

    Real AI training methods may stave off some model collapse, if we ignore existing issues around the cultural homogeneity of training data from across all time periods, or assume the models are sufficiently weighted to mitigate those issues, but it’s by no means settled that collapse is a non-problem.

    You’ve mentioned using data mixing to prevent collapse, but some of the research suggests that even iterative mixing isn’t sufficient dependent on the quantities of real vs synthetic data. Strong Model Collapse (2024), Dohmatob, Feng, Subramonian, Kempe goes into that, and since then there’s been When Models Don’t Collapse: On the Consistency of Iterative MLE (2025) Barzilai, Shamir which presents one theoretical case where collapse won’t occur provided some assumptions hold, but the math is beyond me. They also note multiple situations where near-instant collapse can occur.

    How much data poisoning might affect any of that is not at all clear, it would need to be in sufficient quantity for whatever model to have an effect, but it certainly wouldn’t help. The recent Bixonimania scandal suggests it’s feasible.


  • “model collapse” was demonstrated by repeatedly training generation after generation of models on the output of previous generations

    the best models these days are trained largely on synthetic data - data that’s been pre-processed by other AIs to turn it into stuff that makes for better training material

    You can prevent model collapse simply by enriching the training data with good data - stuff that is already archived, that can’t be “contaminated."

    This feels like an odd juxtaposition.

    If model collapse can be avoided by enriching with uncontaminated data, and model collapse comes from using training data generated by previous generations, doesn’t that imply that:

    1. Either the best models are headed towards model collapse, or,
    2. Models can’t be updated because modern data isn’t usable?