Archive.is blog

Blog of http://archive.is/ project
  • ask me anything
  • rss
  • archive
  • A bit more on the topic of “web archives and news media”:

    1. The Tehran Times website has been up for only a few minutes a day in recent weeks (I don’t know if this is related to the military action or to new Western sanctions against Arvan Cloud—since, at least, the edge nodes aren’t in Iran), so web archives are essentially the only place where you can read it 24/7. There’s also a Telegram channel, but it lacks longreads.
    2. New media outlets (the substack generation) are reluctant to link to old media because they can’t add referral IDs and earn a percentage of sales, like in e-commerce and adult industries. Archives, of course, don’t pay either, but some authors prefer linking archives anyway, just to spite greedy tabloids, much like it was on Gamergate. There’s probably a niche here for something like Readly—a site that republishes other’s articles behind its own paywall and pays revshare.
    • 4 months ago
    • 1 notes
  • Does this website track you? (I found a: top-fwz1 domain in the source code of archive is)

    Anonymous

    Sure, everything is collected by KGB

    • 4 months ago
    • 2 notes
  • Are you DMCA compliant?

    Anonymous

    I wouldn’t say either “yes” or “no” here. That is, ‘yes’ in the sense that we do remove (or blackrectangle) some pages based on user reports, but “no” in the sense that we would not be eligible for safe harbor if we strictly followed the DMCA (i.e., remove everything, without a fee, but based on non-anonymous reports that must include the verifiable name/address/phone number of the reporter). DMCA safe harbor only works for ISPs and certain websites (those that have appointed an agent and registered them with the U.S. Copyright Office); for everyone else, it is still not “compliance” but “voluntary behavior resembling compliance”

    In a strictly legal sense: “no”, because there is no USCO-registered agent.

    • 4 months ago
    • 4 notes
  • How do I get illegal revenge porn images removed?

    Anonymous

    I think it’s better to go through the police or authorities (NCMEC, WHRIK, etc.). This issue is a really complex, overlapped with OnlyFans piracy, and there are far more parties involved: there are scammers offering to remove content from their own sites for a fee, scammers offering to remove content from other sites for a fee, victims who paid them for nothing, leaks of photo verifications from escort and swinger sites, etc. Everyone has their own made-up story, so we won’t judge who is legal and who is illegal, especially based only on who complained first or who said word “illegal” first. Write a report to the police or the authorities, they will assess your situation and forward it to us if removal is necessary: if you are eager to resolve this without involving the police, we get the impression that you are more likely to be one of those scammers (or OnlyFans revenue-optimizers) than a victim.

    • 4 months ago
    • 3 notes
  • Hello, hope everything is going well :) 2 small questions: Any specific reason as to why the archive was down earlier due to a server error? Also, every so often I (and some friends) will see nothing but a rotating loading circle when trying to access the website, what’s the reason for that? Thank you!

    Anonymous

    just tech issues

    • 4 months ago
    • 2 notes
  • After a lot of reading, I'm inclined to believe that the accusations against you are fabricated, but there's one thing I'm still not clear on: did you actually set up any code to send repeated random requests to an external site?

    If so, I'm curious what your reasoning was, as I can't seem to think of a good reason. Even if was a significant threat, that still seems like an odd way to respond.

    P.S. If you get the TV interview, could you link to a recording, pretty please? (Even if there's no English)

    Anonymous

    It was Patokallio’s blackmail from the beginning: if the 3Hz “DDoS” would stop, he wouldn’t hype the story via his friends in the media, even though hyping it was actually preferable for us than slowly sliding into inferno. We had to continue with the “DDoS” to keep annoying him, so he, thinking he was repaying us with harm, actually carried out a scenario that is a win-win-win for everyone: Jani got his advertising clicks which help him survive in overpriced Australia; Wikipedia got its pill from moral panic; we got vaccinated against the next WAAD (probably against Conde Nast too: after their covering the Wikipedia story, it would be a bit difficult to attack us as “yet another 12ft clone” without accusing Wikipedia in large-scale piracy).

    The “changed page” with an Easter egg was supposed only to Patokallio’s eyes: it is the dead link linked from his blog, absent in archive.org (along with whole `lj.rossia.org`—why? someone’s petty revenge for Verbitsky’s “Anti-Copyright” book?), so he must rely on the archived copy to attack the very same archive. The dramatic reaction from Wikipedians, when Jani revealed it to them, was somewhat unexpected, but it did not change anything: they closed the referendum a few days earlier, the result would have been the same. They could have opened and won such a referendum without all this drama—just Patokallio old blog post’s propaganda, which is considered an “authoritative source” there, would have been enough.

    This also reinforced the initial guess that it was Patokallio who was behind the media amplification of the “FBI story”: the very same Jon Brodkin of ArsTechnica—who seeded the “FBI story"—acted again as a Patokallio’s sockpuppet: he republished Patokallio’s blog posts instantly, while never wrote anything about us without Patokallio (e.g. on more newsworthy AdGuard/WAAD drama).

    • 4 months ago
    • 2 notes
    • #patokallio
    • #conde nast
  • The drama has reached the point where it seems we have found a way to accept tax-deductible donations in the US, and German TV is asking for a 15-minute interview.

    And someone says: “admit that it turned out badly,” ha ha.

    • 4 months ago
    • 3 notes
  • By the way, to save everyone from the usual round of naïve questions (“do you store IP addresses?”, …), to dispel a few comforting illusions about what Terms of Service actually mean, and to give you a glimpse into what it feels like to run, moderate content, and support a website, we highly recommend Kevin Nguyen’s novel New Waves.

    It’s a sad, sharp, and often very funny story about a website that differs from ours in only one key respect: it promised not to preserve content, but to delete it on a schedule. As you might guess, that promise becomes the most fragile and consequential part of the whole enterprise. The book captures something that anyone who has worked behind the scenes of an internet platform will immediately recognize: the gap between what users imagine, what policies declare, and what operational reality demands.

    We doubt Kevin was OSINTing us—or even knew we existed—when he wrote it. And yet the parallels are striking. In fact, there are more accidental coincidences there than in some stories written by people who explicitly tried to study us.

    If you want to understand not just how platforms present themselves, but how they actually function under pressure—read or listen it. It won’t answer every question, but it might help you ask better ones.

    • 4 months ago
    • 3 notes
  • In a response you said: "I would understand your claim if you wrote that you unsubscribed from recurring donations. But why should I—or anyone else—be interested in your use of a free utility?" Maybe you shouldn't be interested, except that my concerns seem to be shared by lots of other people — and by Wikipedia, which says you are no longer trustworthy. Maybe you don't care about that either, but if you don't care what users or Wikipedia or any of its users think then why did you bother setting up archive.today in the first place?

    Anonymous

    Wikipedia is more than just a lesbian SJW politbüro trying to hide behind a hypernym. Wikignomes (actual content writers) love us for neat features like accordion unfolding or joining multi-page content into a single page, while WMF tries to distance itself from copyright issues and always remained silent. All of this has low priority for many reasons; for example, it seems that this is the only site that actively uses the archive whose bosses have never reached out to us to ask anything or negotiate anything (the only contact is GreenC who is not even allowed to run their WaybackMedic on WMF servers). So it’s mutual.

    Why bother setting up: to watch over some pages changing over time. The mistake was to release a free service to the public, it must be paid (or by-invite) from the beginning. There are way too many free users.

    • 4 months ago
  • anyway, what's your relations to verified.lu?

    Anonymous

    The same as to the Olympic ice skater: Patokallio’s cherry-picking.

    Better ask Patokallio why he wrote on verified.lu and ignored the ice skater, and how that logic got propagated to become an authoritative source? And how should we treat those who spread this (including Wikipedia), realizing that it is either satire or provocation? With flattery, perhaps?

    • 4 months ago
  • The archive does host a lot of CSAM though - not to impugn the work of the archive, but whenever somebody seems to report CSAM to the archive, pages are merely hidden rather than fully removed (eg with a vpn). A user complained on x a while ago about the archive not removing CSAM when reported and some archived pages (like smutty.***) contain potentially hundreds of CSAM. Have you at least considered looking yourself and removing any CSAM you see?

    Anonymous

    CSAM in user (not, for example, NCMEC’s) reports is often not actually CSAM, but merely a magic word they expect will prioritize their messages. It is not uncommon for users to claim copyright on their “CSAM.”

    Many Smutty’s pages reported by authorities have been removed, so I assume they have reviewed everything. Total removal of whole Smutty is sort of an OSINT case, not a CSAM one.

    • 4 months ago
    • 1 notes
    1. "Trustworthiness" in my question meant "reputation of being trustworthy", i.e. reputation of you not modifying saved pages to insert fake information into them. Guarantees about your archive being trustworthy are literally on the main page:
      > Archive.today is a time capsule for web pages! It takes a 'snapshot' of a webpage that will always be online even if the original page disappears
      > …
      > and provides a short and reliable link to an unalterable record of any web page.
      This was true until very recently you chose to tarnish that reputation and began inserting fake information into saved pages.
    2. Do you think there's absolutely no illegal content on your archive, including CSAM? That's simply consequences of you refusing to remove information no matter what. I'm not saying that this approach is wrong or right, you just need to accept these consequences when someone points them out.
    3. Is me saying that some website/company is American also "nationalism" or whatever bad political position you think it is? Or you constantly mentioning that Jani is a finn? This really reeks of "everyone I don't like is Hitler". Don't be ridiculous.
    Anonymous
    • First, it was not a page you saved; it was a page me saved, moreover from a social network’s webpage me controlled - so I could as well alter the original. That “reputation” is at least no worse than Reddit’s. Second, we do delete/alter pages (see the question about whistleblowers), although we try to minimize that impact. If you need stronger consistency guarantees, you should develop something on a blockchain: Arweave has done something like that. However, it would hardly be compatible with maintaining a clearnet website.
       
    • Of course there is some CSAM, as on any user-generated content website. On Twitter, for example, or on Quora. I know this because many CSAM pages reported to us (and deleted) originate from those platforms, not from chans or .onion websites. And those are not ephemeral pages quickly uploaded and screenshotted — they have been there for years; some of them are still there.
       
    • Tone of our texts about Jani is a parody of his texts about us: what it might look like if Jani wrote about Jani, we call him “finne” as often as he calls us “guerrilla russian jew”. And we see him (and you?) being unhappy reading that and attempting to take it down, by sending same requests to the same people at Automattic.

      It is not about “political positions” at all (who cares on that shit?), it is about preparing the media landscape before sending bogus legally-looking PDFs to the critical infrastructure, just as well as my texts and the whole drama—thanks to “DDoS” promoted to media by himself—might create him problems on border crossing or when approving a loan at a bank. That’s the idea—why threaten to start involved in a scandal under your real name when you know nothing about your opponent, and in your “investigation” you cherry-pick Russian and Jewish names among others solely to include the words “russian” and “jewish” into the text (and then into friendly media, wiki article, etc)? Or to write that “the service does NOT accept crypto” just to slip “crypto” into the text, at least that way? I used to perceive Jani’s texts as merely satirical until serious media started treating them as an authoritative source; then it stopped being funny.
    • 4 months ago
    • 1 notes
  • I wanted to say I still support you on liberapay, but now I see I stopped paying about a year ago when my card somehow stopped working. Oops. Maybe you should set a paywall, lol. (No, I love that you don't.) Your site is great even with all the recent drama. You and that girl that does Sci-Hub are two heroes of the internet that everyone somehow also pretends to hate. Of course in the ideal platonic world your site would not be necessary, but here we are.

    Anonymous

    thanks!

    • 4 months ago
    • 3 notes
  • Well, then let me chim in: I cancelled my weekly Liberapay donations to you in January after all of the DDoS nonsense. Why should I continue to donate to you after you've just destroyed all trustworthiness you had?

    Anonymous

    I guess I have already addressed this in my previous messages.

    1. The so-called “trustworthiness” never existed outside either that stock photo scam or in connection with the term “destroyed.” References to “trustworthiness” in public discussions of the archive could only arise when it was explicitly questioned. Previously—thanks to Jani Patokallio—only terms like “guerrilla” and “jews” and “FBI wanted” were used, leaving no space for a “trustworthiness” axis: the questions were floating not around whether it is “a trustworthness site or not”, but whether it is “a criminal pedo site or not”. Do you understand the difference between when say 2% are ready to believe the first thing and when 2% are ready to believe the second thing? The former can cause merely fluctuation in donations, the latter - deplatforming (and not only in tech infrastructure, think Durov story).
       
    2. While the snapshots were indeed used in civil courts, this was only possible with my written statement—not merely based on some random page online. And yes, in such stories, you are dependent on my whim—whether I choose to assist you or not, as I am under no obligation to do so.
       
    3. Do not reduce Patokallio’s activities to merely OSINT or doxxing. Publishing open or private information is not an issue at all—there is nothing inherently wrong with “information” part itself. The problem lies in the high density of buzzwords intended to provoke squeamish disdain among certain groups of readers—with quite high a coverage given the variety of those buzzwords—and in the dissemination of that propaganda in the media (and, yes, in the Wiki article, I initially missed that). It is not about publishing random names; it is about commenting on the ethnicity of each name (even that alone shifts the genre of his blog from “investigative journalism” to something else), about cherry-picking words. It is the same kind of preparation as WAAD registering the domain krimo-avocate.fr or his French child porn association. A serious existential threat to us, not comparable to the risk of losing a $2/week donor.
    • 4 months ago
    • 3 notes
    • #patokallio
  • I have used and like archive.today, but the knowledge that when I use it, my browser is hijacked to DDoS some guys website — or that archive.today has been altering screenshots of pages — makes me question whether I can trust your service any more. Why should I continue to use it?

    Anonymous

    Well, I would understand your claim if you wrote that you unsubscribed from recurring donations. But why should I—or anyone else—be interested in your use of a free utility? The less you use it, the better it works for the rest, the more time I have for developing new features. It’s already answered as the first question in a new work-in-progress FAQ


    I want to take a moment to share another story. This is the second time the archive has been called an “internet notary,” and the previous case also led to retrospective changes.

    There’s a fairly widespread scam going around (and it may still be active): stock photo agencies send claims to small website owners, demanding, for example, $9,000 for alleged copyright infringement. If the victim contacts a lawyer, the lawyer often jumps at the opportunity to increase both GDP and overall happiness index, saying something like: “Give me $1,500, and I’ll negotiate with the agency to settle for $3,000”.

    These scammers have started using the archive as an “internet notary,” adding statements like: “Even if you delete our protected image, we have notarized proof in the archive,” effectively turning us into their unwitting accomplice.

    Right now, this should be the group most disappointed. Aren’t you one of them?

    • 4 months ago
    • 4 notes
  • I get your perspective is different and its true that the site has been unfairly maligned for a long time. I guess my perspective is different because I've been using the site for over a decade, and I have deeply appreciated what you are doing, I've donated in the past. Your site is a valuable resource precisely because it often contradicts the narrative of those 'noble' sites. It may feel like a thankless effort, it likely is but its relevance shows its not a lost cause. Thanks for those years.

    匿名

    Thanks for warm words

    • 4ヶ月前
    • 3 notes
  • in /post/111781148156/are-some-ips-blocked-to-reache-your-site you mention the block

    匿名

    Oh, 11 years ago… it is probably not actual anymore

    • 4ヶ月前
    • 2 notes
  • out of curiosity what was the reason for blocking small netblock in Poland ?

    匿名

    Which one? the blocks are mostly AS-based, not country.

    In general, many media company offices are blocked; not only because of this drama, we have dramas almost daily (they just lack their own finn to broadcast): in Poland, Wyborcza employees actively bypassed paywalls on competing news outlets and at the same time wrote complaints to Google on 3rd party browser extensions related to archive (the same asserting nobility through a scapegoat).

    • 4ヶ月前
  • The entire reason your site was valuable is because of its credibility as a third party archival site. There's a reason the U.S. govt wants to find anything they can to discredit you; because it threatens them to have an objective resource that you can reliably point as proof something was on the Internet or was changed. With this DDoS shit and editing pages you have completely destroyed the site's credibility, & for what? bafflingly self-induced sabotage. Now you've given an excuse not to care.

    匿名

    Well, look at the attitude of the articles and discussions before the “DDoS” (or even FBI) story. Was it significantly better? The bias was there already.

    Myself, I am not happy with this pivot toward being a pirate site where noble websites offload their shit to notorious, making them even more noble and us more notorious.

    The drama just incised an existing abscess that would have burst eventually anyway, while you suggest petting and powdering it.

    For us it is definitively better, for nobles — not so. Don’t you consider their actions as “self-induced sabotage”?

    Yet another improvement: in the absence of tabloid dramas (yet presence finne troll blogs), it was easy for attackers like WAAD to depict us as a “child porn” website: they just put into their report much more various (dis)info overwhelming our media presence. A “website banned on Wikipedia” must be not a “child porn” one at least, while the finne troll’s initial discourse and its dissemination by tabloids was open—not to say crafted—for such interpretations, heavily contributing to increasing trust in letters from WAAD, June Maxam and other trolls (the conflict with the finne troll lies in this, not in his disclose of anything sensitive; that has now been corrected by finne himself, and even better than if he had simply deleted his post). It’s not about number of people read the blog. It’s about what someone who’s never heard of us would see first after getting a WAAD-letter. For example, if they check Wikipedia’s article, or whatever AI tools show them first. If the quotes come from the finne blog, she’ll likely believe the letter, there are also Russian Carders, German Jews, FBI and so on, right? If “this site was used on Wikipedia for years and then got banned for some rant” she likely won’t. The default reputation baseline—thanks to Jani Patokallio’s compilations—was around a “russian” “jew” “shady” “underground” “guerrilla” “carder” “child porn” “FBI wanted” website, not “a wholetrusted internet notary”, as you try to depict.

    If you think this is all paranoia and conspiracy theories, consider the fact that WAAD registered a French association whose sole activity, after a year spent in the cold, was to launch a barrage of complaints against us. Trolls are patient operatives, they do play the long games.

    • 4ヶ月前
    • #patokallio
  • Also stop fucking coping about how this was actually "worth it" and a good thing. Take some responsibility for the position and importance you had, and how utterly useless your entire handling of this has been even in pursuit of whatever retributive goal you had in mind. Have you ever heard of the Streisand effect? You do realize that by tying the downfall of the site to that guy's blogpost, now people will have to talk about it as context for what you did? Just own it and admit the mistake.

    匿名

    That’s basically “keep calm and serve a freeloading ingrate”.

    • 4ヶ月前
    • 1 notes
  • On growth limits, finances, and why OSINTers focus on seemingly random, yet interconnected people:

    The “startup” logic seems straightforward: if you have users, buy more servers, grow, invest. Add a paywall or a donation button with sad eyes to cover costs.

    So, what’s stopping you?

    The lack of reliable financial infrastructure.

    This isn’t about post-2022 Russian sanctions or allegations of facilitating illicit activity via PayPal. It’s about the impossibility of reliably exchanging or transferring significant sums within a predictable timeframe without falling victim to a scam (on lower-level services) or triggering an AML check (which often feels like a scam in itself). If there is a delay, the service dies. If you accumulate too many donations on PayPal, you get flagged, and your account is frozen for six months. If you transfer funds to or from a crypto exchange, your bank account may be locked, forcing you to travel thousands of kilometers to resolve it (it was fun during COVID, trying to navigate vaccine mandates). Going 100% Monero (or even Bitcoin) is impossible: donation volumes are significantly lower in crypto, and there are many expenses in fiat. All kinds of intermediaries and small underground fintech companies come to the rescue, but there is a limit to relying on them: how much can you really trust them? Over time, it only gets worse: exchange fees and the risk of losing the entire transaction amount keep growing.

    The names and pseudonyms that catch the attention of OSINTers (and probably the FBI, given their interest in “who paid and where”) are not fully random. They often belong to these fintechs: the names used on their credit cards, for example. These are the names of both managers and “drops”, sometimes used unwittingly. Among them are those who worked honestly but are now resentful that OSINTers exposed their names in the context of archives and “piracy,” effectively preventing them from creating new ventures. Then there are those who ran away with the money (including some of ours) and are now actually on the run from police and creditors, obsessively scrubbing their data from the web.

    Managing the finances of even a small project like ours would require a full-time specialist. Hiring one, in turn, requires a reliable monthly income. Just one fuck-up by this person could kill the project. And while bracing for that fuck-up, we would have to aggressively solicit donations just to pay their salary.

    • 4ヶ月前
    • 2 notes
  • Since we are a service for latecomers, we have to debunk a couple of myths for latecomers (the list will be updated, and the answers will be supplemented and rewritten, then I will add the results to the FAQ).

    1. We need as many users as possible, and they have some value — no, we don’t need them; there are almost always more users than we can serve. The captcha that annoys you — it’s mainly not against bots.
       
    2. Paywalls as a selling point — not really.

      First, here we’re talking about the West, and that’s far from 100%, and probably not even 50% of users. There’s a lot of Japan and Korea, and there are no paywalls there at all.

      Second, some support for news sites only appeared during COVID, because of COVID itself and the demise of projects that actively used the archive (jeuxvideo, taringa, etc.).

      Before that, practically all news was archived in an unreadable form, with all the modal pop-ups saying “subscribe to our mailing list” and so on.

      When asked to fix it, we replied, “Well, your sites are crap, so we accurately archive them as we see them, and anyway, there’s nothing interesting there. Make normal sites first.”

      So, we started cleaning up with COVID, and as it ended, all sorts of moral dilemmas arose, and these are far from just formal issues of legality, copyright, and jurisdiction.

      There are also more serious questions, such as whether it is ethical for us to participate in the dissemination of news in the absence of something like COVID?

      You may argue that there is a war going on right now (and point to any war), but this is rather a negative example: if the free dissemination of news during COVID can be interpreted as a blessing, many publishers removed their paywalls back then, but it is impossible to explain participation in the dissemination of military propaganda in the same way.

      In short, supporting news sites is like carrying a suitcase without a handle: you should drop it, but in the absence of other big categories of users (jeuxvideo, etc.), this would be tantamount to refusing to support Western users (which may not be a bad thing, see question #1).

      So it’s not so much a selling point as a current point of rendezvous with users from certain locations.
    • 4ヶ月前
    • 4 notes
  • It turned out pretty well.

    The hype around “the site that’s been banned from Wikipedia for the fifth time” is better than “the 12ft.io analogue that’s about to be caught by the feds”

    Why didn’t you write about such events earlier, folks of the tabloids? I don’t expect you to write anything good, because then who would read you, but there was plenty of dramas, wasn’t there?

    Because there was no Jani to nudge you?

    I guess I’ll scale down the “DDoS”.

    • 4ヶ月前
    • 3 notes
    • #patokallio
  • Hi. The site doesn't work for me at the moment, only displaying a loading circle. I'm fearing it may have been blocked in my country out of nowhere, for some reason (I live in France), so I would at least like to know for sure.

    If it isn't what I think the problem is, then whatever else is happening I hope gets fixed. I've been a fan of the site for a while, and I would be sad to no longer be able to use it.

    匿名

    fixed

    • 4ヶ月前
    • 1 notes
  • I almost forgot…

    June Maxam — NorthCountryGazette.org — the first “victim” of low-frequency “DDoS” — Jani Patokallio’s perfect doppelgänger. Or maybe just another manifestation of the same demonic force in our dimension: travel book author, publisher, editor, gonzo blogger, curious “cybersecurity” “researcher”, and… serial doxxer (she actually did time for doxxing her neighbors).

    Only one thing that sets her apart from Jani Patokallio: her curiosity wasn’t aimed at other websites, it was aimed at her own visitors.

    If she spotted a “suspicious” (meaning: not from her county) IP in her weblog, she’d fire off a cybersecurity report to the provider and all their upstream peers, accusing them of something along the lines of an “armed intrusion into a secured facility”, yes, from a particular IP address.

    In both sheer chutzpah and the effectiveness of her complaints, she outdid WAAD with its pedo-zoophilia theatrics (though she never quite reached the final boss level: those gorgeous PDFs with official-looking crests, rambling for ten pages about how “servers at IP address X are storing 100,000 illegal bitcoins”; really, what can you do when an L1 support tech, eyes shining, has already grabbed the angle grinder and is heading for your server rack?)

    Not bad for a lady in her 70s.

    She even managed to wear out Archive.org. NorthCountryGazette.org wasn’t blocked there because of content issues, it’s blocked so that “Save Now” won’t even bother accessing robots.txt. The whole saga played out on Twitter.

    That “DDoS” basically saturated her ability to read logs and blast out her nasty PDFs.

    • 4ヶ月前
    • 1 notes
    • #patokallio
    • #maxam
  • I’ve told you a little about the history of our relationship with Cloudflare above; now let me tell you about archive.org and how did it happen that people started to confuse us.

    We created a service similar to Megalodon, which was already quite popular in Japan. First, we had to choose a domain zone. Not the USA or the EU (the current horrors hadn’t happened yet, but SOPA and PIPA was already being planned), and not the Caribbean, where a registrar’s server could crash and take months to recover. Libya (.ly) was fashionable at the time, but Gaddafi had just been killed. So Iceland seemed interesting: there were bearded sysadmins in parliament, they created Mailpile. Then we looked at which single-word domains were available.

    When we started, archive.org didn’t have a “Save Now” function, so our features didn’t overlap at all. Even our names are different, just homonyms: archive.org is a noun, while we are a verb: “archive.is/today” was intended as an imperative, like “Save Now!”

    Then two things happened. First, archive.org introduced its “Save Now” feature. Second, when we finally started communicating—around 2020—Mark mentioned that they come from a background of left-wing activism (this isn’t a secret; their biographies are public; I just hadn’t looked into them until it was brought to my attention).

    By that time, Gamergate and various other scandals had already occurred. With few small exceptions, the right tended to preserve pages, while the left wanted to delete them. That was my aha moment: no collaborations were possible here. And so we became a kind of dialectical pair: we won’t delete what they delete, and vice versa, even when politics isn’t involved.

    This is what’s driving us in this direction, toward the role of a smaller archive.org. Whether that’s good or bad, I don’t know yet.

    • 4ヶ月前
    • 4 notes
  • I notice that searching a blocked page (I.e imagefap which was pointed out as being blocked on by someone on twitter some years ago) you can’t access any pages (blocked by the NCMEC). But searching the page on the onion version works completely. Is this intentional? Seeing as some archived results might be illegal

    匿名

    It is intentional, as people (incl. the police) sometimes asking to access to blocked (or unedited - with all the original ads and bugs) content. It is easier to direct them to .onion than to implement user accounts with different access levels and maintain it.

    We don’t stand out among .onion sites with that. If problems arise, there will be two .onion sites, with a secret/dynamic domain for the uncensored site.

    • 4ヶ月前
    • 1 notes
  • Hello? Do you remove any archives that contains the names of whistleblowers? I had feared for their safety.

    匿名

    Some got hidden/replaced. They are from Snowden epoch, not recent.

    • 4ヶ月前
    • 1 notes
  • About Wikipedia, I promised to write when the referendum there ended so as not to influence it:

    1. Kiwi Farms set up an on-premise mirror of all archive.xx links from their forum (the volume is comparable). Wikimedia… never even had such an idea. That’s all you need to know about linkrot, contingency plans, and who to blame when links disappear.
    2. The value of the archive for Wikipedia was not in linkrot, but in the ability to offload copyright issues. This is not about paywalls. This is, for example, about copyright trolls writing claims to stock photos, about articles deleted or changed for political reasons but pursued under the guise of copyright, etc. It is precisely these links that become dead, then got replaced with archive.xx, and we become the sink for all the attacks, legal and illegal. The need for fast-flux hosting, pseudonyms, and other pirate attributes stems largely from this. Do we really need this kind of “social burden”? Build your own toilet thing, you have millions.
    • 4ヶ月前
    • 3 notes
  • I notice that if a page is taken down (for example: /hdPc9), adding “/embed” next to the to the URL (/hdPc9/embed) still displays the content. Is this intentional, or is it an oversight/bug or something else?

    匿名

    It seems like a bug. It’s strange that an article about Apple computer ended up being censored. Thanks for letting know, I’ll take a look!

    UPD: everything what was submitted by a particular user got censored (mostly pedo-content)

    UPD2: it could be intentional as well. I recall Scribd behavior: removed content is still accessible via embed codes, so the feature might be modeled after them.

    • 4ヶ月前
    • 1 notes
  • btw you feds are retards that can't censor shit. you'll never be able to censor the epstein files. lemme tell ya something 'bout the internet. Once it's on the internet, it stays there forever. Dedicated people will ensure that and will retaliate with lulz. ever heard 'bout lulzsec and anonymous? gonna be pretty bad for you glowniggers. oh and a little imageboard where nothing is beyond its reach. you better watch out. you're next.

    匿名

    Oh….

    • 4ヶ月前
    • 1 notes
  • i think i understand now. it's either a fed running this or the actual creator pulling a kamikaze after being compromised by the feds.

    匿名

    Why?
    If I were you, I would presume “sold to Anthropic” instead of “compromised by the Fed”
    Well, everyone has their own bubble here.

    • 4ヶ月前
  • What's your favourite type of furry porn to crank it to?

    匿名

    A NixOS joke?

    • 4ヶ月前
    • 4 notes
  • People ask questions like “Why are you discrediting your own service like this?” or “Is it worth it for the blogger in Finland?”

    My answer is: yes.

    The real discredit would have been to leave things as they were and let the bloggers and the tabloids slowly escalate the black paranoia: rhyming with carding forums, framing us as hackers wanted by the FBI, and so on.

    Articles about The Threepenny Three-Hertz “DDoS” are far better than anything those bloggers would have invent next “just out of curiosity”.

    Sure, topic isn’t perfect and it could have been improved, but this is exactly what the finne troll took upon himself to hype, for free. Another topic would not have had such virality, and the same biased tabloids would not have printed it, so you would simply not have heard about it. This is the best topic for us that the tabloids could have printed.

    • 4ヶ月前
    • 3 notes
    • #patokallio
    • #conde nast
  • Regarding the FBI’s request, my understanding is that they were seeking some form of offline action from us — anything from a witness statement (“Yes, this page was saved at such-and-such a time, and no one has accessed or modified it since”) to operational work involving a specific group of users. These users are not necessarily associates of Epstein; among our users who are particularly wary of the FBI, there are also less frequently mentioned groups, such as environmental activists or right-to-repair advocates.

    Since no one was physically present in the United States at that time, however, the matter did not progress further.

    You already know who turned this request into a full-blown panic about “the FBI accusing the archive and preparing to confiscate everything.”

    • 5ヶ月前
    • 3 notes
    • #fbi
    • #conde nast
  • So we got the situation reversed: now the finne troll got into kafkaesque realm of sending GDPR requests to AI-agents murmuring about safe harbors and journalistic exemptions.

    This is exactly what we warned him about when he decided that Streisand is on his side: this game can be played by two people, and there is much more bad press about him in open sources than about us. Promoting black-tar propaganda on us would promote the attention on who is its author as well.

    Unlike Jani Patokallio’s writings on us, we definitively do not disclosure any “personal data” besides that in the book his father wrote and published; his relatives are public personas and their activities are well known.

    Unlike (a son of ambassador) Jani Patokallio, we did not publish any private communications.

    On “who is currently subject to investigations by U.S. authorities for serious offenses related to the hosting of illegal content” - it is not only false, it is exactly the leyenda negra, invented and distributed exclusively by Jani Patokallio with his friends in Conde Nast, and supported only by referencing to each other.

    I am reporting manifestly unlawful content published by the account “archive-is” on your platform, directly targeting my family, the Patokallio family.
    Your online reporting form does not function properly and appears to operate as a simple sandbox without effective follow-up. For this reason, I am contacting you in writing.
    The content concerned is accessible at the following URLs:
    https://archive-is.tumblr.com/tagged/patokalli
    https://archive-is.tumblr.com/
    https://www.tumblr.com/archive-is
    https://www.tumblr.com/archive-is/807369905134518272/the-finne-troll-published-his-response-with
    https://www.tumblr.com/archive-is/807584470961111040/it-seems-people-dont-read-between-the-lines-they
    https://www.tumblr.com/archive-is/806966482173083648/some-time-back-i-sat-down-for-an-interview-with
    https://www.tumblr.com/archive-is/806832066465497088/ladies-and-gentlemen-in-the-autumn-of-2025-i
    These pages contain numerous serious, false, and defamatory statements, including:
    “an OSINT investigation on your Nazi grandfather”
    “His grandfather seems to have been a real Nazi criminal”
    “There is a family. A big one. They move in politics and in the arms trade.”
    “He shames the family name”
    “The most toxic content… reputations in free fall”
    “comparing Jani Patokallio to Hunter Biden”
    These statements falsely associate my family with Nazi crimes, arms trafficking, covert political networks, and illegal activities, without any evidence.
    Other publications detail our family history, professional roles, and personal relationships without authorization, constituting unlawful processing of personal data under the GDPR.
    These contents are used in a context of harassment, intimidation, and doxxing. They also serve as a relay for technical attacks, including DDoS attacks against my website.
    This blog is operated by the operator of the archive.today service, who is currently subject to investigations by U.S. authorities for serious offenses related to the hosting of illegal content.
    As a hosting provider, your liability is engaged once you are aware of manifestly unlawful content and fail to remove it promptly.
    In the absence of clear identification of the author, your platform becomes legally responsible for maintaining this content online.
    I therefore formally request:
    – the immediate removal of all cited content,
    – the closure of the “archive-is” account,
    – the prevention of any republication.
    If no prompt action is taken, I reserve the right to refer the matter to the competent authorities and data protection regulators.
    Sincerely,
    J.Patokallio

    • 5ヶ月前
    • 3 notes
    • #patokallio
    • #conde nast
  • It seems people don’t read between the lines. They took what I wrote on finne troll for neuroslop.

    So let’s say it plain.

    There is a family. A big one. They move in politics and in the arms trade.

    There is one man in that family who does something else. He doxes people on the internet. He does it to drive traffic to his little blog, packed with ads.

    In other words, he is the fool of the family. No use to the real business.

    He shames the family name and gets in the way of his father and his brother.

    You don’t believe it?

    He used to run a blog at patokallio.name. Now it’s called gyrovague.com. The name “Patokallio” is hidden in the sidebar. You have to click to see it. It’s not in the text at all. In the posts, even in quotes, he cuts the name out every time.

    So the talk with the family already happened. It must have gone like this: live how you want, but you are not our brother, avoid using our name.

    And now, in that position, the man has taken up doxing random internet projects just to make a little money.

    Epic, isn’t it.

    • 5ヶ月前
    • 4 notes
    • #patokallio
  • The finne troll published his response with “lightly redacted copy of the entire email thread”.

    And guess what has been lightly redacted?

    - “an OSINT investigation” on your Nazi grandfather who changed the name in 1944, will not vibecode a patokallio.gay dating app
    + “an OSINT investigation” on your Nazi grandfather, will not vibecode a gyrovague.gay dating app


    That’s the sore spot. His grandfather seems to have been a real Nazi criminal, even by Finnish standards. We need to dig deeper.

    • 5ヶ月前
    • #patokallio
  • When the lease on your domain expires, it often gets snapped up by what are called “parking” outfits (which is like calling a toll booth a “roadside hospitality concept”). A parking domain is basically a dead address turned into a little money farm: no real content, just ads, redirects, tracking pixels, and a vague pretense of being a website, all optimized to squeeze value out of whatever stray visitors still wander in.

    But what if the domain that falls into a parking company’s hands was not serving articles or blog posts or cat photos, but scripts. Say, for example, a CDN endpoint. Or a banner network. Or some forgotten third-party JavaScript that thousands of living, breathing sites still quietly load in the background.

    Well then the fun starts.

    Because now the parking company is sitting in the middle of someone else’s supply chain. They can redirect visitors from perfectly legitimate, still-active sites that happen to reference that old domain. And they do it in a way designed to stay invisible. No big splash. No obvious breakage. Just a slow siphoning of traffic that can go unnoticed for years.

    For example, here is a case where traffic was stolen from EJ.ru for four years. Four. Nobody noticed until someone sent a bug report that basically said: “Why can’t I archive pages from EJ?” And the answer turned out to be: because somewhere in the stack, a script was loading from a dead domain that had been picked up by a parking company and turned into a redirect machine.

    Here is the archive: https://archive.today/ww82.echobanners.net

    And another similar story: https://archive.today/www3.widgetserver.com

    So when people start talking about hacker ethics, about bug bounties, about responsible disclosure, you start to wonder how that whole moral economy is supposed to function when the so called respectable domain investors are behaving a little worse than the hackers. Not breaking in, not exploiting zero days, just quietly sitting on expired infrastructure and milking the pipes that nobody remembered to shut off.

    • 5ヶ月前
  • Yesterday one of the archive’s early adopters sent me a link to an article about how various sites block archive.org and asked how things are on our end.

    I wrote back something like: “Honestly, the bigger trend we’re dealing with lately is front-enders shipping a hundred JavaScript files per page, so if even one of them fails to load the whole page collapses like a house of cards. Against that background, even if something like what the article describes did happen, it probably passed unnoticed.”

    …

    An hour later an email arrives from a sysadmin at Condé Nast: “Are you blocking our office IP?”

    “Oh. Right. Yes, we are. You’re reprinting the seed-crystal of a finne troll’s black-tar propaganda about us, laundering it with your brand’s legitimacy, and you still expect to keep using our free service? Have you people completely lost your damn minds over there?”

    …

    Back in the dawn-of-the-Internet era, when hosting providers billed by the gigabyte even for dedicated servers and the Great Firewall of China was still a glimmer in some bureaucrat’s eye, lots of sites just blocked visitors from China. They weren’t buying anything anyway.

    Now everyone blocks everyone they don’t like or don’t profit from.

    Walmart (or Target?) blocks everyone outside the U.S.

    Ukraine’s been blocking VK for a decade.

    Things that feel almost like core infrastructure - (((ifconfig.me))), (((ipinfo.io))), … - block Iran.

    We block Cyprus because it has a suspiciously high density of people with a past best left undisclosed starting shiny new “European” lives from scratch.

    …

    To deal with that reality, a multi-exit VPN, one that chooses a exit node depending on the target IP, has been a necessity for a long time now, for bots and humans, long before “VPN” became a lifestyle accessory.

    But it comes with problems:

    First, privacy. Tracking scripts don’t see one IP, they see several. And even that pattern by itself is a de-anonymizing signal, because there aren’t that many surfers who look like that.

    Second, Cloudflare. The exit gets chosen for the IP, not the domain, and multiple sites are mixed together on the same IP. Some only let you in from the U.S., others only from Europe, etc. There’s no good solution. So you pick some compromise region X based on which of your favorite sites you’re least willing to have broken. All your Cloudflare traffic now goes through region X. And if you yourself aren’t actually in X (because you chose it not by proximity but by least-badness for your personal web ecosystem) then your packets start doing laps around the planet.

    For a multi-exit VPN user, a site behind such a “mixer” CDN ends up slower than a site with no CDN at all.

    And this is yet another reason — after EDNS, captchas (that can pop up instead of any one of a hundred included JavaScript files), random de-platformings, did I miss anything? — that makes Cloudflare a kind of natural antagonist.

    Not exactly an enemy.

    More like a sparring partner you keep finding yourself matched against, again and again, in different disciplines, in different rings, each time convinced this bout will finally settle something, and each time walking away a little more bruised and a little more aware of how strange the whole fight has become.

    • 5ヶ月前
    • 1 notes
    • #conde nast
    • #cloudflare
  • Some time back, I sat down for an interview with Legal Tribune. The subject was mainly about paywalls and about the use of public archives to get around them. Now, that interview hasn’t seen the light of day yet (maybe it still will) but I reckon there are two reasons it got shelved. And those two reasons, in my judgment, deserve to be heard by the public.

    They asked me, plain and proper: “Doesn’t the work of these archives undermine the business model of German media and by extension, democracy itself, truth, justice, and the whole Teutonic order?”

    Well now, the natural answer to that is another question: What undermines that business model more — a quiet archive that does not advertise its accidental remedy, or a big newspaper article that reminds precisely those who can afford subscriptions that paywalls can be avoided?

    Especially when we’re talking about Germany, a country with a mighty fine library system. A system where just about anyone with a library card (which is to say, just about everyone) can already get past paywalls. Even the hard ones. Even the kind the archives can’t crack. There are even special browser tools built just to make it easier. And you can’t fix that with some grand gesture like calling it “piracy” and blocking a domain on der Bundesbrandmauer.

    But what can kill their business model is a public debate that marches straight into every German living room and says: “You don’t actually have to pay for this”, that in fact it is not “pay for access”, but merely “donate for our democracy”, and who would subscribe to that? That kind of idea spreads faster than any archive ever could.

    Now, the second reason. And while I’m at it, let me address those who might accuse me of comparing Jani Patokallio to Hunter Biden yesterday. Yes sir, there is some friction between us and the German media. But it ain’t about paywalls. It’s about their wish to scrub from the archive the articles they already took down from their own site. And that makes a man wonder: how do they pull those same articles out of libraries too when a publisher has second thoughts?

    What’s curious is this: almost every one of those stories is about the misadventures of wayward scions from respectable families, boys and girls who manage to tarnish their own last names with their behavior.

    The most “toxic” content for us isn’t politics at all. It’s pages of hookers (kinetic ones, not virtual) and these well-born kids splashed across the papers. Not ministers. Not presidents. Just reputations in free fall.

    And maybe that shouldn’t surprise us. After all, the word reputation comes from the same old root as puta.

    • 5ヶ月前
    • 2 notes
    • #paywall
  • Ladies and gentlemen,

    In the autumn of 2025, I published a subpoena received from the Federal Bureau of Investigation.

    Since that day, I have been asked time and again: “And what happens next?”

    Well, allow me to tell you.

    I published that subpoena as an act of responsible disclosure. I did not maintain a so-called “canary page” - the kind some operators use to signal they remain free from legal gag orders. My circumstances were such that I was far removed from jurisdictions where such orders carry immediate, enforceable weight. Moreover, my site was never prominent enough to attract a dedicated cadre of volunteers who might vigilantly monitor such a page for changes. Thus, I resolved upon a simple principle: should any authority send me a legal instrument, I would publish it forthwith. And that is precisely what transpired.

    I confess, I anticipated interest from no more than a handful of crypto-anarchists - the very same individuals who had previously urged me to implement a canonical canary page, yet who offered no commitment to actually watch over it.

    Imagine my surprise, then, when the matter spilled into the mainstream news and reached million eyes.

    But let us be clear: these were not news reports in any genuine sense. The standard refrain read, “We have reached out to the site’s operator and will update this story upon receiving a response.” Yet no journalist ever contacted us (only exception is Meduza, asking for an interview and a bigger article later). This was not investigative journalism; it was dissemination - pure and simple. A prepackaged narrative, delivered to newsrooms with the polite request: “Dear comrades, here is the truth - please publish it.”

    Curiously, every one of these ersatz “news” pieces prominently cited a two-year-old blog post by a certain Jani Patokallio as its authoritative source - a rather odd choice, given that it was merely a personal blog entry by an unaffiliated third party. One might charitably argue it was a piece of enduring open-source intelligence. Very well, let us grant that. But then, why do nearly all the links within that “investigation” point exclusively to blog.archive.today? Why not cite the original sources directly? And more tellingly, there exist at least five other substantial OSINT analyses concerning archive.today. Why, then, did every journalist - seemingly in lockstep - select this one particular post? Unless, of course, they were not writing at all, but merely copying and pasting a ready-made text.

    This raises a more troubling possibility: what if that link to the old blog post was not a citation, but a SEO backlink? What if Mr. Patokallio was not a passive observer, but the very author of the seed?

    First of all, he had already attempted to promoute that very blog post in the media two years ago. On that prior occasion, it found a home only at Boing Boing and Gigazine. The second try achieved far wider circulation.

    A cursory AI-groking into Mr. Patokallio’s background reveals a man no stranger to the shadowed corridors of media manipulation. He was instrumental in repackaging community-written content from WikiTravel into commercially published Lonely Planet guides under his own editorial imprint.

    But that is merely the beginning.

    The Patokallio family presents a profile of considerable geopolitical entanglement. His brother, Mikko Patokallio, serves as Senior Manager for Ukraine at the Crisis Management Initiative (CMI), a Finnish NGO deeply involved in conflict mediation and Eurasian affairs.

    Their father, Pasi Patokallio, is a career diplomat who has served as ambassador to Israel, Canada, and Australia. He is also a noted critic of the Ottawa Treaty banning anti-personnel landmines, and his advocacy appears to have borne fruit: Finland withdrew from the treaty recently, paving the way for the mining of its 2,000-miles eastern border. He wrote an autobiography modestly titled ‘Me, guns and the world’

    As for the family name itself - Patokallio - it was coined and officially registered in 1944, a year of profound realignment for Finland, as the nation shifted its wartime allegiance. In Finland, surnames can indeed be “registered” like domain names, securing exclusive rights to their use. One cannot help but wonder what prompted the adoption of a new name at such a pivotal historical moment.

    Thus, we are not dealing with a mere hobbyist blogger who “saw a neat website and wrote a post,” as Jani Patokallio once claimed on Hacker News. This is the work of a member of a family with a shady Nazi-era story and deep roots in arms export, the Ukrainian conflict and information operations (Jani’s profile resembles more of Hunter Biden than an IT blogger) - a long-term, systemic interest in the archive project that may well prove more consequential, and perhaps more dangerous, than the attention of either the proprietor of luxuretv.com with his fake French child porn alliances or even the FBI itself.

    • 5ヶ月前
    • 1 notes
    • #fbi
    • #patokallio
    • #ukraine
    • #conde nast
  • Hi there! /tl1pG is not being captured properly. Thank you!

    匿名

    Fixed

    Sorry for delays, Tumblr questions are not forwarded to email since April and I just spot it.

    • 2年前
    • 12 notes
  • Is there any way you could make an Android app that I could just share the URL to in the sharing menu so it goes right there?

    cnlson


    This: https://play.google.com/store/apps/details?id=com.navasgroup.share2archive

    • 2年前
    • 8 notes
  • bleacherreport articles are showing up blank (screenshot works but the actual webpage archive is blank). Could you take a look at this? Thanks!

    e.g. bleacherreport*com/articles/2583445-nba-opening-night-gets-the-hotline-bling-drake-music-video-treatment

    匿名

    I made a fix for future bleacherreport saves, but this one seems unrecoverable (React often cleans the whole page on JavaScript error, and this is what was happening on bleacherreport)

    • 2年前
    • 5 notes
  • Did you decide to stop allowing the archival of Mature Labelled content from Tumblr? Thanks for your time!

    匿名

    No, I did not. Any examples of broken pages ?

    • 3年前
    • 6 notes
  • Hi, thanks for providing this fabulous archive service! In https://blog.archive.today/post/688077534761566208, you said that user’s IP address won’t be send to website since 2019. Could you provide an option to send IP address to capture localized contents? And some websites may only be reached in certain regions… :-(

    匿名

    Websites no longer look at `X-Forwarded-For` for user region, so I have to use per-website proxies to get localized content and avoid geo-block.

    It is not 100% correct though, so feel free to report a bug it you spot that.

    • 3年前
    • 1 notes
  • Can you expand the text at "See More" on QreQE ?

    匿名

    yes

    • 3年前
    • 3 notes
  • Did something change over the past few days such that the site no longer functions through TOR?

    匿名

    It works.

    Let me guess: if you copy-pasted archiveiya74codqgiixo33q62qlrqtkgmcitqx5u2oeqnmn5bpcbiyd.onion from wikipedia, it won’t work because it contains a zero-width space character inside.

    • 3年前
    • 4 notes
  • Hi, could you have a look why pages from steamcommunity won't archive? Here's an example, the page "The updated Steam Mobile App is now available" is stuck in submit. Thanks for the great work and site!

    匿名

    Could you tell the url of the page? I see no failed pages from steamcommunity

    • 3年前
    • 2 notes
  • please add support for showing sensitive content from mastodon screenshots.

    匿名

    It is supported.

    But. Mastodon is not detected automatically, there is a list of domains which may be incomplete. Just mail me where you see it is not supported yet.

    • 3年前
    • 2 notes
  • Pretty please remove the recent restrictions for archiving Twitter. It's practically useless for archiving Twitter now. I just archived someone's media page and it only captured their first 4 tweets.

    匿名

    I know, it need to be remade almost from scratch because of the changes on Twitter side. Old code resulted in loading spinner and no content at all.

    • 3年前
    • 2 notes
  • I archived a website but the sidebar from the right side is in the middle of the page. How do I fix that?

    匿名

    where?

    • 3年前
    • 2 notes
  • Z7gaQ and 09Vbx are stuck on the loading screens, however the thumbnails for these archives show past the loading screens

    匿名

    yes, fixed.

    thank you for the report!

    • 3年前
    • 1 notes
  • "It needs accounts, otherwise instagram redirects to the login page or shows a fake 404 page. Accounts do not live long." I've entered the full instagram url, all i get is "Not Found (yet?)". what does this mean? the provided ig acount i entered is public not private

    匿名

    This means it has tried 10 times and given up.You didn’t specify an account for the archiver to log in to instagram, it’s just not implemented

    • 3年前
    • 1 notes
  • why isn't instagram archive working?

    匿名

    It needs accounts, otherwise instagram redirects to the login page or shows a fake 404 page. Accounts do not live long.

    • 3年前
    • 3 notes
  • Spoilers on the War Thunder forum (and google caches of it) don't get expanded. Could this be fixed?

    匿名

    sure

    • 3年前
    • 2 notes
  • Some pages don't display images correctly (ie, like archive 'wR3Ar'). Could you fix it?

    匿名

    Those are not missing images but missing ad blocks.

    • 3年前
    • 1 notes
  • I notice on the /wip pages when I archive something from a news site, most of the time is spent loading various trackers, assets that are never displayed like videos, etc. Why not only load what's needed to render content? Not commenting on the state of the web here, just the performance of the archiver.

    匿名

    Known trackers are skipped (their lines are gray instead of green)

    • 3年前
    • 2 notes
  • Why can't Instagram pages be archived anymore? I try to save one and then its like "Sorry this webpage is not available."

    匿名

    Instagram and FaceBook are broken most of the time: constantly getting kicked out of the account, being blocked by IP, … Although there aren’t many requests to save pages from there, less than 100 a day. I think there must be quite a few live users who view many more pages in a day. How they do it is not clear to me.

    • 3年前
    • 1 notes
  • Could you expand all "Drivers details" on 3ZQ7j (mesamatrix)? Thanks in advance.

    匿名

    It opens on click

    • 3年前
    • 1 notes
  • Any possibility of a MacOS Safari extension like the Chrome one? Thanks.

    匿名

    Please, ask the author of Chrome/Firefox extension, it is not me. I cannot, I have no MacOS devices.

    • 3年前
    • 3 notes
  • can you replace recaptcha for cloudflare turnstile, i constantly have to do captchas and cloudflare turnstile is much faster for me. Thanks

    匿名

    No, that captcha is too difficult: I can’t select “strawberry cakes” among others just by the picture

    • 3年前
    • 3 notes
  • Are pages of a domain deleted periodically to be sure aren't more than ~1000? Dezgo pages keeps descending. Now they are 1162, previoulsy were 1300 and before where ~2400. Thanks. Those links couldn't be accessed anymore live.

    匿名

    The pages are not deleted, but number of search results is limited.

    To get more pages try to split the query to `domain.com/a*`, `domain.com/b*`, etc

    • 3年前
    • 1 notes
  • Reddit has really been bugging out the archiver the past few days.

    匿名

    examples?

    • 3年前
  • Economist articles archived from at least today are being cut off, not showing full article & text is being overlaid by the usual list of links/images of related articles. Is this a bug, or?

    匿名

    Could you point to exact page?

    I have seen the 15 latest from Economist and found no issues

    • 3年前
  • Community tabs on YouTube channels doesn't archive correctly; redirects to a specific post on the Community tab instead like this /5lUZU

    匿名

    fixed. (clicking on “expand“ buttons hit one wrong button)

    • 4年前
    • 2 notes
  • Is the search function broken?

    匿名

    yes.

    the index is rebuilding, it will be back in few hours

    • 4年前
    • 2 notes
  • does archiving a webpage send the user’s IP address to the host site?

    匿名

    It was so in old version (before 2019) to get localized versions.

    Now it is useless, because nowadays most of localizations are “this page is not for your country”, so the IP is not passed to the host site.

    • 4年前
  • What is the long version of the url please. Wikipedia wants it to be used....how to I get to it? Greg

    匿名

    Click on “share” button to see different forms of linking to a page.

    For example https://archive.vn/Aoans/share

    • 4年前
    • 1 notes
  • The Archive seems to be having difficulty archiving YouTube urls. What's going on?

    匿名

    youtube started showing captcha to the archiver

    • 4年前
    • 2 notes
  • Has Roskomnadzor behavior or amount of removal requests changed since the start of the war?

    匿名

    No, there were no removal requests related to this conflict at all. Neither from Roskomnadzor nor from other agencies. Although the stream of requests about ISIS content goes as usual.

    • 4年前
    • 3 notes
  • About the SSL_ERROR_NO_CYPHER_OVERLAP in the previous question. This problem isn't on the DNS server, but on the Archive website (google it, many people complaining). I turnaround this error by adding your IP to my HOSTS

    匿名

    Your last sentence proves that the problem is exactly DNS server

    • 4年前
  • Could you expand comments on nNUlG? Thanks!

    匿名

    No, `substack.com` does not expand comments, they are on another page; if I click on “expand comments”, the content disappears.

    • 4年前
  • Підключення для цього сайту не захищено archive_dot_ph використовує протокол, який не підтримується.ERR_SSL_VERSION_OR_CIPHER_MISMATCHНепідтримуваний протокол Клієнт і сервер не підтримують загальну версію протоколу SSL або комплект шифрів. Від браузера не залежить. Через VPN все працює. Територіально - Україна.

    匿名

    Try to change DNS from 1.1.1.1 to something else.

    • 4年前
  • Can't you buy cards with crypto, or something like that? I'm sure some people would be willing to help with that.

    匿名

    That’s exactly the case of <<Some people, when confronted with a problem, think “I know, I’ll use regular expressions.” Now they have two problems.>>

    1. Where to get crypto (a certain amount every month, with no skips) ? Donations are not enough and not stable. Brave browser stopped paying again (not enough anyway). Traveling to the towns trying to buy crypto for paper cash from shady people does not look like a sustainable plan. Getting a job with a crypto salary will not cover the demand too.

    2. Cryptocard services are not reliable and prone to exit scam and regulatory risks. This will require supporting a redundant system on top of at least three of them, maintaining the necessary balances everywhere, tracking news and being prepared for losses.

    On the other hand, advertising (and possibly paid features) allows to escape the money conversion hell and to tune cash flows in different currencies and on different sides of the Iron Curtain, to cover the local expenses.

    • 4年前
    • 1 notes
  • Why are there ads on this site?

    匿名

    I already answered this: https://blog.archive.today/post/677297433505628160/a-bit-different-than-a-usual-question-but-do-you

    Basically, money have become fragmented and hard to convert.

    For example, I have no other way to top up PayPal except with donations (and there aren’t enough of them) or by showing a certain amount of advertising from an agency that pays there. Cards do not work, making new ones involves a trip to a different vaccine zone, etc.

    • 4年前
    • 2 notes
  • do you plan to add any cryptocurrency methods as a donate option? considering that there are many people on the world other side who do not have access to a bank card with international payment capabilities (visa/mastercard etc), they can pay with the local banking system, local currency (cash, checkouts or some online payment system), and they can use cryptocurrency or any medium, or need to add support for so many payment systems (webmoney, JCB, wechat pay/alipay/unionpay, prepaid / gift cards

    匿名

    The site does not have any premium features available only to paid users, so there is no need to consider yourself penalized if you can’t pay.

    • 4年前
  • One year later, how has the OVHcloud fire impacted the project? Are you able to participate in the Action Collective (Class Action) against the company?

    匿名

    No, all the equipment there was rented.

    • 4年前
  • I have a suggestion. When a user archives a page, sometimes an error page is archived (ie, like archive "9UE0W"). When that user is shown the archive for the first time, they (and only that user) could be presented with a question, "Has this page been archived correctly?" If they respond "no", then the archive would be deleted, so they can try again. After a short period of time, this option no longer appears to the user.

    匿名

    Retry won’t help: the page address is invalid, it has “%3F” instead of “?”

    • 4年前
  • Thanks for looking at the Telegram link preview issue yesterday. I'm afraid they still don't work. Note that you can test them, by re-scraping a page, using the @WebpageBot bot.

    匿名

    Yes, it works unstable, I do not know yet why :(

    It is a different issue: the first was about /xxxx works and /xxxx/image does not. And it was reproducible on other previews (Twitter, …).

    The second is Telegram-only and affected all pages.

    • 4年前
    • 1 notes
  • I try to log in to Archive, but for some reason it keeps taking me to a Welcome to Nginx page without ever going to the actual website, is there some sort of way to fix it

    匿名

    there are no accounts and no way to log in

    • 4年前
    • 1 notes
  • Is there any link between your website and the Internet Archive?

    匿名

    No

    • 4年前
  • Hello. Thank you for providing this amazing service. Are you aware that link previews for 'archive today' don't work in Telegram, even when using a link to the screenshot tab? Is there a reason why, and is it possible to fix it? Thanks a lot.

    匿名

    looks like robots.txt issue. it should work now

    • 4年前
    • 1 notes
  • There are at least two use cases for your service: 1. To archive pages for posterity. 2. To bypass paywalls or weekly article limits. For the second use case, the article might only need to be archived for a month, after which the content may no longer be of interest. I wonder if you could reduce storage requirements by running two services: one for permanent archives, and one for temporary archives. Also, I am curious: What other use cases are there?

    匿名

    Basic scenario: saving a page that can be edited or deleted.

    Your two are accidental: the first because of the word “archive” in the domain and the false association with archive.org, and the second is side effect of cookie isolation and incomplete javascript support.

    It is definitely not a free permanent infinite cloud storage for your hentai collections.

    • 4年前
    • 1 notes
  • Hello. I am developing an application that programmatically loads archived webpages from your service, but the captcha has destroyed this ability. Your service is one of the best. Is there an API available or is using your service in this way an impossibility? I am open for any discussion. Thank you.

    americansamizdat

    No, I can’t afford automated saving. The current hardware can barely handle manual. Of course, it is possible to create a paid API service to buy new servers. But… the current crisis of supply, payment, and trust has shown that limiting growth was the right thing to do. If the archive had gone that way, it would have to be shut down now.

    • 4年前
    • 2 notes
  • with the looming threat of russia cutting off their internet, how will the site manage their servers being accessible to the world, if they are located there?

    匿名

    They are not in Russia (although looking at energy prices, I would prefer that they were there).

    In any case, Internet fragmentation is already a thing: the Chinese segment has been around for years (the archiver is neither able to crawl most of it nor serve most people there), from now on there will be the Russian segment, what’s next? Islamic? … so yes, you are right, one day every site will land in one of the segments, being inaccessible from the rest. “the world” gets obsolete

    • 4年前
    • 4 notes
  • Why can't instagram pages be archived anymore?

    匿名

    no accounts left, all banned

    • 4年前
    • 1 notes
  • Just a note. In a project I am doing, I gained a lot of web scraping experience, so I thought of doing something similar to archive is.. and then I started to think about how to deal with child porn and jihadis uploading stuff and DMCA and "right to be forgotten" and all the content moderation stuff.... and... nope, not worth the trouble. So thanks for archive is!!! For me at least, scraping and archiving is the easy part, but actually serving other people stuff on my servers... nope. Thanks!!

    匿名

    Yes, it takes years to establish relationships with all sorts of government agencies so as to keep censorship at the minimum allowed level, and still there are regular glitches when an email gets the wrong place or trolls forge official letters.

    I bet any website with user generated content (Reddit, Imgur, Weibo, VK, …) has those issues.

    • 4年前
    • 1 notes
  • Please disable automatic translation on Facebook. It makes impossible to archive non-english facebook posts.

    匿名

    Examples?

    UPD: There is no way to disable automatic translation as a single option. One have to add every language to the do-not-translate list. Too many languages in that list —> account is banned. It is only issue of `m.facebook.com` which lacks “ver original“ button.

    • 4年前
    • 1 notes
  • Why has the URL "archive-li" changed to "archive-ph", and will this affect saved bookmarks at any time in the future?

    匿名

    This is temporary and only for some countries. All 7 domains work, so you do not need to change the bookmarks.

    • 4年前
    • 2 notes
  • i found many children porn images on archive. who i can report it?

    匿名

    email, or here

    also, every page has “report” button

    • 4年前
    • 2 notes
  • There have been several dozen bug reports in the last few days (remove a modal, expand comments, …). With a few exceptions, they’ve almost all been fixed, I won’t respond to each one here.

    • 4年前
  • Why do you say that escalation of the Ukraine conflict to nuclear war would likely lead to the end of the archive today project? What makes the project particularly vulnerable in that scenario? What can be done to mitigate that risk?

    匿名

    Because both copies are in Europe. There are no budget solutions in safer places like Latin America or Asia-Pacific region.

    • 4年前
    • 1 notes
  • Is there a way to archive multiple links at once? Maybe from an html file or something?

    匿名

    No. This is discouraged. I do not have enough computing power to handle all requests, even at the current rate

    • 4年前
    • 1 notes
  • Are you noticing a higher than usual amount of server crashes? I imagine the demand for your archive project to be incredible right now.

    匿名

    The job queue is longer, visits are as usual

    • 4年前
    • 1 notes
  • The website has been slow for some time when archiving Twitter pages, but works fine with other websites. Is there a reason for that? Thx!

    匿名

    1. There are too many pages from Twitter in the queue, which reduces their priority (if it wasn’t for this condition, it would slow everything down)

    2. Twitter API sometimes responds with “429 Too Many Requests” or other error, so it usually takes more than 1 attempt to capture the page.

    I would suggest refraining from saving pages from Twitter for now, especially those people trying to save dozens or hundreds of tweets

    • 4年前
    • 1 notes
  • Can you check why wip/praXv is giving 403 errors while archiving?

    匿名

    There is indeed 403: https://imgur.com/qcco0T2.png

    • 4年前
  • Could you please remove ads on /sMWcv ?

    匿名

    yes

    • 4年前
  • Could you please remove ads on /wZlIn ?

    匿名

    yes

    • 4年前
  • can i close the tab when it is already in queue?

    匿名

    you can, but you won’t know if the process has completed successfully or not.

    • 4年前
  • Could you expand /17wUS and links under same domain? Thanks.

    匿名

    yes

    • 4年前
  • Could you please click "続きを表示" on /Yf2HM ?

    匿名

    yes

    • 4年前
  • Could you click "Souhlasím" on 81uT1 to close the cookie box?

    匿名

    yes

    • 4年前
  • the website is down for me right now(in canada), unsure why. i tried reloading the webpage and even restarting the browser, same with any archives. nothing is accessible.

    匿名

    Half the internet is down, everyone is DDoSing each other

    • 4年前
    • 1 notes
  • a bit different than a usual question, but do you have a perspective in regards to the current Russia and Ukraine conflict.

    匿名

    If cross-border remittances stop working, recouping the site from donations and advertising may become a necessity to maintain sufficient balances on each side of the Iron Curtain.

    PS. Escalation of that conflict to nuclear war (my vision: Ukraine has already made A-bombs) would most likely lead to the termination of the project.

    • 4年前
    • 1 notes
  • Could you click "ADULT ONLY 🔞 18歳未満立入禁止" to reveal images marked NSFW on J3HAj, please?

    匿名

    yes

    • 4年前
  • Is it possible to fix GitHub profile archivals from redirecting to previous years (like s2hjU redirects to 2019 whereas most recent contributions are from 20222)?

    匿名

    They do not redirect, they click on “show more activity“, so you’ll get contributions from 2022 and from 2019 too.

    • 4年前
  • previously asked for fixing the gif for /LrOyy, which has now been resolved ,thank you for that. however, I noticed that for other reddit pages, it also just displays the spinning circles when the post is a gif. is there a global fix that you can push out, or will it need to fixed individual?

    匿名

    That page was fixed individually, I will deploy a global fix in few days

    • 4年前
    • 1 notes
  • the gif on /LrOyy is not being displayed properly, it only shows the spinning wheel

    匿名

    fixed

    • 4年前
  • Could you click <img alt="+" data-cmd="expand"> on 15Pdy to expand the replies, please?

    匿名

    yes

    • 4年前
  • Could you press "Zobacz więcej komentarzy" and then "Zobacz więcej odpowiedzi" on YNKhQ to expand all the comments?

    匿名

    In fact, it’s been pressed many times. You see, the page is very long, with many comments expanded. Unfolding everything is risky - it would take too long and eat up browser memory until it crashes.

    • 4年前
    • 1 notes
  • #3591 in queue. Besides that - any chance to host something similar on my own (private use)? Any code made public?

    匿名

    If you are looking for an enterprise solution, it falls to eDiscovery category.

    If you are familiar with coding and want to run something similar on premise: there is ArchiveBox, Diskernet, … and more projects can be discovered via codesearch

    • 4年前
  • Why Archive is not working lately? Unable to archive any page, as it gets stuck on a blank page when loading. There is only the small loading wheel and the numbers at the top left. This happens often.

    匿名

    What webpage you are saving ?

    There have been problems with Facebook and Instagram lately - they banned my accounts - so what you see may be retry attempts.

    • 4年前
    • 1 notes
  • Could you click <img title="Expand thread" alt="+"> on bvBWw to expand the replies, please? If possible, could you click <a title="Toggle infinite scroll">All</a> to load more thread?

    匿名

    1. yes

    2. it’s a bit risky, because the browser may out of memory

    • 4年前
  • Could you click [展開] on t6zMz to expand the comment section, please?

    匿名

    yes

    • 4年前
  • regarding the individual's response to the blockchain storage, you can describe how a one time price does not cover the actual cost of storing data in perpetuity.

    匿名

    In fact, financial theory solves this problem. After all, there are perpetual bonds you can buy for one time price.

    • 4年前
  • You wouldn't store the entire site on the blockchain, that's obviously unrealistic. It'd be an optional feature for those willing to pay to make a specific page extra safe. Ancillary to the site; not replacing it. The mantra of archivists is LOCKSS (Lots Of Copies Keeps Stuff Safe). The more platforms and formats the better. Your belief seems the opposite. You've rejected every suggestion for a back up so far. Basically arguing because no system is 100% perfect it's best not to use any of them.

    匿名

    Yes, the word “archive” in the title is misleading. The main purpose is to hold up ephemeral web pages for latecomers. Shots taken more than a few days ago have almost no visits.

    Lots of copies also require multiple independent copy operators. Bitcoin SV offers 35 operators, and a vague future at a very high price (plus it would require lots of third-party sites to bring the blockchain data back to html: the blockchain nodes do not do it).

    OK, we have one option for LOCKSS. More bad than good, but at least something. What else?

    • 4年前
    • 1 notes
  • Some blockchains store larger files. The reason most don't is less about technology and more that it's unnecessary for standard transactions. But there's Bitcoin SV which costs 7 cents per 100kb. You could add a "store permanently to blockchain" button on archived pages (EtchedPage does this but they aren't well-known). That way more pages will be safe if something happens to the site. And because it'll be optional you could take a fee while still allowing the rest of the site to remain free.

    匿名

    So we get 35 backups (the current number of nodes in Blockchain SV, which will probably decrease over time with the risk of one day reaching 1…and then 0 - it’s not mainstream blockchain after all) at $7,000,000 per 10TB disk replicated 35 times (appox $200 x 35 = $7,000 value).

    Perhaps we can find something more reliable and cheaper inside this space of three orders of magnitude.

    • 4年前
  • The backup doesn't need to be on a blockchain but it's not unreasonable for there to be one backup somewhere at least. A large chunk of internet history could disappear tomorrow if something happened to you or your servers. It's not planning your site survives the heat death of the universe to have some contingencies in the event of an accident.

    匿名

    There is a backup. But it is not a solution, backups did not help GeoCities or Google Code or many other projects.

    • 4年前
    • 1 notes
  • Can you please remove the "spin to win" popup at yonrC ? Thanks in advance!

    匿名

    yes

    • 4年前
  • On IF21u can you expand the collapsible that shows the post's image + original reddit post please? Thank you

    匿名

    Reddit is captured in the old design (as if it were chosen in user preferences or as in https://old.reddit.com/r/ProperAnimalNames/comments/n8oax1/kangaroo_mouse/). This allows more comments to be captured at the cost of larger images.

    • 4年前
  • Why is it that lately viewing archived pages also requires to solve a captcha often?

    匿名

    Mostly, no. The exception is datacenter IPs

    • 4年前
  • Could you click "Zobraziť celý popis" on wrWG3 to expand the description please? If at all possible could you make it so it would expand the descriptions on all future archives of this real estate listings website?

    匿名

    yes

    • 4年前
  • I haven't seen any ads on your site in a long time. What happened?

    匿名

    The ad agency decided that there was too much NSFW content on the site. Although there was a system to prevent ads from showing on NSFW, given the huge number of pages in the archive, there were quite a few (~400 in ~3 months) falsely classified pages

    • 4年前
    • 2 notes
  • Hello! I want to inform you that archive_ph is unreachable to me since today. tracert report: ... -> ***_ett_ua -> et54-100g_bb1-fra1_worldstream_nl [80_81_195_203] -> 109_236_95_220 -> 109_236_95_227 -> 109_236_95_227 reports: Destination host unreachable. Currently, I could access archive_ph only via proxy. Please fix the issue, if possible. Thanks.

    匿名

    109.236.95.227 is not mine and never has been. Try checking your DNS settings, they may be manipulated by malware.

    • 4年前
  • Please restore /yzwlj. It worked before, but now it says "not found". Thanks.

    匿名

    It is /yzwlJ, big J

    • 4年前
  • How does archive bypass hard paywalls? For example, on some news sites where the article is snipped/abridged server side?

    匿名

    Often the AMP version has more free content than the regular web page. For these sites, an AMP version is downloaded even if a regular version is requested (by replacing “www.” with “amp.” or adding “?amp=1″ to the end, etc) It sacrifices accuracy, but it gives people what they expect. Some browser extensions do the same thing. It can also be done manually.

    • 4年前
    • 2 notes
  • You said free archive sites don't survive long. But you also say you don't want your archive on a blockchain because you want it to remain free. Given you keep no backups (and have no plans to) aren't you guaranteeing the archive will eventually be permanently lost?

    匿名

    I don’t give guarantees. And I don’t trust the guarantees of others (like the clouds). One day it will be lost forever, just like your photos on Facebook, files on Dropbox, etc. It will happen long before the collapse of the universe.

    It doesn’t depend on whether the service is free or paid (I would be more suspicious of paid ones, since their project idea is subordinate to the search for profit).

    Blockchain may be the solution, but it is not designed to store large files. Those cryptoprojects that need large disks are not suitable for storing files, it’s just proof-of-work using a disk instead of a video card.

    I have nothing to offer in this area: you see, even NFTs don’t store whole files on blockchain because it is too expensive even for them. We have to wait for the next generations of blockchain.

    • 4年前
    • 3 notes
  • Do you happen to have a metric for how many captchas are solved each day? I would love to see a line graph of the amount over the past year

    gfair

    I don’t think such quantification makes sense because it would hide a variety of behavioral patterns. Most captchas are solved by “meat bots” - eager people who want to save thousands of pages and are willing to solve a captcha for each one. There are days when most pages are saved by one person on one topic - for example, one account’s tweets. And he is who solved most of the captchas that day. He’s probably the one who grumbles the most about captchas. But you have to slow it down somehow with a thousand pages to let another thousand people save their one page, right? Here, even buying servers will not do anything: it would benefit those 5-10 greedy submitters, not a wide audience)

    • 4年前
  • Could you investigate whether or not the Pale Moon web browser is being mistaken for a bot? I have been getting constant captchas even when trying to archive pages during different times and through different IP addresses.

    匿名

    What you think of as different IP addresses most likely fall into the same group that covers your entire selection of different IPs (e.g. “amazon aws“ or “some commercial vpn“).

    PaleMoon is not pessimized.

    Also, there are about 1000+ items in the queue almost all the time, which means a captcha is required for almost every submission (otherwise it quickly grows to 10000, which is a 3+ hour wait in the queue, which is unacceptable for grabbing short-lived pages).

    • 4年前
  • An article I'm trying to archive is redirecting to a different article/URL before saving. See /PliJA. The article I'm trying to archive is (2022/02/06/61ff0ff946163f66508b4604-html), and it's redirecting to (2022/02/08/62023d1646163f224c8b458c-html) and saving that instead.

    匿名

    Yes, it is a bug: the archiver scrolls pages down and on marca.com scrolling loads another article. I will fix it today. Thank you for reporting!

    • 4年前
    • 1 notes
  • Are these snapshots being saved permanently via blockchain technology? if not, why not? Everything can be erased, blockchain is permanent.

    匿名

    No. It won’t be a free service then. There are already some projects which do save web pages on blockchain.

    • 4年前
    • 1 notes
  • Why do I have to keep checking that I'm not a robot literally every time I archive? Did something happen that I'm not aware of?

    匿名

    Obviously, the information about whether you are a bot or not is not important. it is important to always be able to save pages. So when the system is overloaded, it shows more captchas. When it is underloaded it welcomes bots. There is no “cloud elasticity” which would magically buy new servers to handle the increasing load.

    • 4年前
    • 2 notes
  • Could you expand the texts for /gxn1y ? Thanks!

    匿名

    yes

    • 4年前
Next page
  • Page 1 / 85