Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    89comments
  2. Pirating the Pirates (mubi.com)
    211comments
  3. 12,000-year-old Göbeklitepe burials explain scattered bones (archaeologymag.com)
    17comments
  4. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    58comments
  5. 1996 chat room simulator connected to Win95 and System 7 web desktops (lolchat.rip)
    3comments
  6. Sonnet 5.5 (anthropic.com)
    383comments
  7. World Labs Is Joining AMD (worldlabs.ai)
    66comments
  8. Scientists solve 1840s space weather mystery (arstechnica.com)
    35comments
  9. What is the best shape of a city? Modelling effect of urban form on distance (sagepub.com)
    6comments
  10. Hijacking the PS5's RTMP stream (yashgarg.dev)
    65comments
  11. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    172comments
  12. Parley: Federated, decentralised chat that speaks plain IRC (mills.io)
    167comments
  13. It's Time to Investigate the AI Labs (calnewport.com)
    95comments
  14. Humanos – Help Building the Human Operating System (tryhumanos.com)
    3comments
  15. Does Reddit have an astroturfing problem? What the data suggests (petervijeh.com)
    128comments
  16. How to win a beer with high-dimensional statistics (jamiesimon.io)
    —discuss
  17. ESP32S3 cluster running 1.58-bit (BitNet) Language model (github.com/low-zi-hong)
    1comments
  18. California farmers are struggling to sell grapes as demand for wine drops (kqed.org)
    63comments
  19. Nvidia wants to put a watchdog chip next to every AI agent (cnbc.com)
    140comments
  20. Show HN: HN.watch – Videos of all Hacker News posts (hn.watch)
    78comments
  21. Cf: The Agentic CLI for the Cloudflare API (cloudflare.com)
    44comments
  22. First Steps of the PLC Organization – Independent Public Ledger of Credentials (plcred.org)
    19comments
  23. What reversing, modernising old games tells us about the economic impact of AI (isfine.org)
    16comments
  24. Behold the pawpaw (cbc.ca)
    8comments
  25. Updated Google Maps shows destruction of the city of Rafah (twitter.com/aliabunimah)
    107comments
  26. Deutsche Bahn "joke" is no longer funny (jonworth.eu)
    12comments
  27. Show HN: Destroy Any Website with Stickman (spritefusion.com)
    27comments
  28. OpenAI Says It Will Not Release Newest A.I. Model Over Safety Concerns (nytimes.com)
    9comments
  29. Joseph Szabo’s pictures of American adolescents (newyorker.com)
    43comments
  30. What heraldry and Japanese mon can teach about visual-identity generators (benovermyer.com)
    21comments

Ask HN: Who owns archive.is, and why are they trustworthy?

27 pointsby 6y ago
8 comments
I understand the need for anonymity when you're doing this due to the sheer amount of abuse reports, fake and real DMCAs, etc.

But why do people trust it? How do you know the pages you're archiving haven't been tampered with selectively to change history? This is just out of sheer curiosity, and I am not saying they do this.

This is made further interesting because of the following:

- Analytics from various Russian providers, instead of self-hosted (FYI: I consider GA to be equally privacy-violating as Metrika or Mail.ru)

- Large amounts of reverse proxies off questionable or bulletproof hosting providers

- Indefinitely doing this can't necessarily be cheap either at scale, who is paying for this?

- Demanding tracking or else blocking your access to the site, blocking any resolver that doesn't send the first 3 octets of your IP to them (edns-client-subnet)

- Explicitly tracking you in odd ways: they repeatedly load pixels/do DNS preconnect/preload from wildcard subdomains containing a cookied number, IP, country, tracking IDs. View any archived page and ^F "pixel.archive.is"

6y agoHN ↗

Archive.is isn't very important yet. I don't believe they warrant much concern re such questions. It doesn't yet matter very much if they're super trustworthy or not.

At their present scale, going through and manually changing (tampering with) saved content for propaganda (or similar) purposes, would have very little impact. More realistically, it probably has close to zero potential consequential impact. It'd be quite the chore for very little return.

If they become important some day, with dramatically greater scale of usage, then getting answers to these questions might be important.

If they eventually betray trust, they're trivial to replace. Other competing variations of archive.is exist now. It's a relatively easy service to create. Someone should probably challenge them just on the basis of how bad their ui & ux are.

At scale, if they begin abusing their position, it would become well known, they would get a reputation and it'd kill their service. The barrier to competition extremely low.

6y agoHN ↗

Why would you trust them? If you are trying to archive something, make sure to use multiple (separately owned) services so that you don't need to trust them.

6y agoHN ↗

Which services would you recommend? Last time I searched for free general-purpose website archival sites (for personal bitrot prevention), I could find only archive.is and archive.org.

6y agoHN ↗

Some alternatives that may be what you are looking for are archive.st, Webrecorder.io, FreezePage, and ArchiveBox. There is also perma.cc, which is a project of the Harvard Library Innovation Lab intended for academic usage.

6y agoHN ↗

This is made further interesting because of the following:

An additional concern: they've shown signs in the past of being capricious, or at least, easily annoyed by (subjectively) insignificant slights. They continue to block Cloudflare DNS users, last I checked. The "reason" is that Cloudflare doesn't send along the eDNS client subnet, as a way of protecting their users' privacy. [1]

I would argue this means archive.today / is can't be trusted to have the best interests of the community at heart. It's not a public service in the way that archive.org is.

[1] This bad behavior is actually mentioned in their Wikipedia article, along with the additional uncited claim that they throttle users to 20 MB of data per day, upon which they apparently ban your IP address. I haven't verified the latter claim. https://en.wikipedia.org/wiki/Archive.today#Worldwide

6y agoHN ↗

I didn't know about the Cloudflare DNS users being blocked. I started using Cloudflare's DNS two days ago, I just checked and it looks like I can't use the site anymore.

I can use it with Tor after more than a dozen different reCaptcha requests.

6y agoHN ↗

I share most of these concerns, especially since it’s primarily used in the “alt-right” sphere and could easily become a vector to sow discord, either by tampering with content or simply by mapping the communities of users visiting it.

Worth noting, it’s probably not that expensive to run. Most of the hosting services they use would be offering “unmetered” bandwidth, so the cost is probably fixed per month, likely under $1000.

6y agoHN ↗

Also, why do people use archive.is over web.archive.org? One (web.archive.org) is an actual library and gets all the legal protections that entails, while the other doesn’t.