Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    52comments
  2. Pirating the Pirates (mubi.com)
    198comments
  3. 12,000-year-old Göbeklitepe burials explain scattered bones (archaeologymag.com)
    9comments
  4. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    45comments
  5. World Labs Is Joining AMD (worldlabs.ai)
    47comments
  6. Scientists solve 1840s space weather mystery (arstechnica.com)
    23comments
  7. Hijacking the PS5's RTMP stream (yashgarg.dev)
    55comments
  8. Flock Wants the Most Detailed Map of Its Surveillance Cameras Taken Offline (theintercept.com)
    44comments
  9. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    150comments
  10. Best of British Design (best-of-british-design.vercel.app)
    26comments
  11. Parley: Federated, decentralised chat that speaks plain IRC (mills.io)
    154comments
  12. It's Time to Investigate the AI Labs (calnewport.com)
    58comments
  13. What reversing, modernising old games tells us about the economic impact of AI (isfine.org)
    8comments
  14. Joseph Szabo’s pictures of American adolescents (newyorker.com)
    34comments
  15. 3D necroprinting: Leveraging biotic material as the nozzle for 3D printing (science.org)
    1comments
  16. Sonnet 5.5 (anthropic.com)
    346comments
  17. Cf: The Agentic CLI for the Cloudflare API (cloudflare.com)
    43comments
  18. Nvidia wants to put a watchdog chip next to every AI agent (cnbc.com)
    128comments
  19. First Steps of the PLC Organization – Independent Public Ledger of Credentials (plcred.org)
    18comments
  20. Does Reddit have an astroturfing problem? What the data suggests (petervijeh.com)
    97comments
  21. Show HN: Destroy Any Website with Stickman (spritefusion.com)
    26comments
  22. Launch HN: Vespper (YC F24) – SOTA Docx MCP (vespper.com)
    8comments
  23. What heraldry and Japanese mon can teach about visual-identity generators (benovermyer.com)
    20comments
  24. Updated Google Maps shows destruction of the city of Rafah (twitter.com/aliabunimah)
    38comments
  25. Who wrote Elizabeth I's most scathing letters? (smithsonianmag.com)
    18comments
  26. MongoDB CEO resigns to join Meta (reuters.com)
    253comments
  27. Pacing the Frontier is not the actual goal for AI labs (lesswrong.com)
    65comments
  28. When did Google get so weird? (sancho.bearblog.dev)
    1022comments
  29. GrapheneOS – When an app is slow (wirelessmoves.com)
    41comments
  30. ESP32S3 cluster running 1.58-bit (BitNet) Language model (github.com/low-zi-hong)
    —discuss

Show HN: Export HN Favorites to a CSV File

240 pointsby 6y ago
39 comments
6y agoHN ↗

Our plan is for the next version of HN's API to simply serve a JSON version of every page. I'm hoping to get to that this year.

6y agoHN ↗

Hi Dan,

That's great news! Is there a way to be notified (eg, via email) when this comes out?

Thanks.

6y agoHN ↗

If you (or anyone) want to be on an alpha-tester list, email hn@ycombinator.com and we can add you. Send a username and make sure it has the email address you want to be notified at.

6y agoHN ↗

Are there plans for an export tool, e.g. a user downloading all their comments and upvoted submissions? I tend to use the submission upvote button more than the favorite one, and an export tool wouldn't require a user API key for non-private info.

6y agoHN ↗

That should be an easy extension of the JSON API once we have it.

6y agoHN ↗

Any chance of some newer code being released, or the addition of a place to submit patches?

6y agoHN ↗

I'd love to do that someday. But it would be a lot of work.

6y agoHN ↗

That's great! Is there any plan of exposing authenticated content through the API too? Mainly talking about upvoted stories.

6y agoHN ↗

Thanks Simon. I'd originally written this script to export my DVD ratings from Netflix, since there's no API for that either. It was easy to adapt it to HN.

I wanted to show people that it's possible (and easy) to get to your own data!

6y agoHN ↗

This is smart. I'm adding this to my HN favorites.

6y agoHN ↗

Thanks dvfjsdhgfv. If there's sufficient interest, I can easily turn it into a Chrome extension.

(Edit: haha, I see what you just did there. A little recursive humor.)

6y agoHN ↗

Ok, now that's really smart!

Would a client using your API paginate to fetch, say, 50 pages? When I tried it using ?limit=50, I got a 504 error.

Thanks!

(Edit, never mind, I see you explain it in the readme.)

6y agoHN ↗

I had made it pretty quickly, the limit would be too large, thats 50*30, you may need to provide an offset and make a few requests.

6y agoHN ↗

Interesting that it is using x-ray, seems like x-ray is still using PhantomJS as the plugin, is PhantomJS deprecated? Would it be using Puppeteer instead?

6y agoHN ↗

What is HN's GDPR compliant way of requesting a copy of all stored data? Email dang?

6y agoHN ↗

Considering that there is no mention of GDPR in the HN FAQ or on the "legal" page, my guess is that their position is that GDPR does not apply.

According to Article 3 of the GDPR, it applies to:

1. Processing that takes place in the context of processors and controllers that are in the Union, regardless of whether or not the processing itself takes place in the Union.

2. Processing the data of subjects who are in the Union by controllers or processors who are not in the Union if the processing is related to offering goods or services to such subjects in the Union or the processing is related to monitoring the behavior of such subjects that takes place in the Union.

I don't know how HN is structured, but I've not seen any indication that they are in the Union, so #1 probably does not apply.

#2 applies if they are doing processing related to "offering goods or services to such subjects in the Union" or "monitoring the behavior of such subjects that takes place in the Union".

One of the recitals elaborates on the first branch of that:

In order to determine whether such a controller or processor is offering goods or services to data subjects who are in the Union, it should be ascertained whether it is apparent that the controller or processor envisages offering services to data subjects in one or more Member States in the Union. Whereas the mere accessibility of the controller’s, processor’s or an intermediary’s website in the Union, of an email address or of other contact details, or the use of a language generally used in the third country where the controller is established, is insufficient to ascertain such intention, factors such as the use of a language or a currency generally used in one or more Member States with the possibility of ordering goods and services in that other language, or the mentioning of customers or users who are in the Union, may make it apparent that the controller envisages offering goods or services to data subjects in the Union.

Does HN "envisage" offering services to people in the Union? Or are they a site that is merely accessible from the Union without envisaging offering services there?

There's a recital that elaborates on the second branch, too:

In order to determine whether a processing activity can be considered to monitor the behaviour of data subjects, it should be ascertained whether natural persons are tracked on the internet including potential subsequent use of personal data processing techniques which consist of profiling a natural person, particularly in order to take decisions concerning her or him or for analysing or predicting her or his personal preferences, behaviours and attitudes.

Does the data HN stores about its users satisfy this? And if it does, is the behavior being monitored taking place in the Union?

6y agoHN ↗

Could someone convert it to Python-script?

6y agoHN ↗

Part of the advantage of running JavaScript in your browser is that you might already be authenticated and it can use your session. But, fetching your HN favorites doesn't require authentication.

  #!/usr/bin/env python3
  import requests
  from bs4 import BeautifulSoup

  for p in range(1, 17):
      r = requests.get(f'https://news.ycombinator.com/favorites?id=app4soft&p={p}')
      s = BeautifulSoup(r.text, 'html.parser')
      print([{'title': a.text, 'url': a['href']} for a in s.select('a.storylink')])
6y agoHN ↗

Thanks!

One more question: what is the best way stop it when it will reaches last page?

for p in range(1, 17):

Actually p=17[0] is empty (as p=16 is maximum as for now).

Maybe, script should scrap pages from `1` to `infinity` UNTIL it detect next message on page[0]:

app4soft hasn't added any favorite submissions yet.

[0] https://news.ycombinator.com/favorites?id=app4soft&p=17

6y agoHN ↗

In Python, `range(1, 17)` produces the numbers 1-16. I hard-coded it just for your favorites.

A better way to solve it is to look at the `len()` of the results, and stop when it gets to 0:

  p = 1
  while True:
      r = requests.get(f'https://news.ycombinator.com/favorites?id=app4soft&p={p}')
      s = BeautifulSoup(r.text, 'html.parser')
      faves = [{'title': a.text, 'url': a['href']} for a in s.select('a.storylink')]
      print(faves)
      p += 1
      if len(faves) == 0:
          break
6y agoHN ↗

I think this one is a little cleaner. I used some of the ideas in sbr464's code.

    path = 'favorites?id=app4soft'
    while path:
        r = requests.get('https://news.ycombinator.com/' + path)
        s = BeautifulSoup(r.text, 'html.parser')
        print([{'title': a.text, 'url': a['href']} for a in s.select('a.storylink')])
        more = s.select_one('a.morelink')
        path = more['href'] if more else None
6y agoHN ↗

I had written a Python script to get saved (upvoted) links as JSON / CSV a few years ago: https://github.com/amjd/HN-Saved-Links-Export

I'm not sure if it still works as it too relied on HTML scraping. Perhaps I should update it to support favorites too.

Edit: Whoa, it's been 4 years already. I believe HN didn't have favorite feature at the time. That's why I used upvoting as my bookmarks system and created a script to export that data.

6y agoHN ↗

@amjd, thanks for sharing. I upgraded it from Python 2 to Python 3, but when I ran it, I got a 404 error on the `saved` endpoint. Does it work for you?

Edit: I see from the other PR it's called `upvoted`.

Edit 2: I changed it to `upvoted` and now I get a 200 OK, but the code crashed right afterwards on `tree.cssselect()`.

6y agoHN ↗

It's four years old so the page HTML structure must have likely changed. I'll take a look at it when I'm free and update the script. Thanks for your PR, I'll review and merge the two open PRs as well.

Btw, good work on your JS solution. It's great that it just works without requiring any download or installation. :)

6y agoHN ↗

oops - I didn’t realize that favorites != upvoted.

6y agoHN ↗

Can it will be done with "Scrape Similar" Chrome plugin?

6y agoHN ↗

Thanks for the tip. I gave the "Scraper" extension a try, and 1) I got an error, 2) it only seems to scrape 1 page -- it doesn't paginate (or, did I miss something?).

I used the jQuery selector `a.storylink`.

6y agoHN ↗

Is there any way to find most Favorited items on HN ?