Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Android 17 is the first since 3.x to add new APIs without releasing to the AOSP(grapheneos.social ↗)
    78comments
  2. Korea raises data breach fines to 10% of revenue(koreajoongangdaily.com ↗)
    7comments
  3. Cloudflare Quick Tunnels(cloudflare.com ↗)
    190comments
  4. Saving another 100TB of RAM with math (and Rust)(cloudflare.com ↗)
    8comments
  5. Apple releases iPhone Duo simulator and Xcode 27.1 beta(developer.apple.com ↗)
    23comments
  6. Cache-to-Cache: Direct Semantic Communication Between Large Language Models(arxiv.org ↗)
    3comments
  7. Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug(ledger.com ↗)
    33comments
  8. Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash(cactuscompute.com ↗)
    58comments
  9. The Implications of Linguistic Illegibility for LLM Security(arxiv.org ↗)
    8comments
  10. OpenJev(openjev.com ↗)
    231comments
  11. C++26: Trivial infinite loops are no longer undefined behaviour(sandordargo.com ↗)
    140comments
  12. Our brain evolved from two primitive nervous systems that merged: Study(newscientist.com ↗)
    33comments
  13. I vibed a proof of Conway's conjecture(overreacted.io ↗)
    159comments
  14. Border agents can search cellphones without a warrant or reasonable suspicion(lawandcrime.com ↗)
    64comments
  15. The first new cat species discovered in 100 years(nationalgeographic.com ↗)
    14comments
  16. Show HN: Ax-check.com – Can agents use your product?(ax-check.com ↗)
    21comments
  17. Inside ZCode: Silently uploading your Git history to the cloud(ferstar.org ↗)
    83comments
  18. Minimal Phone 2(minimalcompany.com ↗)
    97comments
  19. A heap overflow and SSO misconfiguration to compromise OpenAI internal repos(hacktron.ai ↗)
    193comments
  20. North Korean nuclear test sets off years of earthquakes(science.org ↗)
    133comments
  21. A search-and-inference database from scratch in pure Zig(antfly.io ↗)
    6comments
  22. Cekura (YC F24) Is Hiring(ycombinator.com ↗)
    discuss
  23. US Military had close call after using AI for hallucinated intelligence report(cnn.com ↗)
    215comments
  24. How SpaceX streamlined the Raptor engine(construction-physics.com ↗)
    11comments
  25. Warez: The Infrastructure and Aesthetics of Piracy (2021)(archive.org ↗)
    6comments
  26. "From Geometry to Algebra and Back Again: 4000 Years of Papers" by Jack Rusher [video](youtube.com ↗)
    discuss
  27. Show HN: Scry, programmable internet search w/ congestion pricing(scry.io ↗)
    12comments
  28. How to Write with an LLM(sockpuppet.org ↗)
    211comments
  29. Mathematicians Build Long-Awaited Graph Sandwich(quantamagazine.org ↗)
    15comments
  30. Jemalloc 5.4.0(github.com/jemalloc ↗)
    83comments

Designing Schemaless, Uber Engineering’s Scalable Datastore Using MySQL

59 pointsby 10y agoeng.uber.com
13 comments
10y agoHN ↗

Is it me or could they have done this way more easily by building some indexing and triggering functionality on top of Cassandra? Even two years ago when they started. Instead they built sharding, indexing, triggering and a Cassandra-like data model on top of MySQL.

10y agoHN ↗

What fun is that when you can abuse technology and then flaunt how clever you are for the abuse?

10y agoHN ↗

main lesson - for a new generation of what would at first look seems like OLTP business, the OLTP pieces like transactional triggers and transactional indexes aren't a requirement anymore. I.e. those requirements seems to go the same way - south - as the transactional consistency of search indexes had went several years ago.

10y agoHN ↗

The question I have is schema updates. The biggest pain I have had with things like Mongo is dealing with old data records.

Use case example for Uber:

1. In 2011, a driver joined. They made a bunch of trips

2. In 2012, Uber added more detail about the trip. Information not collected for the 2011 trips.

3. And so on, each year there are 'just a few changes'

Given the above:

In 2016, Uber want to run a query to reward all drivers based on some piece of information that was only present in 2014 on.

At this point the historical trip information from 2011 is in a significantly different format than in 2016.

In a RDB, at least the old columns are there - or if the db was migrated to a new schema ( a pain ) the issue of the missing fields was addressed.

But dealing with data in old formats was an Uber pain. And the lack of visibility into just knowing the schema used to generate that JSON object is a PITA.

God forbid if you had new code that never even knew about the old 2011 format.

Lastly, what happens if a bug slips through and some JSON field is missing, has odd spelling ( capitalization wrong ), etc.

I would love to hear about how old data is handled in schemaless.

My experience with MongoDB was less than pleasant.

10y agoHN ↗

I'm a huge proponent against "NoSQL" for 99% of use cases/scale and I'll admit that you would have this same problem even with a relational database any time that you add a column that you can't auto-populate based off some pre-existing knowledge.

10y agoHN ↗

God, what a name. A hyphen might be in order, as in:

  schema-less

...at first I read it as she-males.

10y agoHN ↗

Same for me. "Schemalessness" is even worse.

I have no idea why this word caught on instead of "aschematic", which is much easier to parse.

10y agoHN ↗

Odd that they chose MySQL, when they were previously using Postgres. In particular, Postgres' JSON support is so extensive (including indexing, which now is even more extensive [1]), and offers performance benefits over MySQL.

The advantage of MySQL in this situation is probably the support for multimaster replication.

[1] http://pgxn.org/dist/jsquery/

10y agoHN ↗

Is it just me or is the reasoning behind the switch from postgres to mysql very vague? They describe a sharded mysql database... Sharding postgres isn't necessarily any more difficult, instagram apparently uses it in a sharded manner with many shards. You'd think storing json in the pretty sweet jsonb column type in postgres would be a nice bonus for querying or indexing on.

I guess someone at uber must really like mysql, a good enough reason as any other I suppose. I'd love to hear about what other reasons as to why mysql turned out to be the choice here, as I've usually gone the other way (mysql to pgsql) for many of the great features and performance pgsql has.

10y agoHN ↗

They might've had really experienced mysql dbas. While postgres does seem to be as nice, in my experience, it's harder to find people who truly understand it, vs people who truly understand mysql. Not saying mysql is better, but there's more (deep) experience for it out there.