- 3comments
- 331comments
- 93comments
- 114comments
- 200comments
- 2comments
- 44comments
- 35comments
- 357comments
- 553comments
- 143comments
- 25comments
- 397comments
- 350comments
- 127comments
- 13comments
- 45comments
- 2comments
- 10comments
- 36comments
- 22comments
- 616comments
- 13comments
- 22comments
- 126comments
- 17comments
- 20comments
- 11comments
- 39comments
- 49comments
I'd personally rather see Haskell become part of the options for https://arrow.apache.org/, but this is still a cool project.
Looks like a pretty cool community building great tools with care. Using a functional language for data transforms sounds like a sane idea, haven't played around with it yet but it's definitely on my list now.
However they claim using Haskell for data science is "fast", which doesn't really mean anything until you have numbers to show. A little benchmark with pandas and polars wouldn't hurt I guess.
Haskell tends to be C-fast.
That's not what I see reported, they say Haskell tends to have bad memory layout generally and takes a 5x or so hit to performance.
Depends what you're doing, but it really can have C-comparable speed. [0]
Unoptimised/naive Haskell might be that slow. But that's true of a lot of languages and isn't particularly interesting to me. Java is slow if you do everything the naive way, too.
[0] https://entropicthoughts.com/on-competing-with-c-using-haske...
Java ain't slow even if you do really stupid stuff. Hell, I would even argue that java is the most resistant to stupid code. Some way over-abstracted everything linked data structure will be faster in java than it is in C.
Haskell's most badly optimised type, is the String.
Java's most badly optimised type, is the String.
Both of them need a string-builder pattern, the default operators don't work around things to do the right thing for you. They expect you to understand how data works.
This is underway: https://github.com/duckdblabs/db-benchmark/pull/180
Cool cool cool!