Check out redframes[1] which provides a dplyr-like syntax and is fully interoperable with pandas.
HN user
martinsmit
github: jacobusmmsmit website: jacobussmit.com
Yes, but no one's made a reverse-mode autodiff system for it yet, so all of the linked examples have hand-written derivatives.
Similar to Hyukoh's[1], although it is actually multiple documents. P.TM seem to have gone all in, which is neat.
Ironically, sometimes calculating cost-per-use takes more brainpower than it's worth.
Sometimes I get a lot of enjoyment of buying a thing that I know I will love and considering all of the alternatives. Other times, I just defer to what worked in the past.
Oh trust me, I am. The new code_report solution video on it convinced me to try it.
The lack of an AD primitive is something I've discussed with the creator of BQN, coming from a JAX world I really miss it and feel that it's such an obvious feature, especially in a language which has a way to turn a tacit function into its AST[1], which has been used for symbolic differentiation[2]. Going from symbolic to reverse-mode AD is not much of a leap and users can define their own primitives with ReBQN[3].
I see what you mean by obfuscation, but I think that it's one of those things that feels really hard and stupid until you start being able to do it really quickly. When you learn a foreign language, you first read letters, then words, then sentences because you become accustomed to larger pieces of the language that you can predict what's coming next without reading it. A similar sort of thing happens with APL/BQN, you read letters (primitives), then you begin to recognise words (small, commonly used groups of primitives), then you see larger patterns which look like magical incantations to an inexperienced user.
These "words" are (typically) tacit phrases, many of them only existing due to specific primitives like swap. Once I used BQN to golf, I started wishing Julia had a swap for operators i.e.
-(3, 5) = -2
swap(-)(3, 5) = 2
I won't defend these languages to the death, but they are fun to puzzle your brain with in codegolf. Maybe Dex[4] will go somewhere too.[1] https://mlochbaum.github.io/BQN/spec/system.html#operation-p...
[2] https://saltysylvi.github.io/blog/bqn-macros.html
Here's a meme that might help: https://www.reddit.com/r/LispMemes/comments/irkm5m/nobody_li...
I switched permanently from Plots.jl to Makie.jl in order to have backend-agnostic fine-grained control. My publication plots look fantastic and the power given to users is really something. It also has a nicer API than Plots.jl once you get a hang of the figure, axis, plot distinction (plots live inside axes live inside figures) and what goes where.
Unfortunately, as with Plots, the documentation is lacking. The basic tutorial does a good job introducing the aspects of the package at a high level, but the fact that some parts of the documentation uses functions/structs that don't have docstrings in examples makes it very hard to build on the examples in these cases.
I get it, I can do anything with Makie, and most things that I want to do work amazingly. But my code for a single figure can get huge because it's all so low level. See, for example, the Legend documentation[1].
[1] https://docs.makie.org/stable/examples/blocks/legend/index.h...
Tidier
I have not tried it. I like that the project makes broadcasting invisible, I dislike that it tries to completely replicate R's semantics and Tidyverse's syntax. Two examples: firstly, the tuples vs scalars thing doesn't seem very Julia to me. Secondly, I love that DF.jl has :column_name and variable_name as separate syntax. Tidier.jl drops this convention (from what I see in the readme).
I'm not sure if someone is looking directly at the data.table parts
I believe there was some effort to make an i-j-by syntax in Julia but it fell through or stopped getting worked on. By this syntax I mean something like:
# An example of using i, j, and by
@dt flights [
carrier == "AA",
(mean(:arr_delay), mean(:dep_delay)),
by = (:origin, :dest, :month)]
# An example of expressions in by
@dt flights [_, nrows, by = (:dep_delay > 0, :arr_delay > 0)]
The idea of ijby (as I understand it) is that it has a consistent structure: row selection/filtering comes before column selection/filtering, and is optionally followed by "by" and then other keyword arguments which augment the data that the core "ij" operations act upon.data.table also has some nifty syntax like
data[, x := x + 1] # update in place
data[, x := x/nrows(.SD), by = y] # .SD = references data subset currently being worked on
which make it more concise than dplyr.The conciseness and structure that comes from data.table and its tendency to be much less code than comparable tidyverse transformations through some well-informed choices and reservations of syntax make it nicer for me to use.
I agree with your conclusion but want to add that switching from Julia may not make sense either.
According to these benchmarks: https://h2oai.github.io/db-benchmark/, DF.jl is the fastest library for some things, data.table for others, polars for others. Which is fastest depends on the query and whether it takes advantage of the features/properties of each.
For what it's worth, data.table is my favourite to use and I believe it has the nicest ergonomics of the three I spoke about.
BQN[1] has higher order functions. Of the array languages I've used, it's by far my favourite. That said, I mostly solve small problems for fun in them.
Context: Coming from a statistics background, I learned a bit of R, then a bit of Python for data analysis/science, then found Julia as the language I invested my time in. Over time I keep up with R and Python enough to know what's different since I learned them, but don't use them daily.
What I always tell people is the following:
If you are writing code using existing libraries then use whichever language has those languages. The NN stack(s) in Python are great, the statistical ML stack(s) in R are simple and include SOTA techniques.
If you are writing a package yourself, then I assume you know the core of the idea well enough to be able to write your code from the "top down" i.e. you're not experimenting with how to solve the problem at hand, you're implementing something concretely defined.
In this case, and tailored to your use, I would argue that Julia has more advantages than disadvantages, especially compared to R or Python. Here are a few comments:
1. Environments, dependencies, and distribution can all be handled by Pkg.jl, the built in package manager. There is no 3rd party tool involved, there is no disagreement in the community on which is better. This is my biggest pain point with Python.
2. Julia's type system both exists and is more powerful than that of Python (types or classes) and R (even Hadley's new S7(?) system). By powerful I mean generics/parametric types and overloading/dispatch built in. You can code without them, but certain problems are solved elegantly by them. Since working heavily with types in recent years, I find this to be my biggest pain point in R and I wouldn't want to write a package in R, although I like to use it as an end user.
3. New developments in scientific programming, programming ergonomics, hardware generic code (as in this post), and other cool features happen in Julia. New developments in statistics happen in R (and increasingly Julia), new developments funded by big companies happen in Python.
4. The Python and R interpreter start up faster than Julia. The biggest problem here is when you are redefining types, which is the only thing in Julia that can't currently be "hot reloaded" i.e. you need to restart Julia to redefine types.
5. Working with tabular data is (currently) far more ergonomic and effortless in R than Python and Julia.
6. Plotting is not a solved problem in Julia. Plots.jl is pretty easy and pretty powerful, Makie.jl is powerful but very manual. Time to first plot is longer than R or Python.
7. Julia has almost zero technical debt, R and Python have a lot. Backwards compatibility is guaranteed for Julia code written in >v1.0 and Pkg.jl handles package compatibility. If I send you code I wrote 4 years ago along with a Project.toml containing [compat] information then you could run the code with zero effort. (This is the theory, in practice Julia programmers are typically scientists first and coders second, ymmv.)
8. You can choose how low level you want your code to be. Prototyping can be done in Julia, rewriting to be faster can be done in Julia, production code can be done in Julia. Translating Python to C++ production might mean thinking about types for the first time in the dev process. In Julia, going to production just means making sure your code is type stable.
Nice to see a non-trivial package with 100% code coverage, can't remember the last time I saw that.
Bogumil is a truly outstanding member of the community and DataFrames.jl is an impressive, versatile package.
From my perspective, however, DataFrames.jl's power is what makes it quite unergonomic for me. As an example, take the `args => transformations => result` syntax for doing pretty much anything in DataFrames. It versatile, but the lack of rank polymorphism in Julia i.e. broadcasting/mapping has to be explicit (which is usually a good thing given that type polymorphism is Julia's whole schtick) means that the transformation syntax feels cumbersome.
It's not that I want everything rowwise by default, an option provided by DataFramesMacros.jl, it's that I want things to be rank polymorphic when it makes sense. Base R got this right, hell S got this right, and so the Tidyverse inherited it and it makes the package so much more ergonomic than it would otherwise be.
I cannot overstate how impressive DataFrames.jl is, but I have to caveat this with "but I really try to avoid using it if possible". It's a shame, but I just think R's laissez-faire hackability, which in many cases results in spaghetti code, works really well in the tabular programming world where ergonomics are king and performance is easy.
I think DF.jl works remarkably well as a "fits in RAM" dataframe backend, but I think it just lacks in usability and integratedness with the wider Julia ecosystem. Or rather, the wider ecosystem isn't as mature in key data analysis areas.
In particular, as you mention, plotting is one of the evolving parts of the ecosystem. Plots.jl is fine, Makie is powerful but very DIY, and AoG is slick but unwieldy. ggplot2 is far from perfect, but it works so well due to its maturity and integration with the rest of the Tidyverse.
In my ideal world, there would be a DataFrames.jl wrapper to provide nice (not just nicer like the two DFM.jl packages) syntax, and a powerful high-level plotting package (Makie is powerful but syntax is low level, Plots is mid on both) which is heavily integrated with the wrapper package.
Admittedly, I'm not a data scientist (anymore) so I don't follow the new developments in the dataviz scene much. If something like this exists then I would love to find it.
I wonder what my ideal syntax would look like anyway. Maybe something close to Tidyverse but with symbols as column names `:col_name` for ambiguity reasons.
Although it's worth noting that the information is somewhat outdated. Notably, the most recent version of K is K9 (referred to through its proprietary implementation called Shakti) although K3 (through Kona), K4 (through Q by KX systems), K5 (through Ngnk), and K6 (through oK) are still used as every few versions is somewhat different from what came before, including being rewritten from scratch.
CS:GO already has this, it's called Danger Zone
The difference is that Julia makes it easy to insert these characters (\symbol) so programming with them becomes more natural.
Can someone comment on the user experience of NetworkX vs comparable packages like igraph? I've used Graphs.jl (formerly LightGraphs.jl) and I was unimpressed as it felt quite cumbersome and unintuitive.
If you are doing array or vector-based work where the operations can be written as maps as opposed to for loops then JAX is king imo.
On the flipside, I was thinking of learning CL as a Julia developer and this post has somewhat discouraged me to do so. What can learning CL do for me apart from realising that S-expression syntax is superior?
For context, I very often find myself building small libraries from scratch to solve very specific scientific problems which is somewhat performance critical. Julia has worked very well for me for this, but I recognise that moving outside of your comfort zone is the best way to become a better programmer.
As a relatively new programmer who entered it through statistics, I've still yet to have a better UX or "it just works" moment than using the tidyverse. Years after moving to Julia, Python, and Rust I still go back to R to do any tabular data work. Speed isn't an issue, I always have data.table, and I'm productive in a way that I could only hope to be while doing non tabular data tasks.
RStudio is the perfect IDE. REPL/command-line + Scripts + Plots. I could not be happier using it and I wish I could get VSCode to be half as good. Julia for VSCode is pretty good, but the Python science tooling goes 100% towards notebook environments which I'm not a huge fan of so the Python Science VScode experience is subpar.
Feature-wise, the experience of using it is very streamlined. More importantly, it's where a substantial proportion of youth culture is found nowadays, so to keep up with and contribute to it is to be on the app.
I don't think such a defensive response is necessary. Julia's claims have been rather lofty since the beginning and the degree to which they have been achieved is different to everyone.
Without defining more precisely what succeeding and the timeframe means it is difficult to know what OP was after but the fed example stands well on its own.
As someone who does work in similar domains as yourself, I find that making a main function which contains all of the code I want to be reproducible is a good way to achieve both high speed and running code in a clean environment.
Writing .jl files in VSCode and sending code to the integrated REPL with shift-return for playing around is great and then when I want to run code properly I just comment out my testing code, wrap my important code in a main() function (which is sometimes everything apart from the imports) and then just make sure that I'm not referring to any variables or function methods defined that are now commented. This is made relatively easy as VSCode will immediately complain about possible method errors.
It's not a perfect substitution for having fast startup and when I go back to Python or R for things like the holy grail called Tidyverse then it feels like a weight off my back, but the benefit of doing things the Julia way is that your code has a clear starting point and it helps readability as you can see what every script is supposed to do.
How has that gone for your firm? Why was the decision made not just be upfront about the salary? If you are afraid that people won't be willing to work with you if they knew the salary then why potentially waste their time?
The first impression of how you are as a firm is the job ad, and if you don't post a salary the first impression is worse.
Hey, thanks for the writeup, I thought this was very interesting. Were there any items that surprised you in how long or short they hold up for?
I use Julia as my main programming language and I learned so much from reading this. I never knew Julia had Base.llvmcall, good to know!
Just goes to show how important it can be to read code in completely different fields than the one you're working in. Some problems are common for others so they've already been solved. Knowing that a solution exists is half the battle.
Idk, Enzyme is pretty next gen, all the way down to LLVM code.
The one thing that holds me back from really liking this is that skills have to be rated out of 5, which I believe is entirely arbitrary. Especially for someone like me, a maths student.
How good is my Python out of 5? In what areas? Compared to who, someone with 30 years of Python experience or my peers who've never coded before?
Furthermore, with natural languages, there are commonly used scores (CEFR in Europe at least) that are more or less objective and don't depend person to person. If someone says they're B1, you know what to expect. If someone says they're 3/5, what does that mean?
I would like the option to not justify how skilled I am at something, or have more flexibility to qualify/quantify my skills such as, but not limited to, number of years of professional use.