I hope that person is paying you for this positive PR spin!
HN user
btwillard
https://brandonwillard.github.io/
As others here are saying, the rules are the same for everyone in that anyone can effectively buy their freedom (even if only gradually over time). The practical difference is that most people can't afford it.
Obviously it depends on exactly what those opinions will end up being. It's hard to dispel the idea that every other big company isn't trying to find a way to do exactly the same things as the other big players, just somewhere else down the line. My anticompetitive-potential detector definitely went off as soon as I read that headline.
OpenAI function calling provides a means of producing valid JSON from a function signature, or a simplified version of a JSON schema definition; however, it has two big downsides:
1. it doesn't work with open-source models, and 2. sometimes the JSON it produces will not even parse.
A similar interface was recently merged into our open-source project Outlines. The approach can be extended to any open-source model and it can guarantee that the output will be parsable JSON.
Constructive feedback, bug reports, feature requests, and questions are greatly appreciated!
In my experience, that's always been the standard among Vim--and even some Emacs--users.
We've considered it, but I haven't started any work on it. Feel free to create an issue to track the status, if only to find out who else might be interested and/or working on it.
We provide this in https://github.com/outlines-dev/outlines.
Our project Outlines provides JSON output in a near optimal way that also works for all types of pre-trained transformer-based LLMs: https://github.com/outlines-dev/outlines
Our approach also extends to EBNF grammars and LALR parsing. There's an example of that in the repository. It builds off of the Lark library, so you can use existing grammar specifications instead of starting from scratch.
In case anyone is wondering, this is essentially a few complaints about the basic transpilation/source-to-source approach taken by Cython and then some promotion for Rust. It unfortunately mixes some general C/C++ complaints in there, too.
We also had an implementation of grammar-driven guidance around the same time: https://github.com/normal-computing/outlines/pull/131. I imagine many others did as well, given all the papers we found on the subject. The point of this and our ongoing work is the availability of very low cost guidance, which was implemented a while ago for the regex case and expanded upon with JSON.
FYI: We've had grammar constraints available in Outlines for a while, but not using the FSM and indexing approach that makes the regex case so fast. My open PR only adds that.
Ha, yeah, in a distant, but really fun, past!
The underlying approach can improve the performance of anything that requires the set of non-zero probability tokens at each step, and anything that needs to continue matching/parsing from a previous state.
Yeah, and our addition to all that is to almost completely remove the cost of determining the next valid tokens on each step.
With our indexing approach, it only costs a dictionary lookup to get the next valid tokens during each sampling step, so very little latency.
It means I have exceptionally high confidence that this will be the biggest thing to hit the economy, society and markets in the last hundred years for good and ill.
How does having ADHD and/or Asperger's mean that? Are you implying that those give people high confidence?
We're working on some of the DSL-related parts of this in https://github.com/aesara-devs
It sounds like you're saying you want to hire people who weren't laid off? If so, what's the reasoning behind this?
Eh, Steve!?
Not a complete replacement, but very cool and related: https://github.com/webyrd/Barliman
Arxiv's submission policy says: "Submissions to arXiv should be topical and refereeable scientific contributions that follow accepted standards of scholarly communication." (https://arxiv.org/help/submit)
If you're implying that he violated ArXiV's submission policy, then you're going to need to stretch the definitions of those words a bit, as you attempted to do.
I agree YouTube does not have this policy. I still find incorrectly claiming a proof of RH on YouTube distasteful, for similar reasons.
I get that you have all these personal opinions/takes, but ArXiV and YouTube don't appear to be justifying them.
I personally know a mathematician who has pointed out the mistakes to him. But also, Polson good enough at math himself that he should be aware of these points. I would criticize him just the same if he knowingly published false economics results.
OK, so where did they post these discussions with Nick so that we can all clearly see that he's a bad or stupid person, as you're implying? Your comments assume that this is all common knowledge or apparent, but it isn't.
Are we just supposed to take your word for it? Well, I happen to know Nick, and, from my experiences with him, I have no reason to believe any of the things you're saying and/or implying.
Do you think he's trying to poison the minds of young mathematicians?
At the very least, you're raising ArXiV and YouTube to a standard they openly do not meet.
Also, did you raise these points to him and he denied your evidence/proofs? Do you know someone who did and you're speaking for them?
It is bad practice to post incorrect results and not retract them when this is pointed out.
And it's good practice to take shots at people whenever you see their name?
Also, no, there's no "practice" that says you can't keep a mistake posted on ArXiV or YouTube. That just sounds like something you've made up to justify attacking someone.
And who's pointing this out to him?
He has continued to update the paper (as recently as last year), it has not been retracted, and he appears to maintain the proof is correct on his YouTube channel.
You know that ArXiV and YouTube aren't considered "journals" and don't have a similar requirement for "retractions". It sounds like you think he should hide his shame; otherwise he deserves your attacks.
We all make mistakes, and no one would really care if he admitted this. What makes the case notable is that he has enough mathematical training that he should really be able recognize his proof is wrong, especially after the errors are uncovered by others. Continuing to assert the proof is correct after errors have been found is bizarre.
People probably shouldn't, and don't, care because he's not affecting them in any way with his RH-related ArXiV and YouTube musing.
More importantly, where are there people showing him how his proof is wrong, and where is he outright denying their points/proofs? It seems like I'm missing a link or two, because I haven't seen any of these things, yet you're referring to them as though they're apparent and damning.
My guess is that he simply doesn't get feedback about this stuff, and he probably doesn't even care that much, because this is just a set of ideas with which he likes to work. I've seen no signs of a charged or "high stakes" math community engagement, and definitely not the kind that deserves these unwarranted personal criticisms.
It's not the discussion that makes you a crank. It's publishing a paper claiming you have solved something that you haven't even begun to understand. It's bypassing peer review. It's ignoring all the literature out there and claiming that is somehow a virtue.
You've effectively said that people can't post things on ArXiV unless they're up to your unstated standards; otherwise, they're just "cranks".
Also, no one is "bypassing peer review" by posting on ArXiV and/or YouTube, nor have I seen anyone claim "a virtue" of any sort. Where are you getting all this? From the abstracts?
It just wastes everyone's time. For example, this paper. It would take a lot of time to sit down and work through the math until I find a specific error in it. But the barren abstract and reference sections strongly suggest that would be a huge waste of my time and energy.
Well, since you've now admitted that you haven't read his work, it seems like everything you said earlier really must be coming from the abstracts alone.
It just wastes everyone's time. For example, this paper. It would take a lot of time to sit down and work through the math until I find a specific error in it. But the barren abstract and reference sections strongly suggest that would be a huge waste of my time and energy.
Who's time is wasted? The random people who volunteer to read his ArXiV submissions and/or YouTube videos? Really, the only waste of time I've seen is your ad hominem comment in a HN post that isn't even about the author you're blatantly criticizing.
Mochizuki pulled this stunt with the ABC conjecture. Wasted YEARS of mathematicians' time just to arrive at the conclusion it was all elaborate mathematical smoke and mirrors. The guy is an egotistical asshole with a messiah complex, and managed to burn a lot of PhD students by chasing his red herring. Not cool.
And that's why we must all attack Mochizuki whenever we see his name, right? Do you know Nick? Is he egotistical? Does he have a messiah complex? Are you just going on a tangent now? Are you just math trolling?
It's odd that people can't work on something in which they're interested and talk about it publicly without being labeled a "crank".
Sorry, there's nothing holy about the Riemann Hypothesis; anyone can add their input to the problem, especially on ArXiV and/or YouTube. You know, that's how discussions can start, and sometimes people want to discuss things with others.
It seems like you might think only certain people should work on these special problems, and only if they do it correctly according to you.
Your comment is just an ad hominem, so why don't you tell us why you're actually mentioning this? Do you think the math is wrong in the relevant material from the article because of his RH-related musings on ArXiV, or do you have an anti-RH-researcher bias that you want to share with the world? Are you just hating on someone?
It's a pretty big investment, but there's a sit/stand/lay workstation here: https://altwork.com/
Hello, I'm the person spearheading this Theano fork! Your comments match my experience with the old Theano very well, so I have to respond.
Apparently, the main new feature for Theano will be the JAX backend.
The JAX transpilation feature arose as a quick example of how flexible Theano can be, both in terms of its "hackability" and its simple yet effective foundation (i.e. "static" graphs). It's definitely not the main focus of the fork, but it is easily the newest feature that stands out at the user-level.
The points you raised about the old Theano are actually the main focus, and we've already made large internal changes that address a few of them directly. At the very least, nearly all of them are on the roadmap toward our new library named "Aesara".
The `Scan` `Op` and its optimizations are definitely going to change, and I have no intention of sacrificing improvements for backward compatibility, or anything else that would constrain the extent of improvements. I too have dealt with the difficulties involved in writing Scan optimizations (e.g. https://github.com/pymc-devs/symbolic-pymc/blob/master/symbo...) and am painfully aware of how unnecessary most of them are.
- The graph building and esp the graph optimizations are very slow. This is because all the logic is done in pure Python. ...
The most important graph optimization performance problems are not actually related to Python performance; they're demonstrably design and implementation induced. That is unless you're talking exclusively about graphs so large they reach the "natural" limits of Python performance by definition. Even then, a nearly one-to-one C translation isn't likely to solve those scaling problems.
For example, the graph optimization/rewriting framework would require entire graphs to be copied at multiple points in the process, and this was almost completely due to some design oddities. We've already made all of the large-scale changes needed in order to remedy this design constraint, so we're well on our way to fixing that. See https://github.com/pymc-devs/Theano-PyMC/pull/158
The rewriting process also doesn't track or use node information very well (or at all), so the whole optimization process itself can take an unnecessary number of passes through a graph. For instance, its "local" optimizations have a "tracking" option that specifies the `Op` types to which they apply; however, that feature isn't even used unless the local optimizations are applied by a `LocalOptGroup`. I've noticed at least a few instances in which these local optimizations are applied to inapplicable `Op`s on each visit to a node. Worse yet, within `LocalOptGroup` those local optimizations aren't applied directly to the relevant `Op`s, even though the requisite `Op` type-to-node information is readily available. In other words, optimizations could be directly applied to the relevant nodes in these cases and dramatically reduce the amount of blind graph traversals performed.
At best, a reimplementation in a language with a better compiler, like C, would largely amount to a questionable brute-force attempt at performance, and the ease of manipulating graphs and developing graph rewrites would suffer. With Aesara, we're going for the opposite. We want a smarter framework and _more_ focus on domain-specific optimizations (e.g. linear/tensor algebra, statistics, computer science) from the domain experts themselves, so code transparency and ease of development really matters. When we need raw performance in specific areas of the code, we'll pinpoint those areas and write C extensions, in standard Python fashion.
... When switching to TensorFlow, building the graph felt almost instant in comparison. ...
Last I checked, TensorFlow had almost no default graph optimizations, aside from some basic CSE and minor canonicalization and algebraic simplifications in the `grappler` module, so it absolutely should be instantaneous. More importantly, TensorFlow isn't designed for graph rewriting, and definitely not at the Python level where rapid prototyping and testing is possible outside of Google.
Otherwise, if you're talking about initially _building_ a graph and not calling `theano.function`, there are no optimizations involved. Latency in that case would be something entirely different and well worth reproducing for an issue. If what you were observing was the effect of calling `theano.function`, the latency was most likely due to the C transpilation and subsequent compilation. That's a feature that necessarily takes time, but produces code that's often faster than TensorFlow even today.
In summary, the changes we're most focused on right now are for developers like yourself who have had to deal with the core of Theano, so, please, stop by the fork and help us make a better `Scan`!
I'm only one person involved, but my primary reason for choosing Theano over TensorFlow has to do with the ability to manipulate and reason symbolically about models/graphs.
In order to improve the performance and usability of PyMC, I believe we need to automate things at the graph level, and Theano is by far the most suitable for this--between the two, at least.
You can find some work along these lines in the Symbolic PyMC project (https://github.com/pymc-devs/symbolic-pymc); it contains symbolic work done in both Theano and TensorFlow.