What is the quality of software that gets shipped? What is the rate of defects and security issues?
What are the support costs once the software is shipped?
HN user
What is the quality of software that gets shipped? What is the rate of defects and security issues?
What are the support costs once the software is shipped?
Sample output for a significant github repository?
With this design, it’s possible to run native SQL selects on tables with hundreds of thousands to millions of columns, with predictable (sub-second) latency when accessing a subset of columns.
What is the design?
How well does this approach work with C++ source code - which is notoriously difficult to parse, given context-dependent semantics?
Tieredsort seems like a good balance between performance and complexity. Enough complexity (yet still relatively simple) to get very good performance.
No one had the motivation to fix it, including management. Many of the developers saw the problem as job security.
Anecdote:
I consulted for a large manufacturing firm building an application to track the logical design of a very complex product.
They modeled the parts as objects. No problem.
I was stunned to see the following pattern throughout the code base:
Class of the object
Instance #1 of the class
Instances 2,,n of the class
I politely asked why this pattern existed.
The answer was "it's always been that way."I tracked down the Mechanical Engineer (PhD) who designed the logical parts model. He desk was, in fact, 100 feet away from mine.
I asked him what he intended, regarding the model. He responded "Blueprint, casting mold, and manufactured parts." - which I understood immediately, having studied engineering myself.
After telling him about the misunderstanding of his model by the software team, I asked him what he was going to do about it. He responded "Nothing."
I went back to the software team to explain the misunderstanding and the solution (i.e. blueprint => metaclass, casting mold => class, and manufactured parts => instances). The uniform response was "It is too late to change it now."
The result is a broken model that was wrong for more than a decade and may still be deployed. The cost of the associated technical debt is a function of 50+ team members having to delineate instance #1 from instances 2,,n for over a decade.
N.B. Most of the software team has a BS (or higher) in computer science.
P.S. Years later, I won't go anywhere near the manufactured product.
In the United States, we all used to take a required course called Civics.
We learned how government and justice worked.
45% slower to run everywhere from a single binary...
I'll take that deal any day!
I think that there are a few critical issues that are not being considered:
* LLMs don't understand the syntax of q (or any other programming language).
* LLMs don't understand the semantics of q (or any other programming language).
* Limited training data, as compared to kanguages like Python or javascript.
All of the above contribute to the failure modes when applying LLMs to the generation or "understanding" of source code in any programming language.
"English as a programming language" has neither well-defined syntax nor well-defined semantics.
There should be no expectation of a "correct" translation to any programming language.
N.B. Formal languages for specifying requirements and specifications have been in existence for decades and are rarely used.
From what I've observed, people creating software are reluctant to or incapable of producing [natural language] requirements and specifications that are rigorous & precise enough to be translated into correctly working software.
IIRC, the percentage of Black students admitted to elite public schools was 10% in 1975.
What happened in the last 50 years to move that number down by a factor of 10 - literal decimation?
Software Engineering is only about 60 years old - i.e. the term has existed. At the point in the history of civil engineering, they didn't even know what a right angle was. Civil engineers were able to provide much utility before the underlying theory was available. I do wonder about the safety of structures at the time.
In my suburban middle school in Northern NJ, everyone (boys & girls) was required to take:
* Wood shop
* Metal shop
* Cooking
* Sewing
* Typing
I never saw an injury.
Learning to work with our hands safely was quite valuable.
AI won't replace software devs for a key reason.
The hardest part about writing software is the requirements & specifications.
Some ambiguous thoughts rendered as an LLM prompt is neither.
The output will only be as precise as the input allows.
Regarding Aro (new C compiler in the Zig toochain), it seems that a C compiler written in Zig shortens considerably the path to supporting constexpr in C using Zig's comptime capabilities.
There is precedent for no federal income tax in the United States.
Prior to the federal income tax (1913), all federal government expenses were covered by duties, tariffs, and levies.
This discussion of graph search is the best thing I've seen in quite some time. Discussing and highlighting the difference between related graph search techniques made getting to A* very intuitive.
To your point, profile your data as you would your code.
A sorted array of bit locations would represent a sparse bit set well enough to start, with O(N) storage and O(log N) access. Once the sets became large and/or dense, another data structure could be considered.
Moral of the story:
First, use an array. If an array doesn't work, try something else.
B-Tree
Because none of the COVID vaccines were vaccines in the traditional sense. They did not prevent infection nor disease.
The definition of 'vaccine' was changed to 'stimulates the immune system'. Likely to avoid litigation.
If the arrays of objects can conform to a single schema (across all scalar attributes), then make a second table to hold the objects in the arrays.
Now you have a two table schema with (at most) one join in a given query.
Check out Vertica. It does a great job at various forms of compression. In addition, DuckDB is an easy way to get started with efficient OLAP queries.
If the attributes are scalar, I would still suggest a column store that supports null values. Column compression will save you much space and give you excellent OLAP query performance.
As the schema evolves, simply add new columns.
For what kind of data? For what kinds of queries?
If the columns are scalar then consider a column store.
I'm reviewing models, at the moment. Model selection will depend greatly on the hardware capabilities at each school. Phi-3 could be a good starting point.
The project is an idea at the moment. My contact in Kenya has direct access to the Principals of the schools that our supported students attend.
My thought is that the teachers would not have to do much. Many of the students already know python and could do self-learning individually or in groups.
A flash drive with llamafile+models and documentation might be all that it would take to get them started - even offline.
Bonus: Using llamafile, the same binary distribution works on MacOS, Linux, and Windows.
Yes. However, it is uncommon.
I have been a Data Engineer, SWE, CTO, and a professional technical interviewer.
If you're willing to post your email address, I'm happy to send you a message, receive and review your resume, and do a brief assessment to determine what you might change to land a role in Seattle.
N.B. I have worked in Seattle since the dotcom era and have seen many of the market changes that have accurred in the city.
Btw, I support some Kenyan high school students and am looking at supplying a few schools with llamafile+models on flash drives for their computer science curricula.
What feedback have you received, regarding your resume? Which kinds of roles are you trying to secure?