Enjoying this a lot ! Possibly too much ;)
HN user
Morendil
[ my public key: https://keybase.io/morendil; my proof: https://keybase.io/morendil/sigs/snY2XSXWydrLeuoAXNBsC_Wm6DYlJfjqHXirA9LIoDc ]
"Research paper", ha. I came across this a few years ago, I call it the "Frankenpaper". See this gist for why I call it that: https://gist.github.com/Morendil/85336bf97211f9f31102ce2ee4e...
This needs some signal-boosting!
I would trust someone with a sharp eye like that to review my code much more than I would trust the author of the original article.
The PDF that requires an email can also be found here without giving one: https://pdfs.semanticscholar.org/feb0/0dd349ad701e1e4045282b...
Not in-depth across the whole set; I picked apart one or two before coming to the conclusion that there was something inherently flawed in the whole approach of modeling study of the phenomenon after the protocol for medical intervention. That won't work; you can't do a double-blind study of TDD vs placebo. It's a conceptual tool, and as such requires knowledge and skill.
"Treating" a convenience sample of students with a 2-hour training session and trying to measure "performance" afterwards isn't going to yield much insight into what goes on in the mind of a seasoned programmer who's used TDD exclusively for a while. In my experience, the effects of such practice most definitely carry over to programming where you don't even try to apply TDD.
What convinced me to try TDD was a combination of naivete and desperation. What kept me at it was the undeniable, if entirely anecdotal, feeling of wrestling back control over the code I wrote. It's a more complex set of techniques than just "writing the test first", but it's hard to tease them apart and introduce them into your coding independently of each other.
Part of it is a commitment to evolving a program in very small steps, at every step having a whole program that works. Part of it is a set of instincts for refactoring, so that this program also has just as much "design" (if such a thing could be quantified) as is strictly necessary for the time being. Part of it is a habit of framing the next capability as a tiny experiment, with the test being the experimental protocol, and focusing your mind on just those interactions within the code that are relevant. Part of it is a way of "slow debugging", of not jumping to conclusions when you encounter unexpected behaviour but drilling down into what made the behaviour surprising, and revising your mental model of the program.
Put like that it's clear that to "simply write unit tests", for instance to check a box in a process model, cannot possibly give you the same benefits as practicing the above interrelated set of skills. But it's also clear that it requires more involved guidance than "just write the test first". You can pick up a lot of it on your own, to be sure, just like many people can learn (say) guitar on their own and get all sorts of nuance in technique from focused practice.
But you have to love doing it, and that's not something easily picked up in academic studies.
I'm a big fan of TDD, but I quit reading this article the second my eyes came across "IBM System Sciences Institute" and the chart that accompanies it.
This is by now one of the most thoroughly debunked memes in our profession. I wrote a couple chapters about it: http://leanpub.com/leprechauns . I wrote a handy little guide for people to know just how much BS was involved in any one citation: https://plus.google.com/+LaurentBossavit/posts/aNKut1QV8pT
If that wasn't enough, there is now an actual negative research result on the so-called defect cost increase: https://arxiv.org/pdf/1609.04886.pdf
It should be quite clear by now that this type of argument doesn't do TDD any favors, it taints it by association with intellectual dishonesty.
It's normal to experience emotions related to what happens at work. It's called "being human". And dealing with emotions - putting words on them, accepting them and moving on - is a key skill and a part of growing up.
What varies, in my experience, is whether and how much it is acceptable to discuss (and thus process) these emotions in the workplace itself. Being able to say "I'm sad / angry / joyful about X" makes a world of difference.
Once I became aware of that I started being able to remedy it. To start with I actively sought and encouraged discussions of the emotional components of whatever work I was a part of. Retrospectives were a great way of having a structured framework for these discussions (as opposed to giving the impression that I wanted to psychoanalyze my colleagues or vice versa).
After a while I noticed, too, that management in some places actively preferred the dehumanizing effect of making emotions undiscussable, because it afforded easier control over people. I started avoiding these places and selecting jobs that accepted and expected me to behave as a human adult.
It's quite common, and known as "plagiarism". It will usually get people discredited quite thoroughly.
For more than you ever wanted to know on the "studies" about 10x programmers, see http://leanpub.com/leprechauns
I'm all too aware of these many citations. A few years ago, I went to the trouble of chasing down most of the papers and books, and evaluating how well each of them supported the claim. To put it mildly, I was underwhelmed.
For just one example, here's my treatment of Grady: http://lesswrong.com/lw/9sv/diseased_disciplines_the_strange...
It's not just me. Here's another author of a book aimed at software professionals who attempted some fact-checking, and came up short: http://www.sicpers.info/2012/09/an-apology-to-readers-of-tes...
I, too, used to argue for practices such as test-driven development, based on the supposedly firm knowledge of the "cost of defects curve". I changed my mind about the cost of defects when I saw how poor the data was. This is me in 2010: http://lesswrong.com/lw/2rc/coding_rationally_test_driven_de... and this is me two years later, recanting: http://lesswrong.com/lw/2rc/coding_rationally_test_driven_de...
However, I haven't (entirely) changed my mind about TDD and similar practices. I do still believe it pays to strive to write only excellent code that is easy to reason about. I like to think that I now have stronger and better thought out reasons to believe that.
I'd put the cost of defects claim in the category "not even wrong".
If someone made quantified claims about "the number of minutes of life lost to smoking one cigarette" I would refuse to take them seriously: I would argue that the health risk from smoking is more complex than that and can't be reduced to such a linear calculation.
This talk about "the cost of a defect" has the same characteristics. I don't mean the above argument by analogy to be convincing in and of itself, and I've written more extensively about my thinking e.g. here: https://plus.google.com/u/1/+LaurentBossavit/posts/8tB2RQoHQ...
But it's a large topic that quite possibly deserve a book of its own.
As for the history of software engineering, it's pretty much the same - to do it properly would entail writing a book, pretty much, and I didn't want to do it on WP unless I could do it properly.
I'm disappointed to see a book aimed at "professional" developers continue to spread outdated and debunked information, such as the NIST "study" adduced as evidence for the imperative necessity of "defect cost containment". See my post on the topic: https://plus.google.com/+LaurentBossavit/posts/8QLBPXA9miZ
Or again the Wikipedia page on the history of software engineering, which is frightfully inadequate. https://plus.google.com/+LaurentBossavit/posts/gpSwoWn4CBK
I'll add my voice to those that have already stated such a book shouldn't start by assuming the SDLC as a reference model: it embodies too many of those outdated assumptions. More in that vein in my own book http://leanpub.com/leprechauns
"Though it promises robot carers for an ageing population, it also forecasts huge numbers of jobs being wiped out: up to 35% of all workers in the UK and 47% of those in the US"
...not only are those forecasts not from the cited source (they're from an Oxford study), they're also not credible:
https://plus.google.com/u/1/+LaurentBossavit/posts/is8vMdyXb...
A fantastic book on the history of the search for gravitational waves: "Gravity's Shadow" by Harry Collins.
Point him to "Lessons Learned in Software Testing" by Bach, Kaner and Pettichord: http://www.amazon.com/dp/0471081124
Also, "manual testing" is a slightly unfortunate monicker for the activity we are discussing. It is bound to generate some degree of incomprehension or even hostility on the part of some people, for no foreseeable benefit. "Testing" will do. It is something you do with your head primarily, your hands being involved to pretty much the same degree that they are in programming (and we don't usually call that "manual programming").
As far as I can tell France universally uses "Flow A" (get your card back first, then your cash). When I started using ATMs, in the late 80s, I remember that "Flow A" and "Flow B" were about equally common, then in a relatively short span of time all the banks switched to A.
It surprises me, reading about it now, that it could be different in any other part of the world. That Flow A is the correct solution is not obvious, but it should be obvious to anyone who's studied human behavior and human error, which should be anyone involved in the design of ATMs. The form of "goal fixation" Jenny mentions is a very common pattern in human errors.
I own one of those. They're rad.
This.
Also, the (rarely surfaced) assumption that whoever is doing the hiring can tell a talented programmer from an average one. This is rarely the case.
The idea of this is not "try to break the robot"
Think of it this way. If you want to learn what constitutes strong chess play, will you learn best from playing a) yourself or b) a much stronger player?
Having a "collaborative" exchange with a chatbot is of the same strength as playing chess with yourself, for the purposes of investigating what "thinking" consists of.
The Turing Test is useful precisely when we are trying to "break the bot" as you put it; in fact, when the bot is pitted against a real human, who in that contest plays the role of the chess master.
Saying that Eugene Goostman "passed the Turing Test" is like crowning me World Chess Champion, based on the amazing record of beating 70% of a random sample of six year olds.
"if you ran into this robot in real life
You wouldn't ever "run into" Eugene Goostman in real life, because it lacks the kind of generalist problem solving ability that would allow it to insert itself into any "real life" situation - an ability that even six year olds possess. It literally couldn't even get out the gate.
Next time it's up, try asking a few of these: http://www.cs.nyu.edu/davise/papers/WS.html
That headline should read: "33% of human judges flunk the Turing Test".
Let's please upvote this to ensure it remains the top comment on this story. Head and shoulders above your usual science reporting.
After a couple years of fantasizing about owning one, I finally took action month before last.
I built a Prusa i3 from parts, single-sourced everything as a kit from RepRapWorld.com; that's more expensive, but much more convenient when you don't know what you're doing.
I didn't count the hours but it probably amounted to a couple weekend's worth of full-time days, all told, most of which consisted of tearing down and rebuilding something I'd already done which didn't quite work. Even so, the mechanical assembly was less trouble than the electronics - I plugged something in the wrong way and ended up blowing a fuse on my electronics, then blew a power supply while trying to "fix" that. Fortunately RRW were extremely helpful, I shipped the board back to them, they fixed the fuse and sent it back to me.
If you're not already well equipped for shop type work, remember to budget for tools.
Next time, if there is a next time, I might try the other popular approach, scavenging parts from discarded printers etc. to keep the costs down - and make the project even more challenging and time consuming!
I wanted to interest my kids (12, 15 and 18) in the project but that part didn't turn out as well as I'd hoped.
Once you've built the thing, there's still lots to do: fine-tuning, upgrading (for instance leveling the print bed is a pain in the ass, so I'm considering buying parts for an automatic leveling add-on), trying fancy plastics, and so on.
One unexpected difficulty was calibration. It's one thing to get the various parts working, to have the extruder actually extrude plastic and move around in the X,Y,Z axes - but to print objects you also need to choreograph all these motions just right: to move the right amount of X and Y while the extruder spits just the right amount of plastic. This requires repeatedly fiddling with the source code for the firmware and uploading that to the Arduino compatible board.
I'm still at this "tuning" stage right now. The first few days were heady as I went from "printing" a blobby mess of plastic thread, to a few decent-looking pieces. It's more of a grind now, as fine-tuning involves lots of guessing what could lead to an improvement, doing a print, and starting over; and these machines are SLOW.
Whether I can do something "actually useful" with it is still an open question. Even printing replacement parts still looks like it could be a challenge, there are so many things to get just right if you want parts that don't break, that print at the exact X,Y,Z dimensions, and so on.
I'll be honest, I'm now past the "honeymoon" stage where these difficulties were delightful, and experiencing some frustration, in particular at the slow speed of iteration. I'm a coder, I like being able to test tweaks in a matter of milliseconds. But that was the point - to get out of my software comfort zone and do something substantial involving hardware.
I've backed the Peachy from Kickstarter and hope to receive that in the next few months, it's a totally different principle (stereolithography like the FormLabs, but with many tweaks that keep the cost way down).
Building your own is a great experience, and whether I keep at it or not, I'm glad that I invested the time and money; now I know exactly what these machines can and cannot do, which gives me a great perspective on the hype from mainstream media and evangelists.
I've come to believe that these problems are unsolvable.
"Unsolvable" is strong. My take is that these problems stem inevitably from approaching the goal ass-backwards.
You're trying to force a particular testing framework on people who don't care about frameworks, supposedly in the name of "better communication". There's no reason to expect that to work.
The way I've advocated doing it, for years now, is to first sit with the people in question, discover how they communicate about business goals that the programming effort is expected to assist, formalize their notation as little as you can get away with and use that for acceptance testing.
This should be a process of active listening, not passive recording. The client should be gently nudged away from speaking in solution-terms, for instance.
This device purports to be able to measure glucose levels non-invasively.
It doesn't. The campaign page specifically disclaims that.
This seems neither new nor groundbreaking, however.
http://dst.sagepub.com/content/8/1/54.short
ETA: a 2011 review of the state of the art: http://knowledgetranslation.ca/sysrev/articles/project21/Ref...
I saw this campaign a few weeks ago, felt very tempted, ultimately decided it was prudent to wait for the final product to be available.
The discussion here has raised some interesting points pro and contra. Much hinges on how well the research state of the art supports the company's claims. My default assumption is that people will jump to conclusions, one way or another, without doing enough homework.
Googling for key words in Healbe's brochure has turned up this academic paper which seems to confirm their method for noninvasive BGL measurement is at least a promising path:
http://www.eejournal.ktu.lt/index.php/elt/article/download/4...
A similar article reveals interesting information about an attempt at noninvasive glucose measurement ten years ago:
http://www.engr.uconn.edu/~mam10069/Docs/NonInvasiveGlucoseM...
"Report shows the device had correlation with actual glucose level by only 35.1% and in some cases it gave potentially dangerous measurements."
(Not exactly a confidence booster, but ten years is enough to improve a lot.)
Half an hour of homework has substantially decreased my trust in people who say it's flat out impossible in principle to do what HealBe claim to be doing.
ETA: further academic paper links I've posted in other comments:
http://dst.sagepub.com/content/8/1/54.short
http://knowledgetranslation.ca/sysrev/articles/project21/Ref...
I don't mean to make fun of you, but: http://badassprogrammers.com/ basically demolishes that line of argument (and in the bargain confirms that Poe's Law does apply in Zed's case).
Particularly poignant is the way they redacted "autistic idiots" to "socially awkward idiots": if you're going play the "politically incorrect" card, play it with panache instead of being hypocritical about it.
You're quite right that people tend to miss the broader point.
Could be my fault, could be that more people - on this site in particular - ought to be familiar with pg's Disagreement Hierarchy: http://paulgraham.com/disagree.html
...or with its even more powerful cousin http://lesswrong.com/lw/85h/better_disagreement/
I take many people being bad at programming as evidence that programming is hard, not that it's easy.
We are in violent agreement. Note that "hard" is not the same as "abstract and intellectually demanding". Running a marathon is hard. Becoming a US Marine is hard. Many things are hard that do not primarily require quasi-mathematical skills.
For that matter, finding gravitational waves is hard but may be just as much about investing enough money (think space-based laser interferometry) or just plain luck.
The ways in which people suck at programming are much more diverse than just a failure of abstraction, otherwise we would all be writing Haskell.
These ways include failure to ask what the user or sponsor wants, failure to make sure we've understood what the user or sponsor said, failure to communicate with other members of the team, failure to question the things we learned in school and always took for granted. All of these are common failure modes in the biz, none of them are a failure of abstraction.
In fact excessive or premature abstraction is a widely recognized failure mode of software engineers, and the myth is responsible in good part for that.