HN user

hackeyed

22 karma
Posts0
Comments15
View on HN
No posts found.
[GET] "/api/user/hackeyed/stories?hitsPerPage=30&page=0": 500 Failed to fetch user stories
DIY Book Scanner 5 years ago

Looks incredibly serviceable and well engineered. I would expect reasonable and consistent results from the rig. The biggest question would come down to the cameras.

With these kind of rigs (two cameras, not computer controlled, no computer display) your big potential sources of error are either accidentally failing to trigger one camera or cameras losing focus on the page (especially if you are at something like the end of a chapter where there is often empty space in the middle of the page where the camera's auto-focus area is). His solution of using the IR remote should significantly reduce the issue of failing to capture on one camera. Cameras exist with manual focus settings, but they are often pricier or too old to reliably find one worth recommending to others. The CHDK alternative firmware for certain cheap Canon cameras generally adds a manual focus option for the less expensive cameras (though the individual features depend on who is making the firmware build you get).

Another option worth investigating is the newest Raspberry Pi camera modules with external lenses. Those should give you manual focus and the ability to build up an automated workflow you like around things like moving files around and any pre-processing you need. An ~9 mega pixel camera gets you 300dpi resolution on a full sheet of A4 paper, which is a lot more than most books.

DIY Book Scanner 5 years ago

Right, step 1 -> get page images, step 2 -> author images into book file. While OCR is obviously useful for search, a rotated phone screen will let you comfortably read a pdf book just fine unless you are talking about something like a textbook, in which case you probably wanted a tablet anyway.

I wrote up a guide on the authoring process using FOSS tools for some Digital Humanities folks a couple years ago: https://github.com/wikey/bookscan

It gives some background on the problem and covers a Scantailor (page crop, rotate, deskew), pdfbeads (compression, book metadata) authoring workflow, with pdftk for some general odds and ends.

For those that are digital subscribers and wish to suggest exactly this turning off of 3rd party tracking for logged-in users, this looks like the most relevant email address: membershipsupport@theguardian.com

Be sure to include your subscriber/membership number in that email, which you can find in your profile https://profile.theguardian.com/

These steps are generally split between the image processing ones, the OCR one, and a final step of combining/compressing pages into a single "bound" pdf/djvu file, at least if you are looking to use FOSS software.

For image processing, take a look at Scantailor (https://github.com/scantailor/scantailor/wiki), which will handle all the image processing steps for you and output images that are ideal for OCR.

I have not done OCR on mixed language text but I will say that tesseract has been under active development for years and does continue to improve.

The best FOSS options for binding all the processed images and OCR output into single files are djvubind (https://github.com/strider1551/djvubind) for djvu output, and pdfbeads (https://github.com/ifad/pdfbeads) fr pdf output. I tried to write up an outline of the whole process and how to use each of the tools here: https://github.com/wikey/bookscan

A lot of those tools have received little development in the past couple of years. They tend to do what they do well and reliably so don't let that put you off, though anyone interested in adding to the developer pool would certainly be welcome.

For more general information and especially background discussion, take a look through the DIY Book Scanner forum: https://forum.diybookscanner.org/

The trouble is that only the most engaged minority of users would be willing to pay for the service directly. What I'm interested to see is whether we can build services where the people who care most end up covering the costs for the rest of the users. So you pay $5 for a facebook-like service that then provides free services to everyone who knows you. You would have all the same funding problems as other public goods but there are plenty of them that make it work. Of course, even if something like a public radio or other public support model would fund a social network sustainably, you would still have the network effect problem in terms of getting people actually switched over.

Optical ballots provide a verifiable trail that can be used to audit the electronic counting machines. Seems that whatever system we use could benefit from some randomly selected mandatory audits/recounts at every level.

Many of the journals currently publishing these kind of Registered Reports only publish special issues, which might help with your concern about journals being given control over what work is done in a field, though it also limits how much of the literature can benefit from the new approach.

Another way of thinking about a journal's incentives is that, by moving the primary peer-review stage before a study has actually been conducted, you greatly expand the number of studies that a journal will interact with, which may provide a larger and more valuable role for entities like journals in world where actual publishing is a trivial matter.

The problem with using the blockchain for this is that there are legitimate reasons to withdraw such a registration ranging from misstyping (s/increase/decrease anyone?) to, depending on your kind of study, the inclusion of personally identifiable information. The solution currently being pursued is to expand pre-registration, a practice from clinical research, into other fields.

In pre-registration a trusted intermediary is used to store a read-only copy of your study design and materials, ideally including a pre-analysis plan that specifies the analysis you will run for hypothesis testing (so that we can cut P-hacking out of the picture at the same time). That intermediary can allow researchers to withdraw registrations while preserving a stub that shows everyone a registration used to be there and why it was removed.

That is how pre-registrations work on the Open Science Framework (https://osf.io), a cross-disciplinary FOSS web tool run by the Center for Open Science.

I realize you are replying to part of the discussion but your comment seems to be a nice point in support of the original article and I would like to see if you agree.

You mention one strand of western philosophy, "the Western Analytic (mostly anglophone) tradition" that believes its work is most closely related to math but then also point out that, while practitioners from that school are open to participation from non-westerners, they often face "conceptual and not just large linguistic barriers" at engaging with such contributions.

You seem to be saying that there may be value in these contributions but that we lack the cultural framework to understand and engage with them. If that is true, even though this is the portion of philosophy most consciously focused on mathematics-like deductive reasoning, doesn't that suggest an insufficiency in our current programs and perhaps argue that we need to expand our coverage if we are going to engage with practitioners in a global field?

They share some areas of focus but philosophy also overlaps with history, religion, literature, and logic. All ideas are human ideas and it is important to understand the historical and cultural assumptions that are baked into those ideas if you are going to really engage with them.

"Philosophers believe at least, that they are doing work more along the lines of mathematics than underwater dance (that is, searching for truth)"

Thinking that you can search for truth while ignoring how most of the world thinks about and addresses these issues is like thinking you can study urban planning without traveling to other towns. While there are some schools of thought that take such a tightly analytic view of philosophy those represent only a single strain of philosophy, and a highly Western continental one at that.

It is worth mentioning that Descarte thought he was doing just this kind of rational searching free of all philosophical conceptions but, for him, the proof of God's existence was obvious as a rational axiom using the same methods as his cogito. I think most of us in the more secular setting of hacker news would consider that portion of his rationalism to be the product of his cultural background.

Just as travel expands your understanding of where you are from, engaging with many traditions of thought is one of the only ways to actually become aware of the assumptions and biases inherent in your own.

Having gone through an undergraduate philosophy major and taken many classes on non-western thought, all offered by other departments, I strongly support the article's position. To have a class called "Ethics" that covers only the historical Western positions deprives students of the majority of the world's thinking on the topic.

What many of the comments here are missing by trying to separate "history of thought" from a more abstract conception of "philosophy" is that people do not develop their views of the world in a vacuum. Philosophy classes are not meditation sessions where students try and summon knowledge from the void, they involve reading the works of great thinkers in the field, analyzing their reasoning and engaging with their ideas.

The philosophers you read provide both the content and the tools you learn for use in your own thinking. As such, only presenting philosophers from a particular tradition biases the experiences of your students whether that bias is cultural as the article points out or even towards a particular school of thought within a culture, like the shift from pragmatics to highly analytic philosophy that took place in American schools during the second half of the twentieth century.

When I asked my department head why we did not, for instance, mention Confucius in our Moral Philosophy class, she explained that it is at least partially a bootstrapping issue. Because none of the faculty at my university had training in non-western traditions or spoke any non-western languages, they did not teach them nor did they feel comfortable advising graduate students doing their research on those traditions and in those languages.

Take a look at the graduate program requirements for the Ivy League philosophy programs and you will see that they all require students to know a second language, but the languages they are told to choose from are: French, German, Latin, Ancient Greek, or Dutch if you are really into Kierkegaard. You might also be able to get approval for Russian. The result is an echo chamber where we only teach western so we only hire western so we only teach western.

I doubt the article's suggestion of renaming departments will do much to change this situation but I agree that it would be a more honest representation of the materials taught and sympathize with the frustration behind their argument.