HN user

ihnorton

1,471 karma

Software engineer at TileDB

isaiah plus hnhiring -at- tiledb -dot- com

Posts25
Comments432
View on HN
devblogs.microsoft.com 3y ago

High-confidence Lifetime Checks for C++ in Visual Studio version 17.5 Preview 2

ihnorton
2pts0
github.com 4y ago

Dtype-next: a Clojure library to aid implementation of high performance systems

ihnorton
3pts1
deno.land 4y ago

Deno Foreign Function Interface API

ihnorton
2pts0
discuss.ocaml.org 4y ago

Why should I use OCaml?

ihnorton
3pts0
yyhh.org 5y ago

Writing C code in Java/Clojure with GraalVM

ihnorton
6pts0
github.com 10y ago

ParallelAccelerator for Julia

ihnorton
14pts0
github.com 11y ago

LuaJIT Language Toolkit

ihnorton
68pts15
www.tillett.info 11y ago

Ebola: What needs to be done right now

ihnorton
1pts0
www.newrepublic.com 11y ago

Harvard, Ivy League Should Judge Students by Standardized Tests

ihnorton
5pts0
www.lifl.fr 12y ago

Typed Lua: An Optional Type System for Lua [pdf]

ihnorton
54pts70
featherweightmusings.blogspot.com 12y ago

Rust for C++ Programmers Part 7: Data Types

ihnorton
93pts38
www.kaggle.com 12y ago

Higgs Boson Machine Learning Challenge

ihnorton
2pts0
www.cmu.edu 12y ago

Homotopy Type Theory receives $7.5 million DoD grant

ihnorton
3pts0
www.blog.juliaferraioli.com 12y ago

Parallel Julia on Google Compute Engine

ihnorton
2pts0
www.slideshare.net 12y ago

C++ Micro Services (light OSGI for C++)

ihnorton
5pts0
webapps.stackexchange.com 12y ago

Can I prevent Google from tracking which links I follow in downloaded PDFs?

ihnorton
1pts0
news.harvard.edu 12y ago

Dawn of a Revolution (Isaacson on Gates)

ihnorton
1pts0
www.nature.com 13y ago

Artificial light, sleep deficiency, and the rise of LEDs

ihnorton
1pts0
chronicle.com 13y ago

Synthetic Biology Comes Down to Earth

ihnorton
2pts0
www.nature.com 13y ago

Retinal vascular biomarkers for Alzheimer’s detection [pdf]

ihnorton
17pts2
www.kitware.com 13y ago

KiwiViewer: native 3D visualization for iOS (Android: WIP)

ihnorton
1pts0
www.donw.org 13y ago

clReflect: C++ reflection with Clang

ihnorton
2pts0
github.com 13y ago

CMake Build System for Python 2.7

ihnorton
3pts0
scienceblogs.com 13y ago

Aaargh! Physicists! Again!

ihnorton
5pts1
www.nytimes.com 13y ago

How Computerized Tutors are Learning to Teach Humans

ihnorton
3pts0

Numbers. I’ve read thousands of resumes over the past few months, screened dozens of applicants, and experienced a wide variety of weirdness and fakes both in resumes and on screen calls. Please note that I’m talking about raw “application box resumes”. Referrals and other semi-vetted sources don’t get this level of scrutiny.

I gave two examples of secondary sources, but what I’m really getting at here is that the numbers and noise are so, so high now (not to mention staffing firm fronts and foreign actors) that I usually need more signal than a solid-looking resume before investing even 30’ in a screening call.

TileDB, Inc. | Full-time | REMOTE | USA, Greece | https://tiledb.com/

TileDB is the database designed for discovery, built to organize, structure, and analyze any data. Our solutions for single-cell and population genomics are used by major pharmaceutical companies and research institutes, and power large public data collections such as the Cellxgene Discover Census. We are actively hiring for several roles building our unified data catalog, scalable computation, and interactive analysis platform.

- Infrastructure Engineer: Kubernetes, Terraform, Argo, Grafana, Prometheus, CloudWatch, GitOps; Golang, Python, C++, or Rust (GMT -8 to -4).

- Frontend/UI developer: Typescript, React; experience with high-performance/high-volume data and visualization applications. (GMT -8 to +1)

We are fully-remote, with optional co-working hubs in Cambridge, MA, New York, NY, and Athens, Greece. Apply today at https://ats.rippling.com/tiledb-careers/jobs or reach out directly (email in profile).

TileDB, Inc. | Full-Time | REMOTE | USA, Greece/EU | https://tiledb.com/

TileDB has recently announced a $34 million Series B fund-raise and is actively hiring for engineers across a range of roles (SRE, backend/distributed systems, database internals, and more). You will have the opportunity to work on innovative technology that creates impact for challenging problems in genomics, geospatial, machine learning, distributed systems, and many other areas.

TileDB Cloud is the modern database, allowing developers and scientists to capture, analyze, and share any data with any tool. We build on a broad foundation of open source, maintaining the TileDB storage engine, libraries for genomics (single-cell and population), geospatial (raster, point clouds, and more), a TileDB visualization engine extending Babylon.js, and much more ([github.com/TileDB-Inc/TileDB](http://github.com/TileDB-Inc/TileDB))

With TileDB, all data — tables, genomics, images, videos, location, time-series — is captured as multi-dimensional arrays. To supercharge this data, TileDB Cloud implements a serverless infrastructure delivering query execution, access control, data and code sharing, and distributed computing at global scale — eliminating cluster management, minimizing TCO, and promoting scientific collaboration and reproducibility.

Website: https://tiledb.com | GitHub: https://github.com/TileDB-Inc/TileDB | Blog: https://tiledb.com/blog

We are actively hiring for several roles including:

- Site Reliability Engineer (k8s, Terraform, automation, Prometheus, CloudWatch, GitOps; Golang, Python)

- Senior+ Software Engineer: Backend and distributed systems (Golang, CGo, k8s, Terraform, MySQL/MariaDB)

- Senior+ Software Engineer: Database internals (parsing, query planning, execution, distributed execution; Rust or C++ experience strongly preferred, other systems languages if paired with exceptional expertise)

- Front-end/UI developer: Typescript, React; ideally some additional mix of AWS/Azure/GCS platform experience, Golang, Docker, or other relevant skills.

Apply today at https://tiledb.workable.com or reach out directly (email in profile).

TileDB, Inc. | Full-Time | REMOTE | USA, Greece/EU | https://tiledb.com

TileDB has recently announced a $34 million Series B fund-raise and is actively hiring for engineers across a range of roles (SRE, backend/distributed systems, Python, C++, and more). You will have the opportunity to work on innovative technology that creates impact for challenging problems in genomics, geospatial, machine learning, distributed systems, and many other areas.

TileDB Cloud is the modern database, allowing developers and scientists to capture, analyze, and share any data with any tool. We build on a broad foundation of open source, maintaining the TileDB storage engine, libraries for genomics (single-cell and population), geospatial (raster, point clouds, and more), a TileDB visualization engine extending Babylon.js, and much more (github.com/TileDB-Inc/TileDB)

With TileDB, all data — tables, genomics, images, videos, location, time-series — is captured as multi-dimensional arrays. To supercharge this data, TileDB Cloud implements a serverless infrastructure delivering query execution, access control, data and code sharing, and distributed computing at global scale — eliminating cluster management, minimizing TCO, and promoting scientific collaboration and reproducibility.

Website: https://tiledb.com | GitHub: https://github.com/TileDB-Inc/TileDB | Blog: https://tiledb.com/blog

We are actively hiring for a number of roles including:

- Site Reliability Engineer (Golang, k8s, Terraform, automation, Prometheus, CloudWatch, GitOps)

- Senior Software Engineer: Backend and distributed systems (Golang, CGo, k8s, Terraform, MySQL/MariaDB)

- Senior Software Engineer: Python Data Science Tooling (Jupyter plugins)

- Senior+ Software Engineer: C++ Database Internals (database query planning/execution, distributed execution; Rust or other systems language ok if paired with exceptional expertise and willingness to write C++)

Apply today at https://tiledb.workable.com or reach out directly (email in profile).

TileDB, Inc. | Full-Time | REMOTE | USA, Greece | https://tiledb.com

TileDB is the database for complex data, allowing data scientists, researchers, and analysts to access, analyze, and share any data with any tool at global scale. Here are some active projects you could work on:

- vector search, utilizing TileDB and TileDB Cloud for seamless scaling: https://tiledb.com/blog/why-tiledb-as-a-vector-database (library: https://github.com/TileDB-Inc/TileDB-Vector-Search)

- single cell genomics: in collaboration with the Chan-Zuckerberg Initiative, we recently released TileDB-SOMA for single cell data, with APIs for both Python and R built around a common storage specification: https://tiledb.com/blog/tiledb-101-single-cell

With TileDB, all data — tables, genomics, images, videos, location, time-series — across multiple domains is captured as multi-dimensional arrays. TileDB Cloud implements a totally serverless infrastructure and delivers access control, easier data and code sharing and distributed computing at global scale, eliminating cluster management, minimizing TCO and promoting scientific collaboration and reproducibility.

Website: https://tiledb.com | GitHub: https://github.com/TileDB-Inc/TileDB | Blog: https://tiledb.com/blog

We offer the ability to work remotely for anyone with legal residence in the US or Greece. We have several open positions aimed at increasing TileDB’s feature set, growth, and adoption. You will have the opportunity to work on innovative technology that creates impact on challenging and exciting problems in genomics, geospatial, machine learning, distributed systems and database internals, and more.

We are actively seeking:

- Javascript library Engineer

- Senior Software Engineer: Backend (Golang, CGo, k8s, Terraform, MySQL/MariaDB)

- Senior Software Engineer: Python API (pybind11, Python, C++, CMake, scikit-build, conda)

- Senior Software Engineer: Python Data Science Tooling (Jupyter plugins)

- Senior Software Engineer: Build (CMake/C++, conda, wheels, and other packaging systems)

- Senior+ Software Engineer: Database Internals (C++, database query planning/execution, distributed execution)

Apply today at https://tiledb.workable.com!

TileDB, Inc. | Full-Time | REMOTE | USA | Greece | https://tiledb.com

TileDB is the database for complex data, allowing data scientists, researchers, and analysts to access, analyze, and share any data with any tool at global scale. We have just launched a vector search library leveraging TileDB and TileDB Cloud for powerful local search and seamless scaling to multi-modal organizational datasets and batched computation: https://tiledb.com/blog/why-tiledb-as-a-vector-database (library: https://github.com/TileDB-Inc/TileDB-Vector-Search)

With TileDB, all data — tables, genomics, images, videos, location, time-series — across multiple domains is captured as multi-dimensional arrays. Our vector search library and other offerings are designed to empower these datasets with extreme interoperability via numerous APIs and tool integrations across the data science ecosystem, eliminating the hassles and inefficiencies of data conversion. TileDB Cloud implements a totally serverless infrastructure and delivers access control, easier data and code sharing and distributed computing at global scale, eliminating cluster management, minimizing TCO and promoting scientific collaboration and reproducibility. TileDB, Inc. was spun out of MIT and Intel Labs in May 2017 and is backed by Two Bear Capital, Nexus Venture Partners, Uncorrelated Ventures, Intel Capital and Big Pi. Recent press-release for the 1.0 release of our Single-Cell API, supporting Python, C++, and R with AnnData and Seurat: https://www.businesswire.com/news/home/20230328005380/en/Til...

Website: https://tiledb.com

GitHub: https://github.com/TileDB-Inc/TileDB

Docs: https://docs.tiledb.com

Blog: https://tiledb.com/blog

Our headquarters are located in Cambridge, MA and we have a subsidiary in Athens, Greece. We offer the ability to work remotely for anyone with legal residence in the US or Greece. We have several open positions aimed at increasing TileDB’s feature set, growth and adoption. You will have the opportunity to work on innovative technology that creates impact on challenging and exciting problems in Genomics, Geospatial, Time Series, and more.

We are actively seeking:

- Javascript library Engineer

- Senior Software Engineer: Backend (Golang, CGo, k8s, Terraform, MySQL/MariaDB)

- Senior Software Engineer: Python API (pybind11, Python, C++, CMake, scikit-build, conda)

- Senior Software Engineer: Python Data Science Tooling (Jupyter plugins)

- Senior Software Engineer: Build (CMake, C++, conda and other packaging systems)

- Senior+ Software Engineer: Database Internals (C++, database query planning/execution, distributed execution)

Apply today at https://tiledb.workable.com!

Cxxwrap was broken for a while there, maybe a few years? Maybe it was cxx

Yes, Cxx: https://github.com/JuliaInterop/Cxx.jl -- it uses Clang to generate JIT'd interop thunks, and does neat things to enable that, like cross-language type inference and inlining (in addition to the REPL).

CxxWrap is a different thing, which AFAIK has been actively maintained for > 5 years now: https://github.com/JuliaInterop/CxxWrap.jl

It is similar to Boost.Python/pybind11/nanobind: bindings are written in C++ and compiled ahead of time into a module that defines Julia entrypoints. Those entrypoints take care of signature selection, translation to/from Julia objects, and lifetime bookkeeping.

The fact that Buck2 is written in a statically-compilable language is compelling, compared to Bazel and others. It's also great that Windows appears to be supported out of the box [1,1a] -- and even tested in CI. I'm curious how much "real world" usage it's gotten on Windows, if any.

I don't see many details about the sandboxing/hermetic build story in the docs, and in particular whether it is supported at all on Linux or Windows (the only mention in the docs is Darwin).

It's a good sign that the Conan integration PR [2] was warmly received (if not merged, yet). I would hope that the system is extensible enough to allow hooking in other dependency managers like vcpkg. Using an external PM loses some of the benefits, but it also dramatically reduces the level of effort for initial adoption. I think bazel suffered from the early difficulties integrating with other systems, although IIUC rules_foreign_cc is much better now. If I'm following the code/examples correctly, Buck2 supports C++ out of the box, but I can't quite tell if/how it would integrate with CMake or others in the way that rules_foreign_cc does.

(one of the major drawbacks of vcpkg is that it can't do parallel dependency builds [3]. If Buck2 was able to consume a vcpkg dependency tree and build it in parallel, that would be a very attractive prospect -- wishcasting here)

[1] https://buck2.build/docs/developers/windows_cheat_sheet/ [1a] https://github.com/facebook/buck2/blob/738cc398ccb9768567288... [2] https://github.com/facebook/buck2/pull/58 [3] https://github.com/microsoft/vcpkg/discussions/19129

In practice it is POSIX-only, which is not viable in an organization where Windows needs to be a first class development and deployment platform. Running under cygwin or WSL is a non-starter for a number of reasons (I've done that myself for tooling in the past, after which I would never impose it at an organizational level).

    Really annoyed by the borrow checker? Use immutable data structures
    ... This is especially helpful when you need to write pure code similar to that seen in Haskell, OCaml, or other languages.
Are there any tutorials or books which take an immutable-first approach like this? Building familiarity with a functional subset of the language, before introducing borrowing and mutability, might reduce some of the learning curve.

I suspect Rust does not implement as many FP-oriented optimizations as GHC, so this approach might hit performance dropoffs earlier. But it should still be more than fast enough for learning/toy datasets.

Skimming the docs, I was surprised to see that there appears to be no built-in support for foreign function calls to C libraries ("It's a bold strategy Cotton...")

This makes more sense in light of a comment by the author in a previous discussion about C ABIs [1]:

Virgil compiles to tiny native binaries and runs in user space on three different platforms without a lick of C code, and runs on Wasm and the JVM to boot. [...] No C ABI considerations over here.

[1] https://news.ycombinator.com/item?id=30705383

Sunsetting Atom 4 years ago

I'm just relying on sending blocks of code to the interactive Python terminal. [...] the latter requires .ipynb files which maybe you don't want to check into git.

I use `#%%` code blocks in the interactive window for this purpose: https://code.visualstudio.com/docs/python/jupyter-support-py

This allows a normal editor workflow, interactive execution, and linear history of executed cells. But the code blocks can still be executed top-to-bottom and exported as a clean notebook for sharing purposes.

Pijul 1.0 Beta 5 years ago

Displacing git is certainly impossible in the short term, but that is not the only conceivable goal. Building a sustainable community does not require git levels of success. A few strategic adoptions can go a long way, as shown by Ocaml, Nix, and other similar projects who have small but passionate and vibrant communities. The key there is a core group of users who prioritize a certain set of values over the advantages (ecosystem size, polish, job opportunities perhaps) of technologies with a larger userbase.

There are also ways to displace git without immediately replacing it. Many projects adopted git only after using git-svn for an extended period, giving teams more time to develop git skills without completely disrupting existing workflows. It looks like a pijul-git bridge is planned, which could facilitate gradual adoption. Beyond patch ergonomics, there are many other areas where a concerted effort to provide a better solution than git could yield a highly attractive tool. One that comes to mind immediately is management of multi-repository codebases and dependencies; Pijul (or one of the other git competitors) could provide a better and more cohesive way to manage multi-repository code changes than is possible with git sub[tree,module,...], which would be an adoption magnet. Zig appears to be pursuing such a mixed adoption strategy: they are providing both a new language and a set of tools which support C and C++ and solve several huge pain points -- such as cross-compilaton -- without requiring direct Zig adoption. (the tools have the benefit that they make Zig projects with C or C++ dependencies much easier to bootstrap, which should also allow faster ecosystem growth).

I just noticed that Deno merged their native FFI support: https://deno.land/manual@main/runtime/ffi_api

I had been watching some issues around this, but lost track, so I'm very excited to see it is available now! This makes Deno _very appealing_ for a wide range of tasks where FFI is a small but non-negotiable necessity (places where I would use Python's ctypes, for example tooling around C libraries; such tooling becomes much more complicated if another toolchain and compilation step is required before lib can be called from a script).

I would personally like to see less indexing of duplicate files! There are many things I’ve searched for which return 100s of results from independent checkin-uploads of big libraries like the Android SDK. It would be great if results were filtered by file similarity regardless of git history (if that is in fact the issue).

When I initially read the description, I thought this could provide a drop-in runtime-linked shared library stub, but looking deeper in the docs and examples it appears there is at least some setup code required on the client side?

Can this log the RPC calls around the target C library? Or potentially even replay calls in isolation? The latter could be expensive for non-trivial programs (require saving all synchronized memory state?), but might be more viable with binary diffs of the shared state if the client side doesn't modify the synchronized memory too much.

I think one could radically change the way Python objects work internally, and have the C foreign function interface (FFI) wrap every object passed to a C extension in an API/ABI-preserving facade

HPy is building an API abstraction layer which is designed to be used with both the CPython API and JITs. However, IIUC they are not proposing any changes to CPython itself, but rather to provide a smaller API surface and fewer JIT impedance mismatches when extensions are built against something other than CPython. The lead developer is a longtime PyPy developer.

https://hpyproject.org/ (https://github.com/hpyproject/hpy)

Nix is likely a nonstarter because as far as I can tell it does not natively support Visual Studio and MSVC.

I suspect Bazel was ruled out because it requires the JVM and it has limited open source uptake relative to CMake (huge open source userbase), and Meson (limited presence in open source scientific software, but adopted by GNOME and systemd).

This is interesting. It will certainly give a boost to Meson, which is in part a better CMake (the compilation model is very similar). It probably makes a lot of sense for SciPy given the alternatives, but from an ecosystem perspective it's not clear whether it is better enough to justify the churn. For example, caching seems like a secondary concern and outsourced to Linux-specific technologies. This is unfortunate given the implications of build time for individual and team (CI) productivity, as well as the environmental considerations of redundant compilation at scale.

I haven't followed Meson closely in about 3 years, but I also got the sense that Windows support sometimes lagged. If that's true, it's going to be a tough sell for the many large C++ projects who adopted CMake almost exclusively due to its support for Windows and Visual Studio.

Crystal 1.1.0 5 years ago

Elegant syntax. Crystal is highly inspired by Ruby and it has a lovely elegant syntax.

Agree 100%. I still haven't written any Crystal, but I read through parts of the code-base, out of curiosity, and it is some of the most pleasant compiler code I've ever read.

You should always be able to update CMake and have it continue to work.

CMake hard-codes the absolute path to itself all over the place. Any time a package manager that relies on symlinks (eg homebrew) updates the CMake package, every build that relies on it breaks. (last time that happened I ended up downloading the cmake.org build and recreating the symlinks because I couldn't afford hours of rebuild)

Julia did that for binary dependencies for a few years, with adapters for several linux distros, homebrew, and for cross-compiled RPMs for Windows. It worked, to a degree -- less well on Windows -- but the combinatorial complexity led to many hiccups and significant maintenance effort. Each Julia package had to account for the peculiarities of each dependency across a range of dependency versions and packaging practices (linkage policies, bundling policies, naming variations, distro versions) -- and this is easier in Julia than in (C)Python because shared libraries are accessed via locally-JIT'd FFI, so there is no need to eg compile extensions for 4 different CPython ABIs (Julia also has syntactic macros which can be helpful here).

To provide a better experience for both package authors and users, as well as reducing the maintenance burden, the community has developed and migrated to a unified system called BinaryBuilder (https://binarybuilder.org) over the past 2-3 years. BinaryBuilder allows targeting all supported platforms with a single build script and also "audits" build products for common compatibility and linkage snafus (similar to some of the conda-build tooling and auditwheel). One downside of "cross-compilation everywhere" is that it's not always well-supported by upstream libraries, so the initial effort to build a library may be higher if patches need to be developed, but that effort is arguably higher-leverage than a per-distro packaging approach (eg https://twitter.com/Blosc2/status/1395425736597585920).

All that to say: "make binaries a distro packager's problem" sounds like a simplifying step, but there are some big caveats. It has been tried before, both in other languages and in Python: the fact that conda and manylinux don't use system packages was not borne out of inexperience. One additional issue is that distro packages are only available with sysadmin consent in shared unix-like environments, which can be very limiting for end-users.