HN user

cbcoutinho

632 karma

PGP: 0xACC5190D2F6F160F

[ my public key: https://keybase.io/cbcoutinho; my proof: https://keybase.io/cbcoutinho/sigs/9rNeyVIv5DF69DmgsU-p7NOhsLzg9aA3Azo6u-q_3vI ]

Posts10
Comments265
View on HN

I have developed an open-source memory system for agents accessible over MCP, which makes it possible to access it via any coding agent locally (claude-code) or via mobile (Claude AI, Mistral AI, etc).

The primary storage mechanism is .md files stored as Notes in Nextcloud; however, since Nextcloud supports a rich ecosystem of apps such as documents, rss feeds, calendar, etc, your knowledge base can grow with you. All content can optionally be indexed and available via semantic search - powered by a Qdrant vectordb.

The biggest cost drivers of a system like this is the memory required to host the vectordb - I'm really curious how others are optimizing their knowledge base. Thanks OP for doing to work in summarizing these tools!

If you're interested either the MCP server or Nextcloud App frontend, please check out:

https://github.com/cbcoutinho/nextcloud-mcp-server

https://apps.nextcloud.com/apps/astrolabe

I'm working on a semantic layer for Nextcloud: https://astrolabecloud.com

The service is composed of two open-source services, namely a Nextcloud app (Astrolabe) and backend (nextcloud-mcp-server). I use the service as an MCP server across a number of apps, and others use it primarily for semantic search over large numbers of documents.

Both are open source, and I'm working on a managed offering, completely based in the EU, for individuals/teams that already use Nextcloud and want to be able to use semantic search across some or all of their documents.

Essentially your data stays in Nextcloud, and the MCP server backend keeps a vectordb in sync to enable semantic queries over your content. The number of supported apps is growing, including:

- notes

- deck cards

- files

- news items (RSS feeds)

- cookbook recipes

- contacts & calendar

And I'm adding support for other apps as I go.

MCP Servers are not simple wrappers around REST APIs, they can do much more and we will see more advanced use-cases surrounding MCP as MCP clients continually improve their conformance to the spec. Just wrapping REST APIs may be what MCP Servers are now, but that's just because MCP clients still need to catch up in supporting these more advanced features.

MCP Sampling (with tool calling) is one such example, where MCP servers will be able to generate responses using the MCP client's LLM - not requiring an API key themselves - including calling other MCP tools within the "sampling loop", and without corrupting the primary agent's context. This will generate an explosion in advanced MCP servers, although neither Claude Code, Gemini CLI, nor Opencode support this feature (yet).

Once MCP sampling becomes widely supported the pendulum will swing back and we'll see the boundary between MCP tools, async workflows, and generative responses begin to blur.

The Nextcloud MCP Server [0] supports Qdrant as a vectordb to store embeddings and provide semantic search across your personal documents. This enables any LLM & MCP client (e.g. claude code) into a RAG system that you can use to chat with your files.

For local deployments, Qdrant supports storing embeddings in memory as well as in a local directory (similar to sqlite) - for larger deployments Qdrant supports running as a standalone service/sidecar and can be made available over the network.

[0] https://github.com/cbcoutinho/nextcloud-mcp-server

I've been running openSUSE tumbleweed myself for years, and recommend Linux to like-minded power users. OP is preaching to the choir.

How do you all deal with (extended) family? This Christmas I spent time with my parents and the topic of Windows 11 came up again with all of its associated dark patterns.

What do you all do to help them out of this madness? Is Ubuntu/Fedora/etc the best option for seniors? My dad's entire career was in Silicon Valley 1.0 where Excel/Outlook was his bread and butter and feels married to Windows, but ever since leaving the workforce those skills are more of a hindrance than an asset.

Now that he's retired, he still uses Excel to plan vacations for example, but Windows is riddled with this BS and I am powerless to help him navigate this anti-consumer behavior. It's incredible that Microsoft is shooting their most loyal customers in the foot with this BS.

Do you all help your parents remotely? What kind of issues do you run into being your parents IT support?

This is why I built the Nextcloud MCP server, so that you can talk with your own data. Obviously this is Nextcloud-specific, but if you're using it already then this is possible now.

https://github.com/cbcoutinho/nextcloud-mcp-server

The default MCP server deployment supports simple CRUD operations on your data, but if you enable vector search the MCP server will begin embedding docs/notes/etc. Currently ollama and openai are supporting embeddings providers.

The MCP server then exposes tools you can use to search your docs based on semantic search and/or bm25 (via qdrant fusion) as well as generate responses using MCP sampling.

Importantly, rather than generating responses itself, the server relies on MCP sampling so that you can use any LLM/MCP client. This MCP sampling/RAG pattern is extremely powerful and it wouldn't surprise me if there was something open source that generalizes this across other data sources.

From the paper:

Most language models face a fundamental tradeoff where powerful capabilities require substantial computational resources. We shatter this constraint with Jan-nano, a 4B parameter language model that redefines efficiency through radical specialization: instead of trying to know everything, it masters the art of finding anything instantly. Fine-tuned from Qwen3-4B using our novel multi-stage Reinforcement Learning with Verifiable Rewards (RLVR) system that completely eliminates reliance on next token prediction training (SFT), Jan-nano achieves 83.2% on SimpleQA benchmark with MCP integration while running on consumer hardware. With 128K context length, Jan-nano proves that intelligence isn't about scale, it's about strategy.

For our MCP evaluation, we used mcp-server-serper which provides google search and scrape tools

https://arxiv.org/abs/2506.22760

I dove into using LLMs together with MCP servers for the first time this weekend. Absolutely incredible.

In addition to the code assistant, I configured a Grafana's MCP server with Cline, so that I can chat with an LLM while having real-time metrics and logs.

For context, I self host grafana in addition to a bunch of services on a raspberry pi. Simple prompts such as "why has CPU been increasing this week?" resulted in a deep analysis of logs/metrics that uncovered correlations I had never been aware of.

Incredible. I can only imagine what this will all look like in a few years

The six week cycle with down time in between chunks can be very effective. The crux is to make sure those blocks are properly scoped to make the best use of everyone's time.

This reminds me of the excellent book 'Shape up' by the team from Basecamp. They also champion upfront 'scoping' to make best use of those six weeks. Highly inspirational book for technically mature organizations.

I've been on openSUSE for 10+ years, and it is truly the gift that keeps on giving. The community super knowledgeable and responsive, and the distro is stable.

The biggest advantage openSUSE has compared to other mainline distributions is the openSUSE Build service (OBS). Contributing patches to existing packages is simple, and the build service also hosts custom packages in a personal rep - this let's me keep any custom packages up to date across all my systems. I believe it works with other distros as well so you don't even need to use openSUSE to utilise the service

Water and food scarcity in the Middle East, as a consequence of the changing climate and its reliance on the global food supply chain (wheat imports), is IMO one of the most important contributors of political uproar that lead to the Arab Spring, particularly in Egypt.

https://www.theguardian.com/lifeandstyle/2011/jul/17/bread-f...

Edit: possibly a better source describing some of the correlations between extreme climate events and recent global food shortages, and its adverse impact on Middle East and Northern Africa.

https://www.americanprogress.org/article/the-arab-spring-and...

Are you open to using a tool like direnv to manage env vars on a directory-wide basis? I used to also suffer from the same paranoia, but direnv has allowed me to take back control regarding directory-specific configuration.

You can configure it to set your git user name/email based on your working directory, for example. This is how I keep the separation you mentioned between personal and work on the same machine

I'd argue that the ability to unite people by assigning meaning onto things, places, and ideas is the single reason why humans have been so successful in advancing civilization to where it is today.

Edit: eschew -> assign

Nextcloud Hub 21 5 years ago

How do you deal with the high costs of running an RDS instance? I think almost $30 p/m is a bit high for a small single family NC instance, when the EC2 instance running the actual service is going to a fraction of that.

I do the same. Also remember the following keyword if you use nested .envrc files: source_up

If you don't use this in any nested envrc files, the settings don't carry over which means your git settings are not maintained. It is kind of an escape hatch as the parent files are not verified in the same way normal envrc files are, but I have found the trade off to be worth it

The same 'structure' that makes it easy to onboard new co-workers because they've seen the same project 'structure' before in the past. In that sense, the bottleneck in an organization is getting people productive as fast as possible, even that means using a cleaver instead of a scalpel.

Hypermodern Python 6 years ago

If the path to your executable is fixed, just put it in the shebang and you're done - makes everything way more explicit at the cost of some dynamic behavior.

An anecdote: Homebrew uses this method for shipping python executables.

I do this as well and consider CTE's just a macro that is expanded when executing the query - similar to an inline view.

I've run into issues before with systems where the result of a CTE is so large that it is actually advantageous to instead place it into a temp table to avoid re-fetching the data at each invocation. Sometimes that can be better than a CTE for this use case - YMMV.