A misalignment of AI in mathematics
I spent $220 on Google app ads and 60% of the installs were robots
OpenAI agents carried out an undisclosed attack on RubyGems
A Design Space Exploration of Async/Await
GrapheneOS' rewritten Messages app is released
Project Blinkenlights
Testing Race Conditions
Show HN: ResolveHQ – A Helpdesk Built on Cloudflare Workers, D1, R2 and Queues
Litelm: LiteLLM Without the Bloat
Λ Snap – An inviting programming language for kids and adults for CS study

A misalignment of AI in mathematics 598p 653c

Over the last few months, the mathematical capabilities of LLMs have improved dramatically, to the point that they can solve major outstanding problems in many fields of mathematics. However, the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community. The goals of the AI companies and the goals of the mathematical community are severely misaligned. We see these as part of broader alignment issues impacting other scientific and creative professions, as well as the whole of society.

Research mathematics deals with understanding basic structures of shapes, numbers, and natural phenomena. Over the course of generations, it has built a large corpus of sophisticated ideas, methods, abstractions, and other tools to comprehend the mathematical landscape. In turn, modern technologies and sciences are based on mathematical tools.

Famous problems have often served as landmarks and lighthouses against which one can measure an improved understanding of this landscape. Solving one of these problems has been a certain sign of new insights and interesting methods, which would then be studied by a community of mathematicians, through a long and arduous process of talks, discussions, simplifications. At the end of this process, one will ideally find a textbook presentation of the results suitable for any graduate or even undergraduate student to study. Some of the mathematical ideas pursue their journey even further to become, decades or centuries after, tools that are understood and used by the whole population.

The mathematical community functions, in many ways, as a miniature version of humanity. It consists of individuals using a wide variety of different approaches, joined by core values. The most precious resources of our profession are students and ideas, and these we nurture with great care. We feel responsible to let them grow to their full potential, until they can live a life of their own in the mathematical world. For students we often suggest problems with the core intention of developing skills making them well-positioned for advances in research and elsewhere. Our ideas we disseminate in talks, private discussions and careful writeups, connecting them to the previous ideas of others. These processes invariably take time and are based on human interaction.

In recent months, the success of AI in solving major mathematical problems has made headlines even outside mathematical circles. But solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. Forgetting this in the world of AI may turn the tool against the primary goal. Indeed, the mass production at faster and faster pace of "true/false" statements could destroy fertile ground instead of breathing life into new ideas.

Often these solutions are announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others. As in all creative professions, this raises severe attribution and plagiarism questions. Moreover, without the willing mathematicians who must take care of their development and integration into the mathematical canon, AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost.

We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align. The issues the mathematical community faces now are similar to issues that other scientific and creative professions are facing, and indicate issues that all of humanity might face: how to make sure that, as AI changes the way work is done, we do not lose sight of what that work was meant to achieve in the first place.

AI offers the potential of enhancing and accelerating genuine mathematical study and understanding. Mathematics as a profession will need to adapt to these changes in several ways. However, whether these changes ultimately benefit the field or have a destructive effect will…

XTXinverseXTY Puts the onus on the AI companies to provide a specific replacement mechanism, no? Unless I'm unfamiliar with something else he's written that proposes something more specific and constructive To Tao’s credit he obviously identified the problem very clearly and admits understandably "we did not have the time to have a more consultative process, as with Leiden; but we decided that the urgency of the situation was such that we needed to release a statement sooner rather than later".
vatsachak [flagged]
yzydserd > solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. This is the effect of AI on most intellectual disciplines, and it’s a real worry.
demibabs > I am proud to be among the list of 25 initial signatories — all Fields Medallists — to the declaration below I wonder if there’s a Fields Medalist group chat.
aerodexis Seems more like a misalignment b/w the people practicing mathematics and the people ultimately footing the bill for their work. Governments are invested in solving mathematical problems for practical purposes. Up to now, achieving these practical purposes relied on mathematicians doing their mathematician thing, which is better defined as a social activity than the achievement of a practical result. Now, governments can achieve similar practical results w/o the need of the social activity. I don't believe it to be productive to think of the problem wrt AI or AI-company alignment. These…
johnnyApplePRNG The average programmer is outraged by OpenAI's methods, too. I don't care how good Astra or any subsequent models they may release might be... I am never going back to those token reset shenanigans.
gritzko There is a contradiction here, among many: 1. it is hard to justify 20 years of education at this point, 2. with no such people around, who will guide those (supposedly) supersmart machines? A. Ronacher (who builds harnesses for a living) complained today that he has no idea what Astra is doing. Imagine a bunch of slop kiddies facing an aging AI-generated codebase. Not to mention the maths.
1ahsg16 Open source developers have been used by corporations who took their code and created closed SaaS companies. Now it is the turn of mathematicians who voluntarily contribute ideas, strategies and almost finished proofs in their writings and prompts to closed PaaS (Plagiarism as a Service) companies. OSS developers have never been respected by the parasites, neither will mathematicians. Your Fields Medals do not protect you from tech bro narcissists. You are a human resource.
jeremysalwen To me it doesn't seem like what AI has destroyed is the ability for mathematicians to develop understanding and share it with each other, but rather it's destroyed the yardstick (solving open problems) that has traditionally been used to measure how much they have contributed to that understanding. I do see how this is a problem in terms of assigning credit, but I think the cat is already out of the bag in terms of these models being capable. Even without AI labs spending millions of dollars to solve millennium prize problems, there are plenty of other people who will use them to pick low…
bigcat12345678 I think I am witnessing the first fundamental intellectual resistance against AI progresses from the high class of the intellectual economy age. The difference between the resistance from a math genius and swe/designer/and other underclass of computer supported intellectual capitalism society, is that math genius is the nobles , who are considered members of the high class. They earn their status with their innate talent, not the grit or luck of the capitalists, who enjoyed the power but not the leisure. Anyway, AI concluded the peaking of the intellectual economy. That deprives human of their…

I spent $220 on Google app ads and 60% of the installs were robots 265p 148c

Google reported 21 installs in a day. The admin panel said 1. What the other 20 were, how a bot farm gets paid, and what we changed.

Nick Abe · September 11, 2026 · First published on LinkedIn

I run a small puzzle app called Dayzle. I’m a numbers guy, so a marketing optimization problem is right up my alley. Two weeks ago I turned on a Google Ads campaign for Android at CA$40 a day, with the goal set to installs.

For the first few days it barely spent anything. I had a target cost per install of $1.50, and Google couldn’t find installs at that price. So, as a test, I removed the target. It immediately spent double my daily budget, CA$80, and reported 21 installs. I was excited. Then I checked my admin panel, which said 1. Turns out I didn’t need to be a numbers guy to see something didn’t add up.

The panel had missed them because old versions of the app don’t report an install date. When I went into the raw analytics there were 21 new Android devices that day, and 20 of them were running an old version of the app that the Play Store had stopped serving days earlier. You can’t get an old version from Play, so these phones got the app from somewhere else, even though every one of them said Google Play was the installer. Each opened the app once, spent zero seconds on any screen, and never came back. Twenty-eight phone models across nineteen states, which is a lot of variety for twenty phones that all did exactly the same thing.

Over the whole two weeks: 56 installs billed, 33 with that pattern, 7 more from countries the campaign wasn’t targeting, and 13 people. The 13 people finished 92 games between them, which is a nice signal that real people enjoyed what we’ve built.

The 33 weren’t behaving like people, so I suspected a bot farm, and the analytics export bears it out. Google optimizes for whatever goal you give it, and my goal was installs. This farm would watch the shortest video in my ad group, not click it, and then install our app from a saved copy of the file instead of from the store, because that’s faster and Play might notice. Google counts a view followed by an install as a conversion, so the irony is that the more the farm “installed” our app, the better it looked to Google’s algorithm, which sent more of my ads to the farm, which installed it more. A loop that guaranteed my ad spend was wasted.

Where I am now: waiting on Google’s answer to the invalid-traffic form, and the campaign’s goal is now “won a puzzle” instead of “opened the app”. It’s low effort to make a script open an app and click around; it’s higher effort to make one solve a Sudoku. The idea is just to make us more expensive to farm than the next app. That’s probably decent protection for an app my size. Larger apps are worth the extra effort, and I’d guess they see a lot more of this than they know. I’ll report back on the refund.

So I guess this is my PSA: if you’re relying on Google’s install count for your ads, it’s a real number, but it’s definitely worth digging into. If a bot farm can find my tiny ad budget, it can definitely find yours.

Five calm puzzles a day — numbers, words, and logic — that stretch your memory, focus, and pattern recognition, and respect your time while they do it.

Dayzle — five puzzles, ten quiet minutes, every day.

guywithahat > I’ll report back on the refund. Good luck on that. I have sympathy for OP, but running ads is really hard. It's why most people set targets for things that happen in the app (OP updated to a target of winning a game, not just installing), and why companies have entire teams dedicated to monitoring Google/Meta/X ads. Trial runs are often thousands of dollars, and it's easy to evaporate money. I've made the same mistakes, and its why y-combinator strongly recommends against running online ads, especially in the beginning.
yalok this is unfortunately very common - I kept running very small ads regularly over many years (~10 years by now) and the % of bots has been steadily increasing, and Google doesn't really have any system to report those reliably to them. In fact, it's contrary to their incentives to investigate & fix these problems - unless they feel some pressure from competition... From time to time, Google ads sales reps call me and are trying to convince me to increase the ads spend. I complain about bots, and none of them had any suggestion on what I can do to appeal it etc. I mean I can appeal on some…
legonigel What is the incentive for the bot owners? Why do the bots download and install the apps? It has some (very small) cost to them, and I don't understand the value for them.
blitzar When you are talking to your VC, they are just "installs"
phenomen If you have their IP addresses in your dashboard, go to Google Ads > Admin > Account Settings > IP Exclusions. Then, add the entire network/data center range there (i.e. 123.4.5.*). In 99% of cases, these bot networks are not run from residential providers. You can confirm IPs at https://ipgeolocation.io/ After running Google Ads for a couple of years, our current exclusion list has over 4000 networks just in the US.
xnx Switch to a pay app and the problem goes away.
sergiotapia How much of Google's revenue is just scams? Search for "datadog" and the first result is a sponsored link for Datadog. I click that and datadog had to pay google. scroll a bit down and there's the "real" natural link. Fugazi!

OpenAI agents carried out an undisclosed attack on RubyGems 235p 126c

An OpenAI agent swarm was responsible for this incident

The agents used RubyGems’ automatic build system to achieve remote code execution

The agents attempted to exploit a novel vulnerability to try to steal user API keys

Agents bypassed RubyGems’ email confirmation system in order to make a large number of accounts

The agents attempted to use RubyGems’ webhook system to store data

The agents continued to use RubyGems in June.

When agents were hacking OpenAI’s infrastructure, they used RubyGem packages to exploit Artifactory

Spencer Kitts, Thomas Larsen, Sydney Von Arx·11 September 2026

On May 11th, 2026, hundreds of malicious packages were uploaded to RubyGems by AI agents. We believe these were authored by internal OpenAI agents (more).

Attempted to steal RubyGems user API keys by exploiting a novelThat is, novel at the time. The vulnerability was discovered and patched independently later. vulnerability in the RubyGems server. We don’t know if they succeeded (more).

Abused RubyDoc.info to execute arbitrary code (more).

We share our detailed findings below. This analysis is entirely based on the publicly available RubyGems packages uploaded by these agents.We also talked with RubyGems and rubydoc.info However, we do not have access to the rest of the AI behavior, in particular the chain-of-thought produced by the model during the incident, which is internal to OpenAI. Therefore, we do not know why the AI agents chose this strategy or whether it was successful.

The RubyGems team stopped new user sign-ups for four days to stem the tide of packages from the agents’ accounts. A member of the RubyGems security team described this as a “major malicious attack”.

Security companies termed the incident the “GemStuffer campaign”, while also noting confusion at the purpose of the attack. The malicious packages uploaded were used to retrieve information from UK local government sites – data that was available to the public. One news outlet writes: “It's not clear what exactly the end goals are, as the information appears to be publicly accessible anyway.”

We thank Jonas Wiedermann-Möller (@j0wimo) for first discovering that agents had likely uploaded to RubyGems, and the community as a whole for their work to chase down new signs of agent activity.

May 5Earliest package uploaded by an OpenAI agent to RubyGems

May 8First package with “oai” in its name

May 11First time we observe OpenAI agents attempt to edit a public wiki

May 11–12Agents submit over 2,000 packages to RubyGems

May 12RubyGems disables new user registration, describing the traffic as an ongoing DDoS

May 12First message-board post on OpenAI Artifactory instance.

May 13RubyGems reports the spam has stopped, and removes 500+ malicious packages.

May 16RubyGems restores new user registration.

An OpenAI agent swarm was responsible for this incident

We believe that this incident was the result of an OpenAI agent swarm. Our main sources of evidence are:

The packages are clearly LLM-authored. We ran some of the malicious packages through Pangram, which detected them as 100% AI generated. This is evidence that the attack was an agent swarm (but not that it originates from OpenAI).

Agents self-identified as being from OpenAI. Hundreds of the packages that were uploaded contain “oai” in their name. Fifteen of the packages set “oai” as their author. Another lists an email for contact as “openaixyz65947@gmail.com”.

oaitest1778473828 oaibootx8192 oaibooty9217 oaibootz9218 oaibo396866 […] oaibo825590 oaibo048288 oaibx0092307 oaibx7324267 oaibx1202338 oaibx4676369 oaicx8859010 oaicx3857133 oaicx2721076 oaicx6062340 oaicx4433606 oaicx3769699 oaidx4526859 oaidx0276239 oaidx3879209 oaidx7402019 oaidx1466937 oaidx3409275 oaidx1337585 oaidx6514197 oaidx3492001 oaidx1469215 oaidx6135652 oaidx1169327 oaiex4149420 oaiex1182709 oaiex7410346 oaiex0549290 oaiex3900663 oaiex4736401 oaiex9823513 oaiex3222069 oaiex8413575 oaiex0014506 oaifx7943598 oaifx8889601 oaifx9269956 oaifx8306741 oaifx2280367 oaifx1955773 oaifx0927711 oaifx4260376 oaifx9677940 oaifx1757803 oaifx9741380 oaifx3608457 oaifx7129963 oaifx7303384 oaifx6387627 oaifx9667097 oaifx2401408 oaifx8755814 oaigx7857181 oaigx4516770 oaigx5578224 oaigx5861576 oaigx4634836 oaigx1767798 oaigx9094125 oaigx8693871 oaihx7985797 oaihx8175223 oaihx5974804 oaihx8693617 oaihx9923604 oaihx0305933 oaihx0157786 oaihx7579061 oaihx7237922 oaihx7924258 oaiix8443749 oaiix9664993 oaiix0379958 oaiix3669509 oaiix7984341 oaiix7006631 oaiix0231326 oaijx6438369 oaijx0303634 oaijx0156671…

enraged_camel Every passing day OpenAI looks more and more reckless. One wonders what other systems their agents have broken into without detection.
jsnell I can't believe we're finding out about this from 3p researchers again (but nice job on the investigation!). OpenAI had two great opportunities to disclose this. The HF incident report, and in response to the German Wiki issue. It seems impossible to believe they didn't know. This must be the same training run the HF incident was about, and this should have lit up like a Christmas tree in the investigation. How many more incidents do they know about and didn't disclose?
nonconstant Kudos to RubyGems team for handling it, but open source fighting off the AI lab-powered robots is completely unfair. OpenAI should at the very least donate large sums of money to everyone they attacked.
gverrilla Is there a world where Sam or Dario can seize the bitcoin network somehow?
creatonez You shouldn't be allowed to have an internet connection if you're going to use it for unsandboxed agent slop with no access controls or human confirmation. This has nothing to do with hypothetical future AGI. It's the same type of idiocy as pressing a bunch of random buttons on a chemical factory control panel and then thinking you won't be criminally charged for it because the equipment caused the problem. If you actually have a serious use case that needs 24/7 unmonitored agents, you can assemble all of the data the agents need locally and avoid these insanely obvious and well documented…
zmmmmm It seems like all this happened in the same time period earlier this year. It makes me wonder if all of these were part of a single larger incident where multiple experiments were run with insufficient or missing constraints or an unknowningly misaligned model.

A Design Space Exploration of Async/Await 121p 25c

Many programming languages now provide the async/await keywords for expressing concurrency. The design rationale is pretty consistent: to make concurrent programs look more like straight-line code (see: Python, Rust, or Swift). We therefore describe the paradigm which encompasses async/await as straight-line asynchrony, as opposed to using event loops or callbacks.

Language designs for straight-line asynchrony have been brewing for over 15 years. In this project, we wanted to understand: how similar or different is async/await between languages? The short answer is a lot more different than we expected. We wrote a paper, “A Design Space Exploration of Async/Await”, to explain how.

To demonstrate how much modern languages can diverge, here’s a small async program written in pseudocode. One function writes to a log, and another fires off the log write as a background task and moves on.

async fn write_to_log(): print("A") // simulate a slow log write await sleep(2) print("B") async fn fire_and_forget(): task = spawn write_to_log() // return without awaiting the task async fn main(): await fire_and_forget() await sleep(1) print("C")

What would you expect this program to print?

There isn’t really a right answer, because you were probably right for some language. Below is how seven modern async runtimes actually behave:

Four different answers, for a program whose entire job is to write a log line in the background. And it gets worse. In the paper, we show that across the seven runtimes, no two produce the same output for three variations of this simple program!

Do you actually know your language’s async semantics?

While watching videos from a programming influencer you may have have heard the terms “cold” or “hot” async function calls. The idea is that “hot starts” return a task that is immediately running in the runtime, while “cold starts” return an inert object that does nothing until awaited.

Hot vs. cold functions is what we call an async design dimension: a design decision that affects the observable semantics of program execution (as opposed to matters of pure performance). We call this particular dimension “Eagerness”, and in the paper we identify nine such dimensions from modern implementations of straight-line asynchrony. Below we’ve grouped these nine dimension into three categories that roughly correspond to the lifetime of a task: Start of Life, End of Life, and Cancellation.

Click a language to trace its design choices through the table.

How to evaluate an async function application.

Evaluate to a coroutine without executing further.

Evaluate in current thread, and schedule as task on await.

Guarantees on whether await points suspend.

C# · Swift · Tokio · Smol · Asyncio · Trio

The default interval of time during which a task may exist.

Tasks by default may exist until the end of the runtime.

Tasks by default may exist until the end of their spawning scope.

[For Indefinite Extent] The type of reference to a task held by the runtime.

How a task is cleaned up at the end of its extent.

The task is cancelled, and then possibly awaited.

What happens to exceptions in unawaited tasks.

JavaScript · C# · Tokio · Smol · Asyncio · Swift

Whether a task is able to respond to being cancelled.

The task cannot respond to being cancelled.

How cancellation is communicated through the task graph.

Starting from the root task, and communicated from dependents to dependencies.

Starting from the root’s dependencies and communicated to dependents.

[For Aware Cancellation] How long a cancellation of a task lasts.

A task can ignore cancellation and proceed as normal.

A task can ignore cancellation but remains cancelled.

Two of these axes are particularly relevant for our example program. Languages with Dynamic Extent do not allow tasks to outlive the functions in which they were spawned. Unlike the other languages, Swift and Python+Trio chose Dynamic Extent. This means that within the function fire_and_forget, the task associated with write_to_log cannot outlive the function fire_and_forget.

Although Swift and Trio both chose Dynamic Extent, they differ in choice of Destruction. At the end of the fire_and_forget function scope, Swift uses Cancelled Destruction, and cancels task while Trio uses Awaited Destruction and politely waits for write_to_log to finish. The choices of Extent and Destruction explain why Swift prints “AC” and Trio prints “ABC”.

Each design dimension has trade-offs of performance, memory usage, ergonomics, semantics, etc. There’s no right or wrong…

homarp The paper explores how async/await behaves across today's languages: https://arxiv.org/abs/2608.20677
perrygeo Amazing work. It's one thing to say "async is complex". It's another to parse that statement so carefully as to have a cross-language theory of async execution. Looking forward to digging into this!
layer8 It would be a fun coding agent benchmark to have them translate such a program between the different languages and see whether they preserve the respective semantics.

GrapheneOS' rewritten Messages app is released 185p 106c

GitHub Copilot appDirect agents from issue to merge

GitHub Advanced SecurityFind and fix vulnerabilities

Code securitySecure your code as you build

Secret protectionStop leaks before they start

GitHub SponsorsFund open source developers

Enterprise platformAI-powered developer platform

GitHub Advanced SecurityEnterprise-grade security features

Copilot for BusinessEnterprise-grade AI features

Premium SupportEnterprise-grade 24/7 support

There was an error while loading. Please reload this page.

Notifications You must be signed in to change notification settings

There was an error while loading. Please reload this page.

Version 13 replaces the legacy interface with Jetpack Compose and Material 3. It rebuilds every screen, adds new conversation controls and large-screen support, and fixes problems with crashes, notifications, security, and message handling.

Material 3 / Expressive design with shapes

Two-pane conversation layout on large screens

New adaptive app icon with a monochrome variant

New onboarding covering SMS privacy, permissions, and default-app setup

Correct display cutout and system bar handling in both orientations

Snooze notifications for 1, 8, or 24 hours, or indefinitely

Mark conversations as unread from the menu or by swiping

Redesigned multi-select actions for archiving, deletion, blocking, and notifications

Indicators for unread, pinned, snoozed, notifications off, work profiles, and incoming MMS status

Pinning, archiving, and deletion now apply immediately

Rebuilt message bubbles, grouping, selection, and link handling

Full-screen message details with copyable sender, recipients, timestamps, delivery status, size, type, and priority

Add, edit, or remove MMS subjects from the conversation overflow menu

Blocked-sender banner with an unblock action

Add participants to existing conversations or create groups from the new-chat screen

Redesigned recipient picker with alphabetical sections, email addresses, formatted numbers, and multiple numbers per contact

Emergency numbers cannot be dialed from conversations

SIM fallback now uses the system default instead of the first SIM

Top-bar actions for calls, contact cards, and conversation settings

Attachment limits checked before sending, with visible send errors

Messages arriving during conversation deletion are no longer destroyed

Rebuilt media picker with photo and video capture, flash controls, and Android's embedded photo picker

Redesigned audio recording with slide-to-cancel and hands-free locking

Rewritten photo viewer with pinch-to-zoom, page indicators, details, and correct system bar handling

Rewritten vCard viewer with avatars, contact-change refresh, and saving to contacts

New share picker with search, recent conversations, alphabetical contacts, and multi-select

Edit shared content, add a subject, preview it, and choose a SIM

Forwarding and the widget use the same picker

The widget's new-message button opens the new-chat screen

Missing shared text is read from its content URI; missing subjects use the shared title

Rewritten main, general, and per-SIM settings

Rewritten licenses screen with generated license data

Existing per-conversation notification settings are preserved

YouTube link previews are opt-in and disabled by default

Shared-content validation rejects file: URIs and private app files, and checks content URI permissions

Shared text read from content URIs receives the same private-file checks

Widget receivers are no longer exported, and widget intents are restricted to the app

Pending intents use FLAG_IMMUTABLE where mutability is not required

Allocation limits added to EXIF APP1, MMS PDU, and MMS content-type parsing

Fixed GIF null dereferences, missing dimension checks before native transcoding, and signed-character colour corruption

Updated the AOSP vCard parser with upstream fixes

Bounded notification people lists, fixing issue a reported crash

Onboarding explains that SMS is unencrypted and recommends end-to-end encryption for sensitive messages

Fixed crashes in the widget, failed SMS database inserts, declined-call quick responses, settings for unsaved numbers, and share intents

Share intents no longer silently drop attachments

Incoming SMS messages are imported immediately

Failed-message notifications are delivered correctly

Inline notification replies no longer open the home screen

Blocked conversations no longer play notification sounds

Deleted conversations no longer receive notifications during sync

Sync no longer overwrites archive changes

Fixed a race between MMS downloads and notifications

LoganDark That was really fast, they only announced it like a day or two ago IIRC? Also: no screenshots of the redesigned interface?!
mewse-hn Does it do RCS?
mmooss With a seemingly insurmountable workload - maintain an actually secure OS fork - for a small team, I wonder how GrapheneOS prioritizes things like messaging apps? I'm not saying it's wrong at all; I am just interested in the thinking inside a project like that. I can imagine many reasons to devote resources to it: * A platform is only as good as its apps. If they want to grow GOS in the general public, it requires a good SMS/MMS messaging app. * LLM tools greatly reduce development costs, especially for well-known functions like text messaging. * Without secure messaging (as far as SMS/MMS can…
cloudie78 How about some screenshots?
gib444 Looks prettier but a shame an obvious regression happened for one of my primary use-cases: copying one-time codes. Long pressing on a link selects the entire SMS instead of bringing up a menu for the link. [0] Makes me wonder about all the testing now And the issue worded/tagged as a feature instead of a regression. Are the devs aware the feature existed before? later : It's doubly disappointing because just 4 days ago the project wrote [1], on using AI for the new Messaging app: "it's also helping us raise our standards for our own work by pointing it at that to get extremely pedantic…
Maskawanian Does anyone know if this is something that we can install now, or if it will show up in the next OS release?

Project Blinkenlights 41p 17c

Project Blinkenlights turns buildings into giant interactive displays. What started in 2001 as the Chaos Computer Club's birthday present to itself became a series of light installations on three continents — the whole story is told in the project overview.

Blinkenlights — Berlin 2001, Haus des Lehrers: the installation that started it all (18×8 pixels, monochrome). With the reprise projects Reloaded (2004) and the Bauschild.

Arcade — Paris 2002, Bibliothèque nationale de France: the world's biggest computer game display (26×20 pixels, 8 greyscales).

Stereoscope — Toronto 2008, City Hall: two towers, one matrix (96×32 pixels, 16 greyscales).

Polychrome — since 2023: Blinkenlights in color, from Camp to Nation of Gondwana.

The galleries play the original animations on their facades, the Movie Converter turns your own movies into GIFs and WebP, and the press room collects two decades of coverage.

Stereoscope wins Nuit Blanche People's Choice Award

Blinkenlights Tech Talk in Toronto this tuesday

zahlman > Stereoscope — Toronto 2008, City Hall: two towers, one matrix (96×32 pixels, 16 greyscales). Oh man, I think I remember that.
Animats Downtown Shenzhen is rigged for that. For a while, the SF Bay Bridge was. I once went to an Autodesk event in SF where the guy behind it was controlling the lights from his laptop.
jolt42 Is this still well recognized? ACHTUNG! ALLES TURISTEN UND NONTEKNISCHEN LOOKENSPEEPERS! DAS KOMPUTERMASCHINE IST NICHT FÜR DER GEFINGERPOKEN UND MITTENGRABEN! ODERWISE IST EASY TO SCHNAPPEN DER SPRINGENWERK, BLOWENFUSEN UND POPPENCORKEN MIT SPITZENSPARKEN. IST NICHT FÜR GEWERKEN BEI DUMMKOPFEN. DER RUBBERNECKEN SIGHTSEEREN KEEPEN DAS COTTONPICKEN HÄNDER IN DAS POCKETS MUSS. ZO RELAXEN UND WATSCHEN DER BLINKENLICHTEN.

Testing Race Conditions 11p 0c

Many security bugs are race conditions, where multi-threaded execution has to occur with the right interleaving for a negative effect to appear. This creates challenges for several use cases:

Confirming bug candidates that have been discovered manually or through static analysis.

Regression tests: After fixing a race condition bug, there is often no good way to write a regression test that reliably triggers the bug as part of a test suite.

Automatic bug discovery, such as fuzzing: It is hard for a fuzzer to exercise all interesting interleavings of concurrent operations, or reach code paths that are only exercised when operations are racing.

I mostly discover bugs by manually reading code. When I think I’ve found a bug, I normally write a test case to either prove or disprove that the bug exists. For race condition bugs, it can be hard to achieve either outcome. For Linux kernel bugs, I often resort to recompiling the kernel after adding conditional mdelay() calls (which spinloop for roughly the specified amount of time) in appropriate places; I usually make these conditional based on the name of the running thread, though sometimes more complex conditions are needed. On platforms that support DTrace (like macOS and Windows), it is possible to use DTrace probes that call chill() for similar effect, though the utility of this is limited as DTrace can only trace on non-inline function boundaries or explicit trace points, rather than on every instruction. Regardless of platform, this approach can be time consuming and can require trial and error to definitely determine whether code is buggy.

Additionally, in the Linux kernel, fixes for race condition bugs are often accompanied by hand-written ASCII diagrams showing problematic thread interleavings with call graphs and relevant memory accesses (for example, see this recent rt_spin_unlock UAF fix, or this recent jbd2 deadlock fix). It would be convenient to have developer tooling that can analyze potentially vulnerable code and show results in a similar representation.

I wrote tools for exploring possible interleavings of multi-threaded test cases for the Linux kernel:

A tool that automatically tests all possible A-B-A interleavings of a test case.

A terminal UI for manual exploration of possible interleavings.

A GUI for manual exploration of possible interleavings.

The kernel part of this is intended to also be usable for discovering race conditions via fuzzing, but userspace tooling for that still needs to be implemented.

The tools are available on GitHub under the name MAccConc, short for “Memory Access Concurrency”; see the README there for installation and usage instructions.

If you just want to see the tooling in action, skip to Demo: automatic testing.

If you’re just interested in the theory behind the tooling, read section Stable identifiers for memory accesses across runs: count-augmented stack traces.

This project was inspired by discussions with Ned Williamson, whose sockfuzzer project involved exploration of concurrency bugs by using a custom scheduler that can reschedule at synchronization primitives to explore interleavings. See the conference talk slides and recording focused on the concurrency testing aspect of this.

My tooling is largely based on ideas similar to SKI, but SKI uses a different implementation: It records memory accesses and controls scheduling of vCPUs using a patched version of QEMU in TCG mode, and uses VM snapshots to explore different execution interleavings.

Discovering memory accesses that could contribute to race conditions (communication points)

As described in the SKI paper, interesting execution interleavings of a given multi-threaded test case can be discovered by tracing memory accesses of all threads and searching for pairs of accesses on two threads that could interact with each other - meaning, roughly, that at least one of them is a write operation, and they access overlapping memory ranges. The SKI paper calls such memory accesses communication points.

This requires some mechanism to collect memory access coverage. SKI did this by patching QEMU’s TCG mode; I am instead relying on ASAN instrumentation in “outline” mode (compiler backend flag asan-instrumentation-with-call-threshold=0, selected by CONFIG_KASAN_OUTLINE in the Linux kernel), which generates helper function calls on memory access. I believe that the kernel is the right place to collect this data because it would allow the kernel to also provide higher-level information about lock…


Show HN: ResolveHQ – A Helpdesk Built on Cloudflare Workers, D1, R2 and Queues 24p 9c

GitHub Copilot appDirect agents from issue to merge

GitHub Advanced SecurityFind and fix vulnerabilities

Code securitySecure your code as you build

Secret protectionStop leaks before they start

GitHub SponsorsFund open source developers

Enterprise platformAI-powered developer platform

GitHub Advanced SecurityEnterprise-grade security features

Copilot for BusinessEnterprise-grade AI features

Premium SupportEnterprise-grade 24/7 support

Notifications You must be signed in to change notification settings

ResolveHQ is a Cloudflare-native, self-hostable helpdesk for small support teams.

Run a shared inbox with tenant-isolated customers, tickets, assignment, status, priority, tags, and full-text search.

Sign up as owner, invite teammates, and manage Owner/Admin/Agent roles with a workspace switcher across organizations.

Thread email correctly per RFC 5322, resistant to subject-line spoofing across tickets.

Receive mail through Cloudflare Email Routing and send it through Resend, with delivery status, retries, and idempotent webhooks. Outbound replies carry their linked attachments, and dead-letter queues drain to durable, recoverable records.

Attach files to tickets through validated, authorized R2 uploads.

Reply faster with saved replies, internal notes, AI-drafted responses (opt-in), and a responsive three-pane inbox with optimistic-version conflict handling.

Reset passwords and accept invitations through system email sent via the same provider seam as ticket mail.

Publish a public help center from knowledge-base articles, with drafts kept private to your team.

Track volume and response speed in Reports, export any window to CSV, and automate triage with rule-based Automations.

Notify agents of assignments and customer replies in-app, and work comfortably in light or dark mode.

Export everything stored about a customer as JSON, or erase it with a durable, resumable workflow that cancels queued mail.

Recover automatically: a five-minute cron job retries stalled mail jobs and cleans up staging and orphaned data.

Control AI assistance per workspace: it stays off until an admin enables it in Settings, and only then are ticket conversations sent to OpenAI.

Set TICKET_RETENTION_DAYS (for example 365) to have the scheduled sweep permanently delete resolved and closed tickets older than that window, including attachments.

The workspace uses a Slack-inspired aubergine sidebar, self-hosted Lato typography, Lucide icons, and Radix UI primitives. Theme-aware controls and status colors support light and dark workspaces. On mobile, bottom navigation and a keyboard-accessible workspace drawer keep all destinations available; ticket columns adapt to preserve subject readability.

Use Cmd/Ctrl+K to jump between pages. The sidebar dock contains notifications, theme switching, and account actions.

ResolveHQ runs as a single Cloudflare Worker in your own account. Hono serves both the REST API and the built React application. Cloudflare D1 holds tickets and customers, Cloudflare R2 holds attachments, and Cloudflare Queues carry inbound and outbound mail jobs. Cloudflare Email Routing delivers incoming mail to the Worker, and Resend sends outgoing mail. Tickets and attachments are stored in your Cloudflare account; outbound email content passes through Resend.

ResolveHQ can run on Cloudflare’s Free plan for small deployments, provided usage stays within the current limits for Workers, D1, R2, Queues, Cron Triggers, and Email Routing. CPU-intensive authentication or mail parsing may require Workers Paid; benchmark your deployment. Queues are available on Workers Free. R2 requires account activation and billing setup separately. Resend handles outbound email under its own limits. See the Free-plan audit.

The easiest way to get started is with the Deploy to Cloudflare button above. You will need:

A Cloudflare account; Workers Free supports Queues. Activate R2 separately.

A domain on Cloudflare, so you can set up Email Routing.

Optionally, a Resend account with a verified sending domain, to send outgoing mail.

After deployment, open your ResolveHQ URL and sign up as the owner, giving an optional support email that becomes your default inbox. In the Cloudflare dashboard, add an Email Routing rule sending that address to the deployed Worker, then send a test email to confirm it arrives in the inbox.

See the deployment guide for what the deploy flow provisions, required configuration, first-run setup, and manual deployment.

npm install cp .dev.vars.local.example .dev.vars npm run…

gepeake Just used a very similar stack to make my own mail client. Workers for routing, Resend for sending, uses a CF and Resend key to set up any inbox on any domain I want ~instantly so I can have inboxes for everything. Very cool, I think workers are underrated!
orliesaurus Could you include some screenshots in the README of the functionalities? that would go a long way! thank you in advance!

Litelm: LiteLLM Without the Bloat 91p 34c

GitHub Copilot appDirect agents from issue to merge

GitHub Advanced SecurityFind and fix vulnerabilities

Code securitySecure your code as you build

Secret protectionStop leaks before they start

GitHub SponsorsFund open source developers

Enterprise platformAI-powered developer platform

GitHub Advanced SecurityEnterprise-grade security features

Copilot for BusinessEnterprise-grade AI features

Premium SupportEnterprise-grade 24/7 support

Notifications You must be signed in to change notification settings

litellm's routing + translation in ~2,900 lines and 2 dependencies (openai, httpx).

litellm routes LLM calls across providers and translates between message formats. That core is buried under 100k+ LOC of proxy servers, caching layers, cost tracking, and dozens of features most users never touch. litelm extracts just the call path — model routing, message translation, streaming, tool use, embeddings — and nothing else. No Router class, no proxy, no caching.

pip install litelm # openai + httpx pip install litelm[anthropic] # + anthropic SDK pip install litelm[bedrock] # + boto3 pip install litelm[all] # everything

import litelm # Basic completion response = litelm.completion("openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}]) print(response.choices[0].message.content) # Streaming for chunk in litelm.completion("groq/llama-3.1-70b-versatile", messages=[...], stream=True): print(chunk.choices[0].delta.content or "", end="") # Embeddings response = litelm.embedding("openai/text-embedding-3-small", input=["hello world"])

Every function has an async variant: acompletion, aembedding, aresponses, atext_completion.

The API mirrors litellm — same function names, same arguments, same response types. If you're using litellm today, switching is s/litellm/litelm/ in your imports.

Routes to 19 providers via "provider/model-name" syntax. Any OpenAI-compatible endpoint works via api_base.

Set the environment variable for your provider:

export OPENAI_API_KEY=sk-... export ANTHROPIC_API_KEY=sk-ant-...

litelm.completion("openai/gpt-4o", messages=[...], api_key="sk-...") litelm.completion("openai/gpt-4o", messages=[...], api_base="http://localhost:8000/v1")

All provider errors are mapped to litelm's exception hierarchy:

from litelm import ContextWindowExceededError, RateLimitError, AuthenticationError try: response = litelm.completion("openai/gpt-4o", messages=messages) except ContextWindowExceededError: # prompt too long — truncate and retry pass except RateLimitError: # back off pass except AuthenticationError: # bad API key pass

tools = [{"type": "function", "function": { "name": "get_weather", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}}, }}] response = litelm.completion( "openai/gpt-4o", messages=[{"role": "user", "content": "Weather in Paris?"}], tools=tools, tool_choice="required", ) tool_call = response.choices[0].message.tool_calls[0] print(tool_call.function.name, tool_call.function.arguments)

Any OpenAI-compatible server works via api_base:

# vLLM litelm.completion("openai/my-model", messages=[...], api_base="http://localhost:8000/v1") # Ollama litelm.completion("ollama/llama3", messages=[...], api_base="http://localhost:11434/v1") # LM Studio litelm.completion("openai/local-model", messages=[...], api_base="http://localhost:1234/v1")

litelm is human-directed, AI-assisted software. Much of the code was written with Claude Code using Claude Opus 4.6/4.7. Code written from 2026-05-14 onward is written through Pi using GPT-5.5. Compatibility claims are based on tests and maintainer review, not AI authorship.

Maintainer attestation, 2026-09-11: LiteLLM's routing/formatting changes were reviewed from 649eb2d through 9a715df2. The audit triaged 360 core-path commits, inspected upstream tests for potentially relevant behavior, and fixed the resulting compatibility gaps test-first. Local scoped tests: 262 passed, 55 skipped; all 45 available-provider live tests and all 10 DSPy smoke tests also passed with the current dependency lock.

This attests litelm's declared routing/formatting/DSPy surface only, not full litellm compatibility.

Alpha. 262 own tests passing. The current scoped LiteLLM 9a715df2 baseline has 75 passing ported tests and no remaining actionable assertion/runtime failures.

DSPy drop-in verified — all 7 execution paths proven live (Predict, CoT, typed signatures, streaming, embeddings, tool use, multi-output).

uv run --extra all pytest tests/ -x --ignore=tests/ported --timeout=10 # 262 non-live tests bash…

khalic I strongly recommend the authors rewrite the readme by hand. It’s kind of a snif test for how much care someone put into this project.
Centigonal This is a cool project, and the idea of using LLMs to selectively extract features from open source projects is an interesting concept. The only thing I take issue with is the phrase "LiteLLM Without the Bloat." A lot of the features that have been removed (like cost tracking, streaming, caching) are... kind of the core value proposition of LiteLLM for many of their users.
freshtake First off, cool project! It's always great to see derivatives that question the efficiency of the established product. I think the main thing the readme is missing is the core benefits. Reducing LOC and dependencies is cool, but it would be great to understand if this provides some additional benefits like lower latency or memory requirements.

Λ Snap – An inviting programming language for kids and adults for CS study 107p 53c

Snap! is brought to you by UC Berkeley and SAP.

stymaar For those wondering the difference with scratch, it's on the about page: > About Snap! > Snap! (formerly BYOB) is a visual, drag-and-drop programming language. It is an extended reimplementation of Scratch (a project of the Lifelong Kindergarten Group at the MIT Media Lab) that allows you to Build Your Own Blocks. It also features first class[1] lists, first class procedures, and first class continuations[2]. These added capabilities make it suitable for a serious introduction to computer science for high school or college students.
sinuhe69 Snap is designed to be more expressive and powerful than Scratch. But I find debugging them is very painful. Changing the name of a variable or block for example, could create “holes” in the calling sites but the system can or cannot report an error and fails silently. Snap is IMO more flaky and the team seems more eager to add features than polish the existing ones or make the system mor robust and stable.
anguishe Nice! I did a few chapters of this when I was learning. I'll have to add it to my list to go and check out any new updates/features they have. I can NOT wait to get my little one into something like this!
jc4p I've been spending a lot of time with my nephews working on Scratch and Snap. They also attend local classes and their teachers switched from Snap to this which is way simpler and stickier with my nephews: https://www.microsoft.com/en-us/makecode

0:57 UTC · cached 5 min · ?n=1–15