<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Cloud-Computing on Publish Assistant</title><link>https://pub.sqrt.fr/vincent/publish-assistant/tags/cloud-computing/</link><description>Recent content in Cloud-Computing on Publish Assistant</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 01 Nov 2024 00:00:00 +0000</lastBuildDate><atom:link href="https://pub.sqrt.fr/vincent/publish-assistant/tags/cloud-computing/index.xml" rel="self" type="application/rss+xml"/><item><title>SoCC 2024 Digest</title><link>https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/socc-2024/</link><pubDate>Fri, 01 Nov 2024 00:00:00 +0000</pubDate><guid>https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/socc-2024/</guid><description>&lt;p&gt;12 papers selected.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="queue-management-for-slo-oriented-large-language-model-serving"&gt;Queue Management for SLO-Oriented Large Language Model Serving&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu &lt;em&gt;et al.&lt;/em&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — A queue management framework that enforces latency SLOs for LLM serving by dynamically routing and prioritizing requests across heterogeneous inference capacity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why notable&lt;/strong&gt; — As LLM deployments move into production clouds, meeting strict time-to-first-token and total latency SLOs becomes critical; this work directly addresses that gap with a practical, deployable solution. It is one of the first papers to treat LLM serving as a cloud SLO-management problem rather than a pure model-optimization problem.&lt;/p&gt;</description></item><item><title>CCGrid 2024 Digest</title><link>https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/ccgrid-2024/</link><pubDate>Mon, 06 May 2024 00:00:00 +0000</pubDate><guid>https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/ccgrid-2024/</guid><description>&lt;p&gt;10 papers selected.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="fair-efficient-multi-resource-scheduling-for-stateless-serverless-functions-with-anubis"&gt;Fair, Efficient Multi-Resource Scheduling for Stateless Serverless Functions with Anubis&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Amit Samanta 0001, Ryan Stutsman&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — Anubis introduces a fair, multi-resource scheduler for stateless serverless functions that achieves efficiency without sacrificing isolation between tenants.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why notable&lt;/strong&gt; — Fairness in serverless resource allocation is an open problem as functions compete for heterogeneous resources (CPU, memory, I/O); Anubis provides a concrete, deployable answer. The work directly addresses a gap in production FaaS platforms where existing schedulers optimize for throughput but ignore per-tenant equity.&lt;/p&gt;</description></item></channel></rss>