Ring.onReadEnd and Ring.ReadDiscard call Reset() when the last readable
byte is consumed, which rewinds Off and End to 0. That is a valid empty
state, but it also moves the write position, and bytes staged past that
position with PeekWrite are addressed relative to it. After the rewind a
Commit hands back whatever bytes happen to live at the start of the
buffer instead of the staged ones.
tcp stages out-of-order segments exactly this way (see reassembly.store,
"the ring write pointer advances in lockstep with rcv.NXT, so the staged
bytes are always where seq implies"). Reading is driven by the
application, so a receiver that drains its stream while a gap is open
loses the staged tail silently: the stream arrives with the correct
length and the wrong contents.
Observed on a Sophgo SG2002 (100Mbit DWMAC, receive ring dropping frames
under load): a 16 MiB download arrived complete with a different
SHA-256, and TLS over the same path failed with "bad record MAC", while
the transmit path was bit-perfect.
Mark the ring empty without rewinding instead: End=0 is what empty
means, and the write position is then Off, which Ring.Write already
supports explicitly.
Two tests, both failing before the change:
- internal: staged bytes survive the ring being read empty.
- tcp: a reassembled stream is byte-identical when segments arrive
reordered (segment K arrived as the contents of an earlier segment).
Co-authored-by: Derek den Haas <i.pestano@easyflor.nl>
ControlBlock.Send refused any outgoing segment carrying data in
FIN-WAIT-1, citing RFC 9293's "no further SENDs from the user will be
accepted by the TCP implementation". That rule bounds what the
application may queue, which Handler.Write already enforces, and not the
retransmission of data the connection has already accepted from it.
The consequence is that write-then-close, which is what nearly every
server does with a response, cannot recover from losing its last data
segment. The FIN occupies a sequence number above that data, so the peer
cannot cross the gap to process the close: it waits for bytes that are
never resent while the sender waits for an ACK that cannot arrive.
Neither side times out at the TCP layer.
FIN-WAIT-2 keeps the restriction and gains the reasoning: it is reached
by our FIN being acknowledged, which acknowledges everything below it, so
no unacknowledged data can remain there.
PendingSegment needs the same distinction, since it decides whether to
offer send-buffer data at all; it gets an unexported State predicate
rather than a new exported one.
Two tests, the second red before the change:
- retransmission after RTO expiry, covering the Handler/LossRecovery
seam that the RTO unit tests do not reach (they exercise the state
machine in isolation, where it behaves correctly).
- retransmission after a close with unacknowledged data.
Co-authored-by: Derek den Haas <i.pestano@easyflor.nl>
* add http/httphi
* begin adding httphi tests
* claude found neat bugs
* add low level Handle function and more tests
* more tests, run go generate
* add Hijacker-like functionality
* improve locking and acquisition of Exchanges in reconfiguring
* several bugfixes, add internal.IntLen, round up http-linux example with new router API
* small nit
* add benchmarks
* add query handling
* remove ForEach pattern, allocates in TinyGo
* massive documentation push and code reordering in files
* Router.Handle returns error after being torn down
* run go fix
* rework Mux interface to receive a string request path
* add MethodFrom
* minor doc nit
* fail on incomplete staging
* add raw buffer access
* add streaming API distinct from Exchange
* begin adding multipart form logic
* finish rounding up multipart form parsing
* remove status type
* first Multipart approach
* begin adding readMultiPart
* add Exchange.ReadMultiparts reimagining of clanker slop
* ai insists with backoffs
* simplify clanker slop
* apply go fix
* add a pattern argument to Mux
* explicit header key/value alloc and add ExchangeConfig
* fix tests after excplicit header alloc change
* fix examples
* run go fix
* expose rawsock as experimental package (will use for external benchmarks)
* remove backoff from form parsing
* @MDr164 suggestions get potential fixes
* apply go fix
* add examples
* add README.md
* fix rawsock tinygo implementation
* apply @MDr164 various fixes
* update documentation on ContentLength methods and fix bug in Form reset on empty body
* fix tests
* add fuzz tests
* run go fix
* io.ErrNoProgress on parsing form spin
* run go fix
* remove backoff assumption from Router
* httphi.Handle rejects unsupported protocols
* go format router.go
* add kvbuffer
* rewrite Cookie with KVBuffer
* mid refactor of KVBuffer into Header
* work on KVBuffer exhausted semantics
* add Go's ServeMux Request.PathValue access semantics to Exchange, Mux and MuxSlice
* add PathValue example
* document all the things; improve req Query semantics; add Form.EnableBufferGrowth
* unexport kvBuffer
* add Exchange.PathValueAppend
* use stdlib in example instead of rawsock
* remove rawsock from http example
* add darwin arch rawsock
* fix example
* rename Router.TeardownGoroutines to Shutdown matching http.Server.Shutdown
* rename types and identifiers
* @MDr164 Content-Type and Transfer-Encoding bug catches
Provide a concrete packet-loss recovery algorithm for the LossRecovery
interface added in #168: the RFC 6298 round-trip-time estimator and single
retransmission timer.
RTO is a pure, reactive state machine. It derives RTT estimates and
retransmission decisions solely from the segments observed through the
LossRecovery hooks and the monotonic time handed in at each boundary, so it
holds no clock and allocates nothing (issue #140). It tracks a shadow of the
send sequence space (snd.UNA/snd.NXT) purely from observed segments, which is
how it manages the timer without reaching into the ControlBlock and how it
distinguishes retransmissions for Karn's algorithm.
Covered:
- §2.2/§2.3 SRTT/RTTVAR/RTO smoothing (integer-shift form)
- §3 Karn's algorithm: one sample in flight, never sample a retransmit
- §5.1-§5.3 timer arm/restart/stop as data is sent and acknowledged
- §5.4-§5.6 timeout response: exponential backoff + go-back-N retransmit
- §5.7 backoff collapse on a valid RTT sample
- RTO clamped to [rtoMin, rtoMax]
RTT introspection (SmoothedRTT, CurrentRTO, Running) lives on the concrete type,
not the interface, per the LossRecovery design. Install with new(RTO) on
ConnConfig.LossRecovery; the connection calls Reset on open so the zero value is
ready to use.
Generated with LLM assistance.
Signed-off-by: Marvin Drees <marvin.drees@9elements.com>
* feat(tcp): add out-of-order segment reassembly
Add an opt-in, bounded out-of-order reassembly buffer so a single lost
segment can be recovered by retransmitting the gap while later segments
are held and delivered once the gap fills.
The receiver also subtracts buffered out-of-order bytes from the
advertised receive window and avoids challenge-ACK aborts for in-window
future data. Reassembly is disabled by default.
Generated with LLM assistance.
Signed-off-by: Marvin Drees <marvin.drees@9elements.com>
* implement review feedback around rx buffer reuse
Signed-off-by: Marvin Drees <marvin.drees@9elements.com>
---------
Signed-off-by: Marvin Drees <marvin.drees@9elements.com>
* fix for #50
* fix remaining test
* revert changes and instead approach problem on full-duplex case
* fix merge messup
* fix @MDr164 note and upgrade documentation while at it
* gate full buffer exit on rx shutdown since no risk of full buffer error
* add log line to note issue is hit
* document CloseRead relationship with Close
* persist issue 57 fix in test
* have go fix output error on CI
* apply go fix required fix
* test no || for go fix command
* test no || for go fix command complete, it is needed
* forgot to fix
- http/httpraw: cap header slice growth to pre-allocated capacity, fix
benchmark to reuse Header across iterations (0 allocs/op)
- internal/ring: replace TODO panic comments with invariant explanations
- tcp/conn: document truncated-frame offset check
- tcp/handler: replace TODO with net.ErrClosed on RST-closed connection
- internet/stack-udpport: filter incoming packets by remote IP address
when configured via SetStackNode; skip filter for IPv4 multicast
destinations (class D) so mDNS and similar protocols work correctly
Generated with LLM assistance.
Signed-off-by: Marvin Drees <marvin.drees@9elements.com>
* add ipv6 to xnet.StackAsync
* dns improvements
* improve DNS workings of StackAsync
* add tentative ICMPv6
* work on prefixes and fix some small bugs, plan UDP/TCP6
* fix bugs in StackAsync and ipv4.Prefix.Contains
* update arpsubtable
* completely remove legacy internet.StackIP for StackIPv4/v6
* ipv4/ipv6 tcp/udp
* add TCP6/UDP6 dialing APIs
* add xnet.Stack6 interface
* more ipv6 integration into StackAsync; various tweaks to lneto and documentation+TODOs
* add stack6 tests
* replace netip.Prefix with ipv4.Prefix where it makes sense
* start working on tracking down tcp buffer bug
* mtu refactor
* modularize test
* more precise testing
* tests fail, but is it the failure we are looking for?
* fix typo in espradio link (#76)
* implement a new backoff abstraction (#75)
* rewrite backoff api
* rewrite tcp.Conn.Write
* keep fixing small things
* much better Conn.Read implementation
* fix critical overflow bug in internal.ConnRWBackoff
---------
Co-authored-by: Joel Wetzell <jwetzell@yahoo.com>
* remove old fuzz corpus and stick to using PCG for seeding fuzz test
* polish up drain logic and reduce action search space
* improve fuzzing debug prints
* more fuzz fixes also formatter capture printer fixes
* refactor StackSeeded test
* fix challenge ack infinite packet bug
* more expressive challenge satisfy
* remove printing
* begin adding icmp client
* rely on anon structs
* add tests
* tests passing
* icmp fleshed out
* rework icmp to include ip addr
* rename StackAsync.Demux/Encapsulate to RecvEthernet and SendEthernet
* remove legacy unreachable TCP tests
* remove uses of deprecated internet.StackNode in preference of lneto.StackNode
* documentation
* rename methods to signal no I/O happening
* tcp: remove retransmit logic entirely; add duplicate ack counting to ControlBlock
* retransmit implemented in nice simple straightforward way
* narrow down retransmission cases
* remove old timing tests
* add ControlBlock retransmit test
* add failing handler test
* reworking payload length semantic meaning in code
* fix establish conn logic
* clean up tests and add TCB dupack generation and test it
* catch pending retransmit satisfy in test
* add fuzz test for control block
* bugfix: be more strict in what is considered dupack
* add IncomingIsDupACK docs
* limit queue of retransmits
* protect retransmit overflow from incorrectly updating nxt
* fix for #58
* add fixes for both retransmit and fuzz failure found
* fix infinite challenge ack due to RST+SYN+FIN corrupt packet in SYNRCV state
* fuzz: previous recovery ack path led to infinite transmit
* fuzz: calculate CRCs when fuzzing to reduce search space to non-CRC error cases
* add local fuzz corpus to tests
* add more cases to fuzz corpus
* work on tracking more heap allocations down
* write own heapless IP AppendFormatAddr functions
* remove potential Error method allocations
* more logging
* tcp: bugfix for seq not in snd/rcv.wnd error
* accept dupe ACKs for window updates and prevent decrement of UNA
* add fixes for incorrect tcp functioning
* remove Conn.Available* methods in favor of Conn.Free* methods due to ambiguous name
* pcap: reuse Frame memory
* slog: reduce heap allocations of addresses; also prevent heap alloc of dhcp options in pcap
* dns: heapless improvement; add StackAsync buffer for more heapless operation; start thinking of errors
* errors: begin standardise errors in lneto
* errors: finish standardization of errors
* fix merge issues
* add more lneto errors to rest of package
* format errors.go
* reduce heap allocations in tcp logging; omit use of AppendFloat which allocates a metric sh*tton
* debugheaplog: better heap statistic logging
* heap: use string for pcap.Frame.Protocol
* add potential to eliminate Flags.String heap alloc, remove incorrect HEAP comments
* add StackAsync.DebugErr and httpraw.SetBytes
* many heap alloc reductions and replacement of bytes.Equal with internal.BytesEqual
* tcp: mss honoring; accept syn with ECE/CWR flags; add RSTQueue type
* tcp: move option logic to own file
* remove prints in pcap
* add MSS send threshold inspired by linux/freebsd/lwip thresh
* claude suggests a way forward
* add timing to capture printer
* add pcap.Flags
* fix ICMP CRC calculation and add test
* bugfix: still send data on half-close state(close-wait)
* fix pcap test