KonstantinSeurer/mesa

Commit Graph

Author	SHA1	Message	Date
Andreas Boll	9246df2280	st/osmesa: Fix a typo in a comment s/suport/support/ Signed-off-by: Andreas Boll <andreas.boll.dev@gmail.com> Reviewed-by: Brian Paul <brianp@vmware.com>	2015-12-09 18:29:18 +01:00
Andreas Boll	7af9930ab4	meta: Fix a typo in a print message s/Unkown/Unknown/ Signed-off-by: Andreas Boll <andreas.boll.dev@gmail.com> Reviewed-by: Brian Paul <brianp@vmware.com>	2015-12-09 18:29:15 +01:00
Andreas Boll	c83e161c91	mesa: Fix typos in print messages s/inconsistant/inconsistent/ s/occurences/occurrences/ Signed-off-by: Andreas Boll <andreas.boll.dev@gmail.com> Reviewed-by: Brian Paul <brianp@vmware.com>	2015-12-09 18:29:11 +01:00
Andreas Boll	5c27cb3da3	glsl: Fix a typo in a comment s/suports/supports/ Signed-off-by: Andreas Boll <andreas.boll.dev@gmail.com> Reviewed-by: Brian Paul <brianp@vmware.com>	2015-12-09 18:26:47 +01:00
Brian Paul	aa9af32752	svga: initialize pipe_driver_query_info entries with a macro To be safe, set all the fields in case the enums ordering/values ever change. Reviewed-by: Charmaine Lee <charmainel@vmware.com>	2015-12-09 09:43:47 -07:00
Brian Paul	ab0651ccfd	mesa: detect inefficient buffer use and report through debug output When a buffer is created with GL_STATIC_DRAW, its contents should not be changed frequently. But that's exactly what one application I'm debugging does. This patch adds code to try to detect inefficient buffer use in a couple places. The GL_ARB_debug_output mechanism is used to report the issue. NVIDIA's driver detects these sort of things too. Other types of inefficient buffer use could also be detected in the future. Reviewed-by: José Fonseca <jfonseca@vmware.com>	2015-12-09 09:43:47 -07:00
Francisco Jerez	595c818071	i965: Resolve color and flush for all active shader images in intel_update_state(). Fixes arb_shader_image_load_store/execution/load-from-cleared-image.shader_test. Couldn't reproduce any significant FPS regression in CPU-bound benchmarks from the Finnish benchmarking system on neither VLV nor BSW after 30 runs with 95% confidence level. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=92849 Cc: Chris Wilson <chris@chris-wilson.co.uk> Cc: Jason Ekstrand <jason.ekstrand@intel.com> Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org> Tested-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 15:12:59 +02:00
Francisco Jerez	3dc97a1586	i965: Document inconsistent units the URB size is represented in. Every other gen the representation of the URB size was changed and previous ones weren't updated. I'd be willing to write a series normalizing this to be KB on all generations if anybody else cares.	2015-12-09 14:00:30 +02:00
Francisco Jerez	228d5a3f75	i965: Hook up L3 partitioning state atom. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:59:03 +02:00
Francisco Jerez	1fc797e8e4	i965: Work around L3 state leaks during context switches. This is going to require some rather intrusive kernel changes to fix properly, in the meantime (and forever on at least pre-v4.1 kernels) we'll have to restore the hardware defaults at the end of every batch in which the L3 configuration was changed to avoid interfering with the DDX and GL clients that use an older non-L3-aware version of Mesa. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net> v2: Optimize look-up of the default configuration by assuming it's the first entry of the L3 config array in order to avoid an FPS regression in GpuTest Triangle and SynMark OglBatch2-7 on most affected platforms. Reviewed-by: Jordan Justen <jordan.l.justen@intel.com>	2015-12-09 13:57:40 +02:00
Francisco Jerez	09d9638dd0	i965: Add debug flag to print out the new L3 state during transitions. Reviewed-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:05 +02:00
Francisco Jerez	acc77947ca	i965: Implement L3 state atom. The L3 state atom calculates the target L3 partition weights when the program bound to some shader stage is modified, and in case they are far enough from the current partitioning it makes sure that the L3 state is re-emitted. v2: Fix for inconsistent units the context URB size is expressed in. Clamp URB size to 1008 KB on SKL due to FF hardware limitation. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Acked-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:05 +02:00
Francisco Jerez	95ad0bd33b	i965: Calculate appropriate L3 partition weights for the current pipeline state. This calculates a rather conservative partitioning of the L3 cache based on the shaders currently bound to the pipeline and whether they use SLM, atomics, images or scratch space. The result is intended to be fine-tuned later on based on other pipeline state. Note that the L3 partitioning calculated for VLV in the non-SLM non-DC case differs from the hardware defaults in that it doesn't include a DC partition and has twice as much RO cache space -- This is an intentional functional change that improves performance in several bandwidth-bound benchmarks on VLV (5% significance): SynMark OglTexFilterAniso by 14.18%, SynMark OglTexFilterTri by 7.15%, Unigine Heaven by 4.91%, SynMark OglShMapPcf by 2.15%, GpuTest Fur by 1.83%, SynMark OglDrvRes by 1.80%, SynMark OglVsTangent by 1.71%, and a few other benchmarks from the Finnish system by less than 1%. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:05 +02:00
Francisco Jerez	fa1300f75e	i965: Implement selection of the closest L3 configuration based on a vector of weights. The input of the L3 set-up code is a vector giving the approximate desired relative size of each partition. This implements logic to compare the input vector against the table of validated configurations for the device and pick the closest compatible one. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:05 +02:00
Francisco Jerez	353abb294b	i965: Define and use REG_MASK macro to make masked MMIO writes slightly more readable. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:05 +02:00
Francisco Jerez	fa043698d2	i965/hsw: Enable L3 atomics. Improves performance of the arb_shader_image_load_store-atomicity piglit test by over 25x (which isn't a real benchmark it's just heavy on atomics -- the improvement in a microbenchmark I wrote a while ago seemed to be even greater). The drawback is one needs to be extra-careful not to hang the GPU (in fact the whole system). A DC partition must have been allocated on L3, the "convert L3 cycle for DC to UC" bit may not be set, the MOCS L3 cacheability bit must be set for all surfaces accessed using DC atomics, and the SCRATCH1 and ROW_CHICKEN3 bits must be kept in sync. A fairly recent kernel is required for the command parser to allow writes to these registers. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Acked-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:05 +02:00
Francisco Jerez	6907175a4f	i965: Implement programming of the L3 configuration. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Acked-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:05 +02:00
Francisco Jerez	b22bebe966	i965: Import tables enumerating the set of validated L3 configurations. It should be possible to use additional L3 configurations other than the ones listed in the tables of validated allocations ("BSpec » 3D-Media-GPGPU Engine » L3 Cache and URB [IVB+] » L3 Cache and URB [*] » L3 Allocation and Programming"), but it seems sensible for now to hard-code the tables in order to stick to the hardware docs. Instead of setting up the arbitrary L3 partitioning given as input, the closest validated L3 configuration will be looked up in these tables and used to program the hardware. The included tables should work for Gen7-9. Note that the quantities are specified in ways rather than in KB, this is because the L3 control registers expect the value in ways, and because by doing that we can re-use a single table for all GT variants of the same generation (and in the case of IVB/HSW and CHV/SKL across different generations) which generally have different L3 way sizes but allow the same combinations of way allocations. v2: Use slice count from the devinfo structure instead of the gt number to implement get_l3_way_size(). Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Acked-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:05 +02:00
Francisco Jerez	a403ad4f5a	i965: Add slice count to the brw_device_info structure. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:05 +02:00
Francisco Jerez	c8ff045fdb	i965/gen8: Don't add workaround bits to PIPE_CONTROL stalls if DC flush is set. According to the hardware docs a DC flush is sufficient to make CS_STALL happy, there's no need to add STALL_AT_SCOREBOARD whenever it's present. Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:05 +02:00
Francisco Jerez	2405b75bc9	i965: Define state flag to signal that the URB size has been altered. This will make sure that we recalculate the URB layout anytime the URB size is modified by the L3 partitioning code. Reviewed-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:04 +02:00
Francisco Jerez	4841cab01a	i965: Keep track of whether LRI is allowed in the context struct. This stores the result of can_do_pipelined_register_writes() in the context struct so we can find out later whether LRI can be used to program the L3 configuration. v2: * Split change of gen check in can_do_pipelined_register_writes (jljusten) Reviewed-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:04 +02:00
Francisco Jerez	50c2713726	i965: Adjust gen check in can_do_pipelined_register_writes Allow for pipelined register writes for gen < 7. v2: * Split from another patch and adjust comment (jljusten) Reviewed-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:04 +02:00
Francisco Jerez	5912da45a6	i965: Define symbolic constants for some useful L3 cache control registers. Reviewed-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net>	2015-12-09 13:46:04 +02:00
Dave Airlie	e307cfa7d9	radeonsi: handle loading doubles as geometry shader inputs. This adds the double code to the geometry shader input handling. Reviewed-by: Michel Dänzer <michel.daenzer@amd.com> Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-09 17:04:04 +10:00
Dave Airlie	8c9e40ac22	radeonsi: handle doubles in lds load path. This handles loading doubles from LDS properly. Reviewed-by: Michel Dänzer <michel.daenzer@amd.com> Cc: "11.0 11.1" <mesa-stable@lists.fedoraproject.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-09 17:03:38 +10:00
Dave Airlie	cce3864046	r600: handle geometry dynamic input array index This fixes: glsl-1.50/execution/geometry/dynamic_input_array_index.shader_test my profanity. We need to load the AR register with the value from the index reg Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-09 15:07:53 +10:00
Dave Airlie	38542921c7	r600g: fix geom shader input indirect indexing. This fixes: gs-input-array-vec4-index-rd The others run out of gprs unfortunately. Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-09 15:07:47 +10:00
Dave Airlie	e97ac006d7	r600g: fix outputing to non-0 buffers for stream 0. This fixes: arb_transform_feedback3-ext_interleaved_two_bufs_gs arb_transform_feedback3-ext_interleaved_two_bufs_gs_max transform-feedback-builtins If we are only emitting one ring, then emit all output buffers on it. Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-09 15:07:01 +10:00
Edward O'Callaghan	1f61447ce1	r600: Add ARB_copy_image support [airlied: update relnotes] Reviewed-by: Marek Olšák <marek.olsak@amd.com> Signed-off-by: Edward O'Callaghan <eocallaghan@alterapraxis.com> Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-09 14:41:46 +10:00
Edward O'Callaghan	d13ac27200	r600g: allow copying between compatible un/compressed formats See: `commit e82c527f1fc2f8ddc64954ecd06b0de3cea92e93` which is where a block in src maps to a pixel in dst and vice versa. e.g. DXT1 <-> R32G32_UINT DXT5 <-> R32G32B32A32_UINT Reviewed-by: Marek Olšák <marek.olsak@amd.com> Signed-off-by: Edward O'Callaghan <eocallaghan@alterapraxis.com> Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-09 14:40:32 +10:00
Ilia Mirkin	f920f8eb02	nv50/ir: fix cutoff for using r63 vs r127 when replacing zero The only effect here is a space savings - 822 programs in shader-db affected with the following overall change: total bytes used in shared programs : 44154976 -> 44139880 (-0.03%) Fixes: `641eda0c` (nv50/ir: r63 is only 0 if we are using less than 63 registers) Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org>	2015-12-08 23:15:29 -05:00
Ilia Mirkin	44260d9080	nv50/ir: prefer to color mad def and src2 with the same color This allows us to use the short encoding, and potentially fold immediates in later on. total instructions in shared programs : 6379731 -> 6367861 (-0.19%) total gprs used in shared programs : 728502 -> 728683 (0.02%) total local used in shared programs : 9904 -> 9904 (0.00%) total bytes used in shared programs : 44661008 -> 44154976 (-1.13%) local gpr inst bytes helped 0 51 7267 20306 hurt 0 232 125 274 Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-08 23:15:29 -05:00
Ilia Mirkin	c1c1248b94	nv50/ir: reduce degree limit on ops that can't encode large reg dests Operations that take immediates can only encode registers up to 64. This fixes a shader in a "Powered by Unity" intro. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-08 23:15:29 -05:00
Ilia Mirkin	99581ca393	nv50/ir: only unspill once ahead of a group of instructions We already semi-did this but the list of uses as unsorted, so it was unreliable. Sort the uses by bb and serial, and don't unspill for each instruction in a sequence. (And also don't unspill multiple times for a single instruction that uses the value in question multiple times.) This causes a minor reduction in generated instructions for shader-db (as few programs spill) but more importantly it brings determinism to each run's output. On SM10: total instructions in shared programs : 6387945 -> 6379359 (-0.13%) total gprs used in shared programs : 728544 -> 728544 (0.00%) total local used in shared programs : 9904 -> 9904 (0.00%) local gpr inst bytes helped 0 0 322 322 hurt 0 0 0 0 Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-08 23:15:29 -05:00
Ilia Mirkin	0f647bd65b	nv50/ir: check if the target supports the new offset before inlining Fixes: `abd326e81b` (nv50/ir: propagate indirect loads into instructions) Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=93300 Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-08 23:15:29 -05:00
Dave Airlie	a13b14930d	llvmpipe: fix fp64 inputs to geom shader. This fixes the fetching of fp64 inputs to the geometry shader, this fixes the recently posted piglit's arb_gpu_shader_fp64/execution/gs-fs-vs-double-array.shader_test arb_vertex_attrib_64bit/execution/gs-fs-vs-attrib-double-array.shader_test Reviewed-by: Roland Scheidegger <sroland@vmware.com> Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-09 13:56:39 +10:00
Matt Turner	3a7f95b3aa	nir: Optimize useless comparisons against true/false. Reviewed-by: Jason Ekstrand <jason.ekstrand@intel.com> [v1] Reviewed-by: Eric Anholt <eric@anholt.net> [v1] v2: Move new rule to Boolean simplification section Add a a@bool != true simplification Suggested-by: Neil Roberts <neil@linux.intel.com>	2015-12-08 15:41:08 -08:00
Matt Turner	9e9e6fc8f1	glsl: Switch opcode and avail parameters to binop(). To make it match unop(). Reviewed-by: Ian Romanick <ian.d.romanick@intel.com>	2015-12-08 15:39:47 -08:00
Matt Turner	dd3c16c94b	glsl_to_tgsi: Skip useless comparison instructions. Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-08 15:38:03 -08:00
Matt Turner	eca846e7ae	glsl: Relax qualifier ordering restriction in ES 3.1. ... and allow the "binding" qualifier in ES 3.1 as well. GLSL ES 3.1 incorporates only a few features from the extension ARB_shading_language_420pack: the relaxed qualifier ordering requirements and the binding qualifier. Cc: "11.1" <mesa-stable@lists.freedesktop.org> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2015-12-08 15:36:57 -08:00
Matt Turner	79da7220db	glsl: Use has_420pack(). These features would not have been enabled with #version 420 otherwise. Cc: "11.1" <mesa-stable@lists.freedesktop.org> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-08 15:36:57 -08:00
Matt Turner	c200e606f7	glsl: Allow binding of image variables with 420pack. This interaction was missed in the addition of ARB_image_load_store. Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=93266 Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu> Reviewed-by: Ian Romanick <ian.d.romanick@intel.com>	2015-12-08 15:36:57 -08:00
Eric Anholt	f61ceeb3fd	vc4: Enable MSAA. We still have several failures in the newly enabled tests in simulation: sRGB downsampling is done as if it was just linear, stencil blits are not supported on MSAA either, and derivatives are still not supported (breaking some MSAA simulation shaders). So, other than sRGB downsampling quality, things seem to be in good shape.	2015-12-08 10:09:52 -08:00
Eric Anholt	fc4a1bfb88	vc4: Add support for mapping of MSAA resources. The pipe_transfer_map API requires that we do an implicit downsample/upsample and return a mapping of that.	2015-12-08 09:49:56 -08:00
Eric Anholt	6b4dfd53ae	vc4: Add support for texel fetches from MSAA resources. This is the core of ARB_texture_multisample. Most of the piglit tests for GL_ARB_texture_multisample require GL 3.0, but exposing support for this lets us use the gallium blitter for multisample resolves. We can sometimes multisample resolve using just the RCL, but that requires that the blit is 1:1, unflipped, and aligned to tile boundaries.	2015-12-08 09:49:55 -08:00
Eric Anholt	a97b40dca4	vc4: Add support for multisample framebuffer operations. This includes GL_SAMPLE_COVERAGE, GL_SAMPLE_ALPHA_TO_ONE, and GL_SAMPLE_ALPHA_TO_COVAGE. I haven't implemented a dithering function yet, and gallium doesn't give me a good chance to do so for GL_SAMPLE_COVERAGE.	2015-12-08 09:49:54 -08:00
Eric Anholt	edc3305de7	vc4: Add a workaround for HW-2905, and additional failure I saw with MSAA. I only stumbled on this while experimenting due to reading about HW-2905. I don't know if the EZ disable in the Z-clear is actually necessary, but go with it for now.	2015-12-08 09:49:54 -08:00
Eric Anholt	edfd4d853a	vc4: Add support for drawing in MSAA.	2015-12-08 09:49:53 -08:00
Eric Anholt	e7c8ad0a6c	vc4: Add kernel RCL support for MSAA rendering.	2015-12-08 09:49:53 -08:00
Eric Anholt	568d3a8e32	vc4: Rename color_ms_write to color_write. I was thinking this was the only MSAA resolve thing, so it should be noted separately, but actually load/store general also do MSAA resolve.	2015-12-08 09:49:52 -08:00
Eric Anholt	bf92017ace	vc4: Allow RCL blits to the edge of the surface. The recent unaligned fix successfully prevented RCL blits that weren't aligned inside of the surface, but we also want to be able to do RCL blits for the whole surface when the width or height of the surface aren't aligned (we don't care what renders inside of the padding).	2015-12-08 09:49:52 -08:00
Eric Anholt	fb4877dbab	vc4: Add disabled debug printf for describing blits. I keep typing variants of this while debugging RCL blits for MSAA.	2015-12-08 09:49:51 -08:00
Eric Anholt	2792d118f1	vc4: Fix check for tile RCL blits with mismatched y. This was a typo in `3a508a0d94` that didn't show up in testcases at that moment.	2015-12-08 09:49:51 -08:00
Eric Anholt	1529f138ff	vc4: Fix compiler warning from size_t change. I missed this when bringing over the kernel changes.	2015-12-08 09:49:50 -08:00
Jason Ekstrand	18069dce4a	i965: Make uniform offsets be in terms of bytes This commit pushes makes uniform offsets be terms of bytes starting with nir_lower_io. They get converted to be in terms of vec4s or floats when we cram them in the UNIFORM register file but reladdr remains in terms of bytes all the way down to the point where we lower it to a pull constant load. Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2015-12-07 21:51:23 -08:00
Jason Ekstrand	813f0eda8e	i965/nir_uniforms: Replace comps_per_unit with an is_scalar boolean Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2015-12-07 21:51:23 -08:00
Jason Ekstrand	22c273de2b	i965/nir: Remove unused indirect handling The one and only place where the FS backend allows reladdr is on uniforms. For locals, inputs, and outputs, we lower it away before the backend ever sees it. This commit gets rid of the dead indirect handling code. Cc: "11.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2015-12-07 21:51:23 -08:00
Jason Ekstrand	abb569ca18	i965/state: Get rid of dword_pitch arguments to buffer functions Cc: "11.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2015-12-07 21:51:23 -08:00
Jason Ekstrand	05bdc21f84	i965/vec4: Use a stride of 1 and byte offsets for UBOs Cc: "11.0" <mesa-stable@lists.freedesktop.org> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=92909 Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2015-12-07 21:51:23 -08:00
Jason Ekstrand	13ad8d03f2	i965/fs: Use a stride of 1 and byte offsets for UBOs Cc: "11.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2015-12-07 21:51:23 -08:00
Jason Ekstrand	e3e70698c3	i965/vec4: Use byte offsets for UBO pulls on Sandy Bridge Previously, the VS_OPCODE_PULL_CONSTANT_LOAD opcode operated on vec4-aligned byte offsets on Iron Lake and below and worked in terms of vec4 offsets on Sandy Bridge. On Ivy Bridge, we add a new *LOAD_GEN7 variant which works in terms of vec4s. We're about to change the GEN7 version to work in terms of bytes, so this is a nice unification. Cc: "11.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2015-12-07 21:51:23 -08:00
Ben Widawsky	6ef8149bcd	i965: Fix texture views of 2d array surfaces It is legal to have a texture view of a single layer from a 2D array texture; you can sample from it, or render to it. Intel hardware needs to be made aware when it is using a 2d array surface in the surface state. The texture view is just a 2d surface with the backing miptree actually being a 2d array surface. This caused the previous code would not set the right bit in the surface state since it wasn't considered an array texture. I spotted this early on in debug but brushed it off because it is clearly not needed on other platforms (since they all pass). I have no idea how this works properly on other platforms (I think gen7 introduced the bit in the state, but I am too lazy to check). As such, I have opted not to modify gen7, though I believe the current code is wrong there as well. Thanks to Chris for helping me debug this. v2: Just use the underlying mt's target type to make the array determination. This replaces a bug in the first patch which was incorrectly relying only on non-zero depth (not sure how that had no failures). (Ilia) Cc: Chris Forbes <chrisf@ijw.co.nz> Reported-by: Mark Janes <mark.a.janes@intel.com> (Jenkins) References: https://www.opengl.org/registry/specs/ARB/texture_view.txt Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=92609 Signed-off-by: Ben Widawsky <benjamin.widawsky@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2015-12-07 18:47:04 -08:00
Nicolai Hähnle	d5a5dbd71f	radeonsi: last_gfx_fence is a winsys fence Cc: "11.1" <mesa-stable@lists.freedesktop.org> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2015-12-07 21:15:59 -05:00
Ilia Mirkin	f97f755192	nvc0/ir: fix up mul+add -> mad algebraic opt, enable for integers For some reason this has been disabled for integers ever since codegen was merged, despite there being emission code for IMAD. Seems to work. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-07 18:49:28 -05:00
Ilia Mirkin	1d708aacb7	gk110/ir: fix imad sat/hi flag emission for immediate args According to nvdisasm both the immediate and non-imm cases use the same bits. Both of these flags are quite rarely set though. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org>	2015-12-07 18:49:28 -05:00
Kenneth Graunke	87a1166310	i965: Add brw_device_info::min_ds_entries field. From the 3DSTATE_URB_DS documentation: "Project: IVB, HSW If Domain Shader Thread Dispatch is Enabled then the minimum number of handles that must be allocated is 10 URB entries." "Project: BDW+ If Domain Shader Thread Dispatch is Enabled then the minimum number of handles that must be allocated is 34 URB entries." When the HS is run in SINGLE_PATCH mode (the only mode we support today), there is no minimum for HS - it's just zero. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-12-07 14:48:55 -08:00
Chris Forbes	42ca675cc9	i965: Add state bits for tess stages Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-12-07 14:48:55 -08:00
Chris Forbes	80ea18d1a1	i965: Add backend structures for tess stages Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-12-07 14:48:55 -08:00
Chris Forbes	5340f37902	i965: Set core tessellation-related limits Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-12-07 14:48:54 -08:00
Kenneth Graunke	a9e6a56a02	i965: Request lowering of gl_TessLevel* from float[] to vec4s. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-12-07 14:48:54 -08:00
Kenneth Graunke	7a17356800	i965: Create new files for HS/DS/TE state upload code. For now, this just splits the existing code to disable these stages into separate atoms/files. We can then replace it with real code. v2: Bump the render atoms in this patch so it compiles (in my branch, I'd bumped it in an earlier patch). 61 seems to be the minimum that works, which doesn't match the old value + the number of atoms I added in this patch, so apparently we had some slop before. v3: Actually disable the DS unit on Gen8+. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Kristian Høgsberg <krh@bitplanet.net> [v1] Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-12-07 14:48:54 -08:00
Ilia Mirkin	63b850403c	gk104/ir: sampler doesn't matter for txf We actually leave the sampler unset for OP_TXF, which caused the GK104+ logic to treat some texel fetches as indirect. While this works, it's incredibly wasteful. This only happened when the texture was > 0 (since sampler remained == 0). Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org>	2015-12-07 16:22:54 -05:00
Marek Olšák	32f05fadbb	radeonsi: disable DCC on Stoney Cc: 11.1 <mesa-stable@lists.freedesktop.org> Reviewed-by: Alex Deucher <alexander.deucher@amd.com>	2015-12-07 22:01:08 +01:00
Sonny Jiang	2618886600	winsys/amdgpu: addrlib - port a Fiji bug fix Fiji: Fixed tiled resource failures Signed-off-by: Sonny Jiang <sonny.jiang@amd.com> Reviewed-by: Alex Deucher <alexander.deucher@amd.com> Reviewed-by: Michel Dänzer <michel.daenzer@amd.com> v2: fix a compile failure (typo) - Marek	2015-12-07 21:58:42 +01:00
Sonny Jiang	338d7bf053	winsys/amdgpu: addrlib - port Checks mip 0 for czDispCompatible Signed-off-by: Sonny Jiang <sonny.jiang@amd.com> Reviewed-by: Alex Deucher <alexander.deucher@amd.com> Reviewed-by: Michel Dänzer <michel.daenzer@amd.com>	2015-12-07 21:58:42 +01:00
Sonny Jiang	676bc25140	winsys/amdgpu: addrlib - port fix error for workaround for 1D tiling Signed-off-by: Sonny Jiang <sonny.jiang@amd.com> Reviewed-by: Alex Deucher <alexander.deucher@amd.com> Reviewed-by: Michel Dänzer <michel.daenzer@amd.com>	2015-12-07 21:58:42 +01:00
Christian König	a2c5200a4b	st/va: disable MPEG4 by default v2 The workarounds are too hacky to enable them by default and otherwise MPEG4 doesn't work reliably. v2: add docs/envvars.html, CC stable and fix typos Signed-off-by: Christian König <christian.koenig@amd.com> Reviewed-by: Emil Velikov <emil.l.velikov@gmail.com> (v1) Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu> (v1) Cc: "11.1.0" <mesa-stable@lists.freedesktop.org>	2015-12-07 20:34:17 +01:00
Christian König	ca3e2b76c0	st/va: move HEVC functions into separate file v2 v2: actually copy all of it Signed-off-by: Christian König <christian.koenig@amd.com>	2015-12-07 20:34:17 +01:00
Alejandro Piñeiro	3d260cc653	mesa: remove _mesa_tex_target_is_array _mesa_is_array_texture provides the same functionality and: 1. it returns bool instead of GLboolean 2. it's not related to the texture format (texformat.c) 3. the name's a little shorter v2: remove _mesa_tex_target_is_array instead (Brian Paul) Reviewed-by: Brian Paul <brianp@vmware.com>	2015-12-07 20:31:20 +01:00
Alejandro Piñeiro	b16e0ff34e	i965: use _mesa_is_array_texture instead of _mesa_tex_target_is_array Both methods provide the same functionality, so one would be removed. v2: use _mesa_is_array_texture and not the other way (Brian Paul) Reviewed-by: Brian Paul <brianp@vmware.com>	2015-12-07 20:30:24 +01:00
Ilia Mirkin	db072d2086	gk110/ir: fix imul hi emission with limm arg The elemental demo hits this case. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org>	2015-12-07 13:30:17 -05:00
Brian Paul	32a6e081c3	svga: use the debug callback to report issues to the state tracker Use the new debug callback hook to report conformance, performance and fallbacks to the state tracker. The state tracker, in turn can report this issues to the user via the GL_ARB_debug_output extension. More issues can be reported in the future; this is just a start. v2: remove conditionals around pipe_debug_message() calls since the check is now done in the macro itself. v3: remove unneeded dummy %s substitutions Acked-by: Ilia Mirkin <imirkin@alum.mit.edu>, Reviewed-by: José Fonseca <jfonseca@vmware.com>	2015-12-07 08:57:49 -07:00
Brian Paul	5effc3ae74	gallium/util: check callback pointers for non-null in pipe_debug_message() So the callers don't have to do it. v2: also check cb!=NULL in the macro Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu> Reviewed-by: José Fonseca <jfonseca@vmware.com>	2015-12-07 08:56:51 -07:00
Abdiel Janulgue	b19546abf3	i965: Add defines for gather push constants v2 (Francisco Jerez): - Rename HSW_GATHER_CONSTANTS_RESERVED to HSW_GATHER_POOL_ALLOC_MUST_BE_ONE. - Rename BRW_GATHER_* prefix to HSW_GATHER_CONSTANT_*. Reviewed-by: Francisco Jerez <currojerez@riseup.net> Signed-off-by: Abdiel Janulgue <abdiel.janulgue@linux.intel.com>	2015-12-07 14:58:12 +02:00
Timothy Arceri	9214664aed	mesa: move GLES checks for SSO input/output validation This function is unfinished there is a bunch more validation rules that need to be applied here. We will still want to call it for desktop GL we just don't want to validate precision so move the ES check to reflect this. Reviewed-by: Tapani Pälli <tapani.palli@intel.com>	2015-12-07 21:41:14 +11:00
Timothy Arceri	ad02621854	mesa: move GL_INVALID_OPERATION error to rendering call The validation api doesn't trigger this error so just move it to the code called during rendering. Reviewed-by: Tapani Pälli <tapani.palli@intel.com> Cc: Kenneth Graunke <kenneth@whitecape.org>	2015-12-07 21:41:09 +11:00
Timothy Arceri	4dd096d741	mesa: move pipeline input/output validation inside _mesa_validate_program_pipeline() This allows validation to be done on rendering calls also. Fixes 3 dEQP-GLES31.functional.separate tests. Cc: "11.1" <mesa-stable@lists.freedesktop.org> Reviewed-by: Tapani Pälli <tapani.palli@intel.com> Cc: Kenneth Graunke <kenneth@whitecape.org>	2015-12-07 21:41:05 +11:00
Timothy Arceri	da1a01361b	glsl: re-validate program pipeline after sampler change Cc: "11.1" <mesa-stable@lists.freedesktop.org> Reviewed-by: Tapani Pälli <tapani.palli@intel.com> Cc: Kenneth Graunke <kenneth@whitecape.org> https://bugs.freedesktop.org/show_bug.cgi?id=93180	2015-12-07 21:41:00 +11:00
Dave Airlie	41e82f4f96	r600: apply SIMD workaround to cayman also. At last on ARUBA this is required to stop tessellation hanging in heaven. This removes one of the SIMDs from use by the HS/LS. Reviewed-by: Edward O'Callaghan <eocallaghan@alterapraxis.com> Tested-by: Edward O'Callaghan <eocallaghan@alterapraxis.com> Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-07 18:57:34 +10:00
Dave Airlie	6bf6bdbc2b	r600: fix regression introduced with ring emit changes. This was adding one after a CUT which broke end primitive	2015-12-07 05:44:55 +00:00
Dave Airlie	fc276bda22	r600: remove stale tessellation comment pointed out by Marek. Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-07 11:04:48 +10:00
Dave Airlie	33404f1415	r600: enable tessellation for evergreen/cayman (v2) This enables tessellation for evergreen/cayman, This will need changes before committing depending on what hw works etc. working are CAYMAN/REDWOOD/BARTS/TURKS/SUMO/CAICOS v2: only enable on evergreen and above.	2015-12-07 09:59:02 +10:00
Dave Airlie	a2885d9cf9	r600g: reduce number of ps thread on caicos this allows tess apps to start Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-07 09:59:02 +10:00
Dave Airlie	fe64a0c8bf	r600g: adjust ls/hs thread counts for sumo these stop tess hangs here. Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-07 09:59:02 +10:00
Dave Airlie	e7ce9e3bb8	r600/asm: enable nstack check for tess ctrl/eval shaders. This just makes sure they register at least one stack usage frame like vertex shaders. Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-07 09:59:02 +10:00
Dave Airlie	bb44c1f036	r600/asm: handle lds read operations. Reads from the queue shouldn't be merged for now read operations. Reads from the queue shouldn't be merged for now, or put in T slots. Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-07 09:59:02 +10:00
Dave Airlie	8ec2cb13e5	r600/asm: add LDS ops and barrier to the once per group restriction. LDS ops must be scheduled in X slot, and barrier should be on its own in a group. Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-07 09:59:02 +10:00
Dave Airlie	18871ac576	r600: move VGT_VTX_CNT_EN into shader stages atom. This should be enabled for tessellation shaders as well. Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-07 09:59:02 +10:00
Dave Airlie	958d617d98	r600: enable tcs/tes dumping for R600_DUMP_SHADERS. Trivial patch just to enable dumping more. Signed-off-by: Dave Airlie <airlied@redhat.com>	2015-12-07 09:59:01 +10:00

1 2 3 4 5 ...

68339 Commits